Skip to main content
Updated April 2026 20 min read For developers and advanced users

Guide Claude AI the complete
Models, Transformers, Inference API & Spaces

Everything you need to know about HuggingFace — the open-source AI platform — its models, the Transformers library step by step, the free Inference API, and how to build AI apps.

200K
Context Window
Constitutional
AI Safety
#1
Coding Benchmarks

What is Hugging Face?

Hugging Face is the main hub for open machine learning — often described as the GitHub of AI. It hosts hundreds of thousands of models and datasets along with the libraries and hosting to run them.

Those are five separate products with different costs, limits and failure modes, and most confusion about Hugging Face comes from treating them as one thing. A model that is free to download is not free to serve; a demo that works in a Space will not survive production traffic.

Read the licence before you read the benchmark

This is the section that saves people real money and it is the one most guides skip entirely.

"Open" on the Hub covers a wide range of terms, and the download button does not check any of them. A model's licence is a field on its page, and it can be:

Two things to check beyond the label. Fine-tunes inherit: a model trained on top of a restricted base carries the base's terms whatever the uploader wrote, and the Hub is full of derivatives whose licence field is optimistic. And the dataset has its own licence, which matters if you train on it.

Practical habit: before anything reaches a product, record the model name, the exact revision, the licence and where you read it. That file takes minutes and is the difference between answering a question and starting an investigation.

Quick start

Install the library and run a model in a couple of minutes:

# installation
pip install transformers torch

# basic usage — an extractive question-answering model
from transformers import pipeline

qa = pipeline("question-answering", model="deepset/roberta-base-squad2")
result = qa(question="Where is the Louvre?",
            context="The Louvre is a museum located in Paris, France.")
print(result["answer"])  # Paris

The first call downloads the weights, which is why it is slow and why your disk fills up later — see the caching section.

Choosing a model, properly

Sorting by downloads tells you what is popular, which correlates with being in a tutorial rather than with being right for you. A model page answers better questions if you know what to look for:

Then run your own small evaluation. Twenty examples from your actual data, scored by you, will separate two candidate models better than any leaderboard — because the leaderboard measures something else.

Will it run? The arithmetic

The most common beginner question has a simple answer. Memory needed for the weights is roughly parameters × bytes per parameter:

Add headroom on top for the context and the runtime — in practice assume you need meaningfully more than the weights alone, and more again for long inputs. If a model does not fit, the options in order of ease are: a quantised version (someone has usually published one), a smaller model from the same family, or renting a GPU.

Quantisation costs some quality, and how much depends on the model and the task. It is usually a far better trade than people expect — a 4-bit larger model commonly beats a 16-bit smaller one at the same memory budget.

Where to actually run it

Four options, with the trade-offs that matter:

Tokens, and keeping them out of your code

The API examples you will find online, including older versions of this page, paste the token into the source. Do not.

import os, requests

API_URL = "https://api-inference.huggingface.co/models/mistralai/Mistral-7B-Instruct-v0.1"
headers = {"Authorization": f"Bearer {os.environ['HF_TOKEN']}"}

def query(payload):
    r = requests.post(API_URL, headers=headers, json=payload, timeout=60)
    r.raise_for_status()
    return r.json()

print(query({"inputs": "Explain machine learning in two sentences."}))

Two habits worth adopting at the same time. Create fine-grained tokens scoped to what the job needs — read-only for pulling a model, write only where you genuinely publish — rather than one token that can do everything to everything. And if a token ever reaches a commit, treat it as compromised and rotate it; removing the line does not help, because the history keeps it.

The security setting nobody reads

Two things on the Hub execute or load code from a stranger's repository, and both deserve a deliberate decision.

trust_remote_code=True tells the library to run Python that ships with the model repo. Plenty of legitimate models need it, because their architecture is not in the library yet. It is also arbitrary code execution on your machine, from a repository anyone can create. Use it only for models from an organisation you have a reason to trust, and pin the revision so the code cannot change under you.

Weight formats. Older checkpoints ship as Python pickles, which can execute code when loaded. safetensors exists precisely to avoid that and is now the norm — prefer it, and be suspicious of a repository that only offers the old format.

The same caution applies to datasets with loading scripts. None of this means the Hub is dangerous; it means it is a public registry, and public registries are treated the way you treat any dependency you did not write.

Pin the revision, or your output will change

Model repositories are mutable. Weights get updated, tokenisers get fixed, files get renamed, and occasionally a model is removed or made gated. Code that names only the model will silently pick up whatever is there today.

from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "mistralai/Mistral-7B-Instruct-v0.1"
REV  = "commit-sha-from-the-model-page"   # pin it

tok   = AutoTokenizer.from_pretrained(REPO, revision=REV)
model = AutoModelForCausalLM.from_pretrained(REPO, revision=REV)

This is the single most useful line in this guide for anything running in production. Without it, a Tuesday morning where the model behaves differently has no explanation and no way back.

Where the disk went

Models are large and the cache is silent about it. A few weeks of experimenting produces tens of gigabytes in a directory you have never looked at.

# where the cache lives
export HF_HOME=/path/with/space

# see what is there, and clear what is not needed
pip install "huggingface_hub[cli]"
huggingface-cli scan-cache
huggingface-cli delete-cache

Worth setting HF_HOME deliberately on any shared or space-constrained machine before the first download rather than after the disk fills.

Key capabilities

The smaller task-specific models are the underrated part of this list. A compact classifier or embedding model, running locally in milliseconds for nothing, beats calling a large general model for jobs like routing, tagging or deduplication — and those jobs are most of what production systems actually do.

Evaluating two candidates without a research budget

Leaderboards answer a question you did not ask. They rank general capability on public benchmarks, and public benchmarks leak into training data over time, which is part of why a model that tops a table can disappoint on your material. The evaluation that decides anything is small, private and specific to your task.

It takes an afternoon and it is the most valuable afternoon in the project. Collect twenty to fifty real inputs from your own data, covering the ordinary cases and the awkward ones you know about. Write down what a correct output looks like for each — not a score, a description, since for most tasks there is no single right answer and you will be judging rather than matching strings.

Then run every candidate over the same set, save the outputs, and read them side by side without knowing which model produced which. That last part matters more than it sounds: knowing which is the popular model changes how generously you read its answers.

What you are looking for is rarely average quality, because the candidates are usually close. It is the shape of the failures. One model may be slightly worse overall but wrong in ways you can detect and handle; another may be marginally better and wrong in ways that look exactly like being right. For anything going near a customer, the first is the better system even when it loses on the average.

Keep the set. When you change model, change version, or change a prompt in six months, running it again takes minutes and tells you whether you improved anything — which is otherwise unknowable.

The costs that arrive later

Open weights are free to download, and people reasonably conclude the approach is cheap. Some of it is; these are the parts that are not, and knowing them early changes architecture decisions.

The honest summary: self-hosting wins clearly on privacy, on control, and at sustained high volume. Below that, it frequently loses on total cost while feeling cheaper, because the licence fee was the only number anybody compared.

Fine-tuning: usually not first

Fine-tuning is the most requested and least often correct answer. Before it, in order:

  1. A better prompt, with examples. Free, immediate, and enough surprisingly often.
  2. Retrieval, if the problem is that the model does not know your facts. Fine-tuning teaches behaviour, not knowledge — training on your documents is a poor way to make a model recall them.
  3. A different or larger model. Often cheaper than a training run plus the serving of a custom model.
  4. Then fine-tune, when you need a consistent format, tone or a narrow classification, and you have the data.

If you do: parameter-efficient methods such as LoRA train a small adapter rather than the whole model, which fits on modest hardware and produces a file of megabytes instead of gigabytes. Data quality dominates quantity — a few hundred carefully checked examples typically beat tens of thousands of scraped ones. And hold out a test set before you start, because a fine-tune always looks excellent on the data it was trained on.

Practical uses

Spaces, and what free means

Spaces host a demo app for nothing, which is genuinely useful for showing work and for internal tools. The free tier is CPU-only, it sleeps when idle and takes time to wake, and it is public unless you make it private.

Two practical notes: put any key in the Space's secrets rather than in the repository, and remember that a public Space with an API key in its code is a key you have given away. If a demo becomes something people rely on, move it off the free tier rather than discovering its limits during a client call.

Quick tips

Ready to start?

Sign up to HuggingFace for free and access hundreds of thousands of ready-to-use AI models.