Guide Claude AI the complete
Models, Transformers, Inference API & Spaces
Everything you need to know about HuggingFace — the open-source AI platform — its models, the Transformers library step by step, the free Inference API, and how to build AI apps.
What is Hugging Face?
Hugging Face is the main hub for open machine learning — often described as the GitHub of AI. It hosts hundreds of thousands of models and datasets along with the libraries and hosting to run them.
- Model Hub — the repository: Llama, Mistral, Whisper, Flux and a very long tail.
- Transformers — the Python library that loads and runs most of them in a few lines.
- Spaces — hosted demo apps (Gradio, Streamlit), free at small scale.
- Inference providers and endpoints — calling models over HTTP without owning a GPU.
- Datasets — data for evaluation and fine-tuning.
Those are five separate products with different costs, limits and failure modes, and most confusion about Hugging Face comes from treating them as one thing. A model that is free to download is not free to serve; a demo that works in a Space will not survive production traffic.
Read the licence before you read the benchmark
This is the section that saves people real money and it is the one most guides skip entirely.
"Open" on the Hub covers a wide range of terms, and the download button does not check any of them. A model's licence is a field on its page, and it can be:
- Permissive — Apache 2.0 or MIT. Commercial use, modification, redistribution. No practical constraints for most products.
- A community or bespoke licence — several major model families ship under their own terms, which typically permit commercial use but add conditions: acceptable-use rules, attribution requirements, naming obligations for derivatives, and sometimes a threshold above which you need a separate agreement.
- Non-commercial — often a Creative Commons variant with NC. Fine for a prototype, not for anything that earns money, and this catches people late.
- Research only, or weights released with no licence at all — which means you have no permission, not that permission is implied.
Two things to check beyond the label. Fine-tunes inherit: a model trained on top of a restricted base carries the base's terms whatever the uploader wrote, and the Hub is full of derivatives whose licence field is optimistic. And the dataset has its own licence, which matters if you train on it.
Practical habit: before anything reaches a product, record the model name, the exact revision, the licence and where you read it. That file takes minutes and is the difference between answering a question and starting an investigation.
Quick start
Install the library and run a model in a couple of minutes:
# installation
pip install transformers torch
# basic usage — an extractive question-answering model
from transformers import pipeline
qa = pipeline("question-answering", model="deepset/roberta-base-squad2")
result = qa(question="Where is the Louvre?",
context="The Louvre is a museum located in Paris, France.")
print(result["answer"]) # Paris
The first call downloads the weights, which is why it is slow and why your disk fills up later — see the caching section.
Choosing a model, properly
Sorting by downloads tells you what is popular, which correlates with being in a tutorial rather than with being right for you. A model page answers better questions if you know what to look for:
- Base or instruction-tuned? A base model completes text; it does not follow instructions. Most people reaching for an LLM want the instruct or chat variant, and the difference explains a lot of disappointing first experiments.
- Size, in parameters, which decides whether you can run it at all. See the arithmetic below.
- When was it last updated, and is the repository maintained or abandoned mid-experiment?
- Does the model card describe the training data and the intended use? A card with neither is a model nobody will be able to explain if it behaves oddly.
- Are the numbers on the card independently verifiable? Self-reported benchmarks on an uploader's own page are marketing, not evaluation.
- Language coverage. Most open models are heavily English-weighted, and performance in other languages varies far more than the headline scores suggest. Test on your own text rather than trusting a multilingual label.
Then run your own small evaluation. Twenty examples from your actual data, scored by you, will separate two candidate models better than any leaderboard — because the leaderboard measures something else.
Will it run? The arithmetic
The most common beginner question has a simple answer. Memory needed for the weights is roughly parameters × bytes per parameter:
- 16-bit (the common default) — 2 bytes each. A 7-billion-parameter model is about 14 GB.
- 8-bit — about 7 GB for the same model.
- 4-bit — about 3.5 GB, which is what makes a 7B model run on ordinary consumer hardware.
Add headroom on top for the context and the runtime — in practice assume you need meaningfully more than the weights alone, and more again for long inputs. If a model does not fit, the options in order of ease are: a quantised version (someone has usually published one), a smaller model from the same family, or renting a GPU.
Quantisation costs some quality, and how much depends on the model and the task. It is usually a far better trade than people expect — a 4-bit larger model commonly beats a 16-bit smaller one at the same memory budget.
Where to actually run it
Four options, with the trade-offs that matter:
- Your own machine. Free, private, and limited by your hardware. Best for development and for anything where data cannot leave.
- The free hosted inference. Excellent for prototyping, rate-limited, and subject to cold starts — a model that has not been used recently takes time to load before your first response. Not something to build a product on.
- Dedicated endpoints. You pay for a machine that stays warm. Predictable latency, and you are billed for the hours it exists rather than the requests it serves — so an idle endpoint is the classic surprise bill. Set it to scale to zero if the traffic is bursty and you can tolerate the cold start.
- A commercial API from a model vendor. Often the right answer, and worth saying on a page about open models: if you do not need the weights, do not need privacy guarantees, and are not fine-tuning, a hosted frontier model is usually cheaper per unit of quality than operating your own.
Tokens, and keeping them out of your code
The API examples you will find online, including older versions of this page, paste the token into the source. Do not.
import os, requests
API_URL = "https://api-inference.huggingface.co/models/mistralai/Mistral-7B-Instruct-v0.1"
headers = {"Authorization": f"Bearer {os.environ['HF_TOKEN']}"}
def query(payload):
r = requests.post(API_URL, headers=headers, json=payload, timeout=60)
r.raise_for_status()
return r.json()
print(query({"inputs": "Explain machine learning in two sentences."}))
Two habits worth adopting at the same time. Create fine-grained tokens scoped to what the job needs — read-only for pulling a model, write only where you genuinely publish — rather than one token that can do everything to everything. And if a token ever reaches a commit, treat it as compromised and rotate it; removing the line does not help, because the history keeps it.
The security setting nobody reads
Two things on the Hub execute or load code from a stranger's repository, and both deserve a deliberate decision.
trust_remote_code=True tells the library to run Python that ships with the model repo. Plenty of legitimate models need it, because their architecture is not in the library yet. It is also arbitrary code execution on your machine, from a repository anyone can create. Use it only for models from an organisation you have a reason to trust, and pin the revision so the code cannot change under you.
Weight formats. Older checkpoints ship as Python pickles, which can execute code when loaded. safetensors exists precisely to avoid that and is now the norm — prefer it, and be suspicious of a repository that only offers the old format.
The same caution applies to datasets with loading scripts. None of this means the Hub is dangerous; it means it is a public registry, and public registries are treated the way you treat any dependency you did not write.
Pin the revision, or your output will change
Model repositories are mutable. Weights get updated, tokenisers get fixed, files get renamed, and occasionally a model is removed or made gated. Code that names only the model will silently pick up whatever is there today.
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "mistralai/Mistral-7B-Instruct-v0.1"
REV = "commit-sha-from-the-model-page" # pin it
tok = AutoTokenizer.from_pretrained(REPO, revision=REV)
model = AutoModelForCausalLM.from_pretrained(REPO, revision=REV)
This is the single most useful line in this guide for anything running in production. Without it, a Tuesday morning where the model behaves differently has no explanation and no way back.
Where the disk went
Models are large and the cache is silent about it. A few weeks of experimenting produces tens of gigabytes in a directory you have never looked at.
# where the cache lives
export HF_HOME=/path/with/space
# see what is there, and clear what is not needed
pip install "huggingface_hub[cli]"
huggingface-cli scan-cache
huggingface-cli delete-cache
Worth setting HF_HOME deliberately on any shared or space-constrained machine before the first download rather than after the disk fills.
Key capabilities
- Text — translation, summarisation, classification, extraction, generation.
- Vision — classification, detection, segmentation, image-to-text.
- Audio — Whisper for transcription, and speech synthesis models.
- Multimodal — models that take images and text together.
- Fine-tuning — adapting a model to your own data.
The smaller task-specific models are the underrated part of this list. A compact classifier or embedding model, running locally in milliseconds for nothing, beats calling a large general model for jobs like routing, tagging or deduplication — and those jobs are most of what production systems actually do.
Evaluating two candidates without a research budget
Leaderboards answer a question you did not ask. They rank general capability on public benchmarks, and public benchmarks leak into training data over time, which is part of why a model that tops a table can disappoint on your material. The evaluation that decides anything is small, private and specific to your task.
It takes an afternoon and it is the most valuable afternoon in the project. Collect twenty to fifty real inputs from your own data, covering the ordinary cases and the awkward ones you know about. Write down what a correct output looks like for each — not a score, a description, since for most tasks there is no single right answer and you will be judging rather than matching strings.
Then run every candidate over the same set, save the outputs, and read them side by side without knowing which model produced which. That last part matters more than it sounds: knowing which is the popular model changes how generously you read its answers.
What you are looking for is rarely average quality, because the candidates are usually close. It is the shape of the failures. One model may be slightly worse overall but wrong in ways you can detect and handle; another may be marginally better and wrong in ways that look exactly like being right. For anything going near a customer, the first is the better system even when it loses on the average.
Keep the set. When you change model, change version, or change a prompt in six months, running it again takes minutes and tells you whether you improved anything — which is otherwise unknowable.
The costs that arrive later
Open weights are free to download, and people reasonably conclude the approach is cheap. Some of it is; these are the parts that are not, and knowing them early changes architecture decisions.
- Serving is the real expense. A GPU that stays warm is billed by the hour whether or not anyone is using it. At low or bursty traffic, a per-token commercial API is frequently cheaper than a machine sitting idle overnight.
- Somebody has to operate it. Updates, driver versions, a model that stops loading after a library upgrade, monitoring. This is ordinary infrastructure work and it does not appear in any comparison of model quality.
- Egress and storage. Multi-gigabyte weights pulled repeatedly by a build pipeline cost bandwidth and time; cache them deliberately rather than downloading on every deploy.
- Experiments. Fine-tuning runs, and the ones that fail, are billed the same as the ones that work.
The honest summary: self-hosting wins clearly on privacy, on control, and at sustained high volume. Below that, it frequently loses on total cost while feeling cheaper, because the licence fee was the only number anybody compared.
Fine-tuning: usually not first
Fine-tuning is the most requested and least often correct answer. Before it, in order:
- A better prompt, with examples. Free, immediate, and enough surprisingly often.
- Retrieval, if the problem is that the model does not know your facts. Fine-tuning teaches behaviour, not knowledge — training on your documents is a poor way to make a model recall them.
- A different or larger model. Often cheaper than a training run plus the serving of a custom model.
- Then fine-tune, when you need a consistent format, tone or a narrow classification, and you have the data.
If you do: parameter-efficient methods such as LoRA train a small adapter rather than the whole model, which fits on modest hardware and produces a file of megabytes instead of gigabytes. Data quality dominates quantity — a few hundred carefully checked examples typically beat tens of thousands of scraped ones. And hold out a test set before you start, because a fine-tune always looks excellent on the data it was trained on.
Practical uses
- Transcription — Whisper, locally, for material that should not leave your machine.
- Classification at volume — sentiment, routing, tagging, with a small model that costs nothing per call.
- Embeddings — for search and deduplication; local embedding models are fast, cheap and good.
- Document extraction — pulling fields from invoices and forms.
- Image captioning — alt text and catalogue descriptions at scale.
Spaces, and what free means
Spaces host a demo app for nothing, which is genuinely useful for showing work and for internal tools. The free tier is CPU-only, it sleeps when idle and takes time to wake, and it is public unless you make it private.
Two practical notes: put any key in the Space's secrets rather than in the repository, and remember that a public Space with an API key in its code is a key you have given away. If a demo becomes something people rely on, move it off the free tier rather than discovering its limits during a client call.
Quick tips
- Filter by task in the sidebar rather than searching names — it surfaces the small specialised models you did not know existed.
- Downloads and likes measure popularity, not suitability. Read the model card.
- Look for a quantised version of any large model before assuming you cannot run it.
- Gated models need you to accept terms on the page first; the download fails confusingly otherwise.
- Pin revisions everywhere, and record the licence alongside the model name.