Skip to main content
Level: Advanced Updated: September 2026

Embeddings — semantic vectors

The technology behind semantic search and RAG. How to turn text into numbers that represent meaning — and how to use it in practice.

What an embedding is

An embedding is a representation of text (a word, sentence or document) as a vector of numbers — a list of hundreds or thousands of numbers that encode its meaning. The core idea: texts with similar meaning get vectors that are close in the space, and different texts get distant vectors.

For example, "dog" and "puppy" will be close; "dog" and "car" will be far apart. The trick: closeness is measured by meaning, not by identical words. So embeddings-based search finds relevant results even when you didn't use the exact same words — this is what's called semantic search.

In short

An embedding = a "numeric fingerprint" of meaning. Close in the space = similar in meaning. It's the basis for RAG, semantic search, recommendations and classification.

How similarity is measured — cosine similarity

To know how close two vectors are, you usually use cosine similarity — a measure that examines the angle between the vectors (not the distance). The result ranges from -1 to 1:

In practice, a semantic search engine computes the embedding of the query, and returns the passages with the highest cosine similarity. This computation is done quickly by a Vector Database even over millions of vectors.

Models & dimensions

You don't produce embeddings yourself — you use a dedicated embedding model. The choice affects quality, speed and cost:

The number of dimensions is the length of the vector. More dimensions = more "resolution" but more storage and compute. Important: you must never mix embeddings from different models in the same index — they aren't in the same space.

Main uses

  1. RAG — the basis for retrieving information before the model answers. The most common use.
  2. Semantic search — searching a site/product by meaning, not just keywords.
  3. Classification — grouping texts into categories by similarity.
  4. Clustering — discovering topics/groups within a collection of texts.
  5. Recommendations — "similar articles," "related products."
  6. Deduplication — detecting duplicate or very similar content.

Chunking — splitting documents correctly

You don't embed a whole document as one vector — you split it into chunks and embed each one. This is critical for RAG quality: a chunk that's too big dilutes the meaning; too small loses context.

Code example — basic semantic search

from openai import OpenAI
import numpy as np
client = OpenAI()

def embed(texts):
    r = client.embeddings.create(model="text-embedding-3-small", input=texts)
    return [d.embedding for d in r.data]

def cosine(a, b):
    a, b = np.array(a), np.array(b)
    return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))

docs = ["how to reset a password", "support opening hours", "refund policy"]
doc_vecs = embed(docs)

query = "I forgot my password"
q_vec = embed([query])[0]

scores = [(cosine(q_vec, dv), d) for dv, d in zip(doc_vecs, docs)]
scores.sort(reverse=True)
print(scores[0])   # ('...0.86...', 'how to reset a password')

Note: the query didn't contain the word "reset," but the search found the right passage — because it understood the meaning. In production, this computation is done efficiently at scale by a Vector DB.

What embeddings do not capture

Knowing the blind spots is more useful than knowing the mechanism, because every retrieval system you build will fail in exactly these ways.

An embedding compresses a passage into a fixed-length summary of what it is broadly about. Compression means loss, and the losses are systematic:

Two of these have direct fixes rather than workarounds. Identifiers and exact terms are what keyword search is for, which is why hybrid retrieval is not an optimisation but a correction. Recency and validity are what metadata filtering is for. Neither is solved by a better embedding model.

Questions and answers do not look alike

The most commonly missed detail in this whole topic, and it silently degrades retrieval in a way nobody notices because everything still returns results.

A question — "how do I cancel?" — and the passage that answers it — "To end your subscription, open Settings and…" — share little vocabulary and are not obviously similar texts. Embedding both the same way and comparing them assumes a symmetry that is not there.

Many embedding models are trained for this asymmetry and expect you to say which side you are embedding, usually through a short prefix or a separate mode — one for queries, another for documents. Use the wrong one, or omit it, and the model is doing a different job from the one it was trained for. The results are plausible, slightly worse, and there is no error message.

So: read the model card before you index anything, and check whether it expects an instruction or prefix, and what it should be. Then apply it consistently — a corpus indexed one way and queried another is a subtle, expensive mistake to discover months later.

Where a model is symmetric, this does not apply. The point is to know which you are using rather than to assume.

Dimensions, storage and the arithmetic

Dimension count is treated as a quality dial and it is also a cost decision you can compute before committing.

Storage is roughly vectors × dimensions × bytes per number. A million chunks at 1,536 dimensions in 32-bit floats is around six gigabytes before any index overhead; at 3,072 dimensions it doubles. Memory-resident indexes are what make search fast, so this is frequently your infrastructure constraint rather than a rounding error.

Two levers most people do not know they have:

The practical order: measure retrieval quality at full dimensions on your own data, then reduce until quality starts to move. Most corpora tolerate far more reduction than the default suggests.

Chunking decides more than the model does

Teams compare embedding models and leave chunking at whatever the library defaulted to. The second choice usually matters more.

The stated 300–800 token range with overlap is a reasonable starting point. What improves on it:

Filtering beats better embeddings

If your corpus spans products, customers, languages, regions or years, the largest available quality improvement is usually not a better model — it is not searching the irrelevant parts at all.

Restricting a search to one product's documentation before any vector comparison eliminates a whole category of wrong answer, deterministically, rather than ranking it lower and hoping. It is also faster and cheaper, since fewer vectors are compared.

This requires the metadata to exist at index time, which is why the previous section insists on it. Retrofitting a tenant id onto an index built without one means re-embedding the corpus.

A useful design rule: anything you might ever want to filter by should be metadata, not something you hope the embedding captured. Dates, owners, document types, permissions, language, version.

Changing the model means re-embedding everything

Vectors from different models are not comparable — a point the page makes and whose operational weight is worth spelling out.

Switching embedding models is not a config change. It means re-embedding your entire corpus, rebuilding the index, and verifying that retrieval quality actually improved rather than assuming it. For a large corpus that is a real cost in money and time, and it is why the choice deserves a proper evaluation up front rather than "whatever the tutorial used".

Two things that make a future migration survivable:

And when you do migrate, run both indexes in parallel against your evaluation set before switching. A newer model with a better benchmark score can be worse on your particular domain.

Choosing a model without trusting a leaderboard

Embedding benchmarks rank models on public datasets. Those datasets are not your documents, and the gap between a benchmark and your corpus is often larger than the gap between the top ten models.

The things that actually decide it for a given project:

The test that settles it costs an afternoon: take your evaluation set, index the same corpus with two or three candidates, and compare recall. Whichever wins on your documents is the right answer regardless of where it sits on a public table.

You may not need a vector database

The reflex is to reach for specialised infrastructure. For a great many projects it is unnecessary, and the simpler option is faster to build and easier to reason about.

At modest corpus sizes — the low tens of thousands of chunks — comparing a query against every vector with a straightforward numerical library is fast enough for interactive use, needs no extra service, and is exactly correct rather than approximate. Many internal tools and documentation searches never outgrow this.

A dedicated vector store earns its place when you have a genuinely large corpus, need filtered search at scale, want persistence and concurrent writes, or need the operational features — replication, backups, access control — that a database provides and a file does not. Several ordinary databases now offer vector search as an extension, which is often the right middle ground: no new system, and your metadata is already next to your vectors.

Start with the simplest thing that works on your actual corpus size. It is easier to add infrastructure when you need it than to remove it once everything depends on it.

Measuring whether retrieval works

Everything above is a choice, and choices need a number attached or you are guessing.

Build a small evaluation set: thirty to fifty real questions, each labelled with which chunks genuinely answer them. Tedious, and it converts every later decision from an argument into a measurement.

Then measure recall at k — how often a correct chunk appears in the top k results. That single number tells you whether your chunking, your model and your filtering are working, and it is the ceiling on everything downstream: a reranker cannot promote what retrieval never returned, and a language model cannot answer from a passage it was never given.

Re-run it whenever you change chunk size, embedding model, prefixes or filtering. Retrieval systems are full of changes that help one query and quietly break five, and the set is the only thing that notices.

Common mistakes

Next step

Understand embeddings? Now connect them to storage and retrieval with a Vector DB and RAG.