Embeddings — semantic vectors
The technology behind semantic search and RAG. How to turn text into numbers that represent meaning — and how to use it in practice.
What an embedding is
An embedding is a representation of text (a word, sentence or document) as a vector of numbers — a list of hundreds or thousands of numbers that encode its meaning. The core idea: texts with similar meaning get vectors that are close in the space, and different texts get distant vectors.
For example, "dog" and "puppy" will be close; "dog" and "car" will be far apart. The trick: closeness is measured by meaning, not by identical words. So embeddings-based search finds relevant results even when you didn't use the exact same words — this is what's called semantic search.
An embedding = a "numeric fingerprint" of meaning. Close in the space = similar in meaning. It's the basis for RAG, semantic search, recommendations and classification.
How similarity is measured — cosine similarity
To know how close two vectors are, you usually use cosine similarity — a measure that examines the angle between the vectors (not the distance). The result ranges from -1 to 1:
- 1.0 — completely identical in meaning
- ~0.8 — very similar
- ~0 — unrelated
In practice, a semantic search engine computes the embedding of the query, and returns the passages with the highest cosine similarity. This computation is done quickly by a Vector Database even over millions of vectors.
Models & dimensions
You don't produce embeddings yourself — you use a dedicated embedding model. The choice affects quality, speed and cost:
- OpenAI —
text-embedding-3-small(1536 dimensions, cheap and fast) andtext-embedding-3-large(3072 dimensions, more accurate). A good default for most. Model ids and offerings change — check the provider's current list rather than copying a name from a guide. - Open source / local — families like BGE, E5 and nomic. Free and private, running locally or via Hugging Face.
- Multilingual — if your content is not English, check the model was trained with real coverage of your language rather than trusting a multilingual label. Quality varies far more across languages than headline benchmarks suggest.
The number of dimensions is the length of the vector. More dimensions = more "resolution" but more storage and compute. Important: you must never mix embeddings from different models in the same index — they aren't in the same space.
Main uses
- RAG — the basis for retrieving information before the model answers. The most common use.
- Semantic search — searching a site/product by meaning, not just keywords.
- Classification — grouping texts into categories by similarity.
- Clustering — discovering topics/groups within a collection of texts.
- Recommendations — "similar articles," "related products."
- Deduplication — detecting duplicate or very similar content.
Chunking — splitting documents correctly
You don't embed a whole document as one vector — you split it into chunks and embed each one. This is critical for RAG quality: a chunk that's too big dilutes the meaning; too small loses context.
- Typical size: 300–800 tokens per chunk, with an overlap of ~10–15% so you don't cut sentences in the middle.
- Smart splitting: prefer natural boundaries (paragraphs, headings) over arbitrary cuts.
- Metadata: store a source, title and date for each chunk — useful for filtering and citation.
Code example — basic semantic search
from openai import OpenAI
import numpy as np
client = OpenAI()
def embed(texts):
r = client.embeddings.create(model="text-embedding-3-small", input=texts)
return [d.embedding for d in r.data]
def cosine(a, b):
a, b = np.array(a), np.array(b)
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
docs = ["how to reset a password", "support opening hours", "refund policy"]
doc_vecs = embed(docs)
query = "I forgot my password"
q_vec = embed([query])[0]
scores = [(cosine(q_vec, dv), d) for dv, d in zip(doc_vecs, docs)]
scores.sort(reverse=True)
print(scores[0]) # ('...0.86...', 'how to reset a password')
Note: the query didn't contain the word "reset," but the search found the right passage — because it understood the meaning. In production, this computation is done efficiently at scale by a Vector DB.
What embeddings do not capture
Knowing the blind spots is more useful than knowing the mechanism, because every retrieval system you build will fail in exactly these ways.
An embedding compresses a passage into a fixed-length summary of what it is broadly about. Compression means loss, and the losses are systematic:
- Negation. "The plan includes support" and "the plan does not include support" land close together, because they concern the same subject. This is the failure that surprises people most and it is structural rather than a model weakness.
- Identifiers. Order numbers, SKUs, version strings, error codes, people's names. These carry meaning by being exact, and exactness is precisely what compression discards. SKU-4471 and SKU-4417 are nearly identical vectors.
- Numbers and quantities. "Under ten employees" and "over a thousand employees" are about the same topic and embed accordingly.
- Recency and validity. A vector has no notion of whether a document is current, superseded or wrong. A policy from 2019 retrieves exactly as well as this year's.
- Specificity within a topic. A passage about refunds scores well for a delivery question, because both are customer-service text.
Two of these have direct fixes rather than workarounds. Identifiers and exact terms are what keyword search is for, which is why hybrid retrieval is not an optimisation but a correction. Recency and validity are what metadata filtering is for. Neither is solved by a better embedding model.
Questions and answers do not look alike
The most commonly missed detail in this whole topic, and it silently degrades retrieval in a way nobody notices because everything still returns results.
A question — "how do I cancel?" — and the passage that answers it — "To end your subscription, open Settings and…" — share little vocabulary and are not obviously similar texts. Embedding both the same way and comparing them assumes a symmetry that is not there.
Many embedding models are trained for this asymmetry and expect you to say which side you are embedding, usually through a short prefix or a separate mode — one for queries, another for documents. Use the wrong one, or omit it, and the model is doing a different job from the one it was trained for. The results are plausible, slightly worse, and there is no error message.
So: read the model card before you index anything, and check whether it expects an instruction or prefix, and what it should be. Then apply it consistently — a corpus indexed one way and queried another is a subtle, expensive mistake to discover months later.
Where a model is symmetric, this does not apply. The point is to know which you are using rather than to assume.
Dimensions, storage and the arithmetic
Dimension count is treated as a quality dial and it is also a cost decision you can compute before committing.
Storage is roughly vectors × dimensions × bytes per number. A million chunks at 1,536 dimensions in 32-bit floats is around six gigabytes before any index overhead; at 3,072 dimensions it doubles. Memory-resident indexes are what make search fast, so this is frequently your infrastructure constraint rather than a rounding error.
Two levers most people do not know they have:
- Truncation. Several current embedding models are trained so that the first N dimensions remain useful on their own — you can cut a 3,072-dimension vector down substantially and lose less quality than you would expect. Check whether yours supports this; it is a large saving for a small accuracy cost.
- Quantisation. Storing each number in fewer bits, or even as binary, shrinks the index dramatically. Combined with a re-ranking pass over the top candidates in full precision, it is how large systems stay affordable.
The practical order: measure retrieval quality at full dimensions on your own data, then reduce until quality starts to move. Most corpora tolerate far more reduction than the default suggests.
Chunking decides more than the model does
Teams compare embedding models and leave chunking at whatever the library defaulted to. The second choice usually matters more.
The stated 300–800 token range with overlap is a reasonable starting point. What improves on it:
- Prefix every chunk with its heading path — document title, section, subsection — before embedding. A chunk reading "This does not apply to annual plans" is meaningless alone and precise when preceded by "Billing › Refunds › Exceptions". This is the cheapest retrieval improvement available and almost nobody does it.
- Split on structure, not on character count. Headings, paragraphs, list boundaries. A chunk that begins mid-sentence embeds badly and reads badly when shown as a citation.
- Retrieve small, return large. Index precise small chunks for matching, then feed the model the surrounding section once a chunk wins. You get precision in retrieval and context in generation, which single-size chunking cannot give you at once.
- Keep tables and code intact. Splitting a table across chunks produces two useless fragments. Detect and keep them whole even when oversized.
- Store the source, the section and the date as metadata on every chunk. This is what makes filtering and citation possible later, and adding it retrospectively means re-indexing everything.
Filtering beats better embeddings
If your corpus spans products, customers, languages, regions or years, the largest available quality improvement is usually not a better model — it is not searching the irrelevant parts at all.
Restricting a search to one product's documentation before any vector comparison eliminates a whole category of wrong answer, deterministically, rather than ranking it lower and hoping. It is also faster and cheaper, since fewer vectors are compared.
This requires the metadata to exist at index time, which is why the previous section insists on it. Retrofitting a tenant id onto an index built without one means re-embedding the corpus.
A useful design rule: anything you might ever want to filter by should be metadata, not something you hope the embedding captured. Dates, owners, document types, permissions, language, version.
Changing the model means re-embedding everything
Vectors from different models are not comparable — a point the page makes and whose operational weight is worth spelling out.
Switching embedding models is not a config change. It means re-embedding your entire corpus, rebuilding the index, and verifying that retrieval quality actually improved rather than assuming it. For a large corpus that is a real cost in money and time, and it is why the choice deserves a proper evaluation up front rather than "whatever the tutorial used".
Two things that make a future migration survivable:
- Keep the source text, chunked, alongside the vectors. If all you stored was embeddings, you cannot re-embed without redoing the extraction and chunking too.
- Record which model and version produced each vector, so a partial migration is detectable rather than silently mixing spaces — which produces subtly wrong results with no error.
And when you do migrate, run both indexes in parallel against your evaluation set before switching. A newer model with a better benchmark score can be worse on your particular domain.
Choosing a model without trusting a leaderboard
Embedding benchmarks rank models on public datasets. Those datasets are not your documents, and the gap between a benchmark and your corpus is often larger than the gap between the top ten models.
The things that actually decide it for a given project:
- Domain match. A model trained mostly on web text handles web text well and may do noticeably worse on medical notes, legal drafting or code. Domain distance costs more than a few benchmark points.
- Your languages, tested on your own content. A multilingual label covers a wide range of actual quality, and cross-lingual retrieval — a question in one language finding an answer in another — either works well or barely at all depending on the model.
- Input length limits. Models truncate silently past their maximum. If your chunks exceed it, you are indexing the first part of each chunk and believing you indexed all of it.
- Hosted or local. An API call per chunk is simple and has a per-token cost and a privacy implication; a local model is free per call and needs hardware. At corpus scale this is a real decision, because indexing a million chunks is a million calls.
- Stability. A hosted model that is updated can change your vector space. Check whether the provider versions its embedding models, because a silent update means your new vectors no longer match your old ones.
The test that settles it costs an afternoon: take your evaluation set, index the same corpus with two or three candidates, and compare recall. Whichever wins on your documents is the right answer regardless of where it sits on a public table.
You may not need a vector database
The reflex is to reach for specialised infrastructure. For a great many projects it is unnecessary, and the simpler option is faster to build and easier to reason about.
At modest corpus sizes — the low tens of thousands of chunks — comparing a query against every vector with a straightforward numerical library is fast enough for interactive use, needs no extra service, and is exactly correct rather than approximate. Many internal tools and documentation searches never outgrow this.
A dedicated vector store earns its place when you have a genuinely large corpus, need filtered search at scale, want persistence and concurrent writes, or need the operational features — replication, backups, access control — that a database provides and a file does not. Several ordinary databases now offer vector search as an extension, which is often the right middle ground: no new system, and your metadata is already next to your vectors.
Start with the simplest thing that works on your actual corpus size. It is easier to add infrastructure when you need it than to remove it once everything depends on it.
Measuring whether retrieval works
Everything above is a choice, and choices need a number attached or you are guessing.
Build a small evaluation set: thirty to fifty real questions, each labelled with which chunks genuinely answer them. Tedious, and it converts every later decision from an argument into a measurement.
Then measure recall at k — how often a correct chunk appears in the top k results. That single number tells you whether your chunking, your model and your filtering are working, and it is the ceiling on everything downstream: a reranker cannot promote what retrieval never returned, and a language model cannot answer from a passage it was never given.
Re-run it whenever you change chunk size, embedding model, prefixes or filtering. Retrieval systems are full of changes that help one query and quietly break five, and the set is the only thing that notices.
Common mistakes
- Mixing models. You must not search with one embedding model against an index built with another.
- Bad chunks. Most RAG problems are actually chunking problems, not model problems.
- Ignoring normalization. If the DB doesn't normalize, compare with cosine, not dot product.
- Not storing metadata. Without a source per chunk you can't cite or filter.
- Semantic search only. Combining with keyword search (hybrid) usually gives better results — see Advanced RAG.
Next step
Understand embeddings? Now connect them to storage and retrieval with a Vector DB and RAG.