Skip to main content
AI Engineering Context Engineering
Level: Advanced Updated: September 2026

Context Engineering

The hot discipline of 2026: not just what you ask the model, but what you put in its context window. That's what separates a system that works from one that gets confused.

What Context Engineering is

If Prompt Engineering is about phrasing the instruction well, Context Engineering is the broader step: managing everything that goes into the model's context window on each call — the instructions, the retrieved knowledge, the history, the tool results and the examples. In 2026, as systems became multi-step agents, this became the single most critical skill.

The core idea: the model is only as good as the context you gave it. Throw too much irrelevant information at it — it gets confused and expensive. Give it too little — it guesses. Context Engineering is finding the right mix: the right information, in the right amount, in the right order.

The saying that stuck in 2026

"Most LLM failures in production aren't model failures — they're context failures." Too much, too little, or irrelevant information in the window.

The context window and its problems

The context window is the amount of text (in tokens) the model can "see" in a single call. Even when it's huge, there are three real problems:

The conclusion: more context isn't necessarily better. The goal is relevant, concise context, not maximal.

Parts of the context

Good context is made of several parts, each needing management:

  1. System prompt / instructions: role, rules, format. Relatively stable.
  2. Retrieved knowledge (RAG): the most relevant passages from the knowledge base — not everything, only the relevant.
  3. Conversation history: previous messages. The part that grows and needs management.
  4. Tool results: outputs from tool calls (search, API). Can be enormous.
  5. Examples (few-shot): 1–3 examples that steer the model.
  6. The current query: what the user just asked.

Key techniques

1. Selective retrieval

Instead of pushing a whole document, retrieve only the relevant passages with RAG. This is the basic tool for shrinking context without losing information.

2. Compression & summarization

When the history or a tool result is too big — summarize it before putting it in the context. For example, in a long conversation: summarize the first 10 messages into a paragraph, and keep only the latest ones in full.

3. Ordering & structuring

Put the critical information at the start or end (not the middle). Use clear tags/headings (<knowledge>...</knowledge>) so the model distinguishes between parts. Clear structure improves accuracy.

4. Memory management

Not everything needs to be in every call. Keep "long-term memory" (user preferences, facts) separately, and retrieve only what's relevant to each interaction — exactly like RAG, but over memory.

5. Tool result pruning

A tool that returns 5,000 lines of JSON — don't insert all of it. Filter/summarize before returning it to the model. This is one of the big problems in agents.

The big challenge: context in agents

In a multi-step agent, the context accumulates — each step adds a thought, a tool call and a result. After 15 steps, the context is bloated, expensive, and suffers from context rot. This is why agents "fall apart" on long tasks.

Strategies that work:

A context budget you can actually enforce

Everything above is advice about proportion — relevant, concise, well-ordered — and advice about proportion is useless until it becomes a number. The teams that get this right treat the context window as a budget with named line items, decided in advance, rather than as a bucket that fills up until something breaks.

Pick a working ceiling well below the model's actual limit. Not because the model cannot handle more, but because the limit is where you get an error, and you want a policy that takes effect long before that. Then divide it: a fixed allowance for the system prompt, an allowance for retrieved passages, one for history, one for tool results, and headroom for the answer itself. The allowances need not be equal and the right split depends entirely on the job — an extraction task wants most of its budget in retrieval, a long support conversation wants it in history.

What makes it a budget rather than a wish is deciding, ahead of time, what gets dropped when a section overflows — and dropping it in the code rather than hoping it does not happen. Retrieval overflow means fewer passages, lowest-scoring first. History overflow means the oldest turns get compacted, not truncated mid-sentence. Tool-result overflow means the result is summarised or written to a file and replaced by a reference. The rule for each section should be written down somewhere a person can read it, because when output quality drops a month from now, the first question will be what the system chose to leave out, and "it depends on what fit" is not an answer you can debug.

One line item people forget: the answer. If you fill the window to the brim with input, the model has nowhere to write. A long structured response needs thousands of tokens of room, and a system that budgets only its input will hit truncated outputs intermittently, under exactly the conditions — a rich, complex question — where the answer mattered most.

Caching changes the arithmetic

Prompt caching is the single biggest practical influence on how context should be ordered, and it quietly contradicts some of the advice above. Providers cache a prefix: if the beginning of your request is byte-for-byte identical to a recent one, that portion is served from cache far more cheaply and faster. The moment one character differs, the cache stops matching from that point onward — everything after it is recomputed.

This turns ordering into an economic decision, not just an attention one. Anything stable belongs at the front: the system prompt, the tool definitions, the few-shot examples, the long reference document that never changes between calls. Anything volatile belongs at the back: the current query, the newest turns, the freshly retrieved passages. A system that injects a timestamp or a session id into the top of its system prompt has, without anyone noticing, disabled caching for every request it ever makes.

The tension with lost-in-the-middle is real but smaller than it looks. Lost-in-the-middle is about where the critical information sits, and the critical information is usually the query and the retrieved passages — which are volatile, so they go at the end anyway. The two rules mostly agree. Where they genuinely conflict — a large, stable reference document that is also the thing the model must attend to closely — the fix is not to move it but to point at it: keep the document in the cached prefix and put a short, specific instruction near the end telling the model what to look for in it.

Caching also rewards stable formatting in ways that feel pedantic and are not. If your history is rendered by joining messages with a separator, and that separator changes, or messages get renumbered as the conversation grows, or a summary replaces earlier turns in place — each of those invalidates the prefix. Appending is cache-friendly; rewriting the beginning is not. Design the assembly so that a new turn adds to the end and changes nothing before it, and compact on a schedule you control rather than on every call.

Compaction is lossy in a specific direction

Summarising history to keep it small is the standard move, and it works. It is also the most common source of a particular kind of agent failure: the system repeats work it already did, or contradicts a decision it already made, with complete confidence and no sign that anything is missing.

That happens because summaries are systematically biased toward what succeeded. Asked to summarise a long exchange, a model writes down the conclusions and drops the process — and the process is where the negatives live. "We tried the batch endpoint and it rejected records over 50 fields" becomes nothing at all, so twenty steps later the agent tries the batch endpoint again. Exact identifiers go the same way: an order number, a file path, a chosen variable name gets paraphrased into "the relevant record," which is unusable for anything that has to reference it precisely.

The practical answer is to stop treating compaction as a single summarisation and split it in two. Keep a short, append-only ledger, in plain text, outside the summarised history: decisions taken, constraints discovered, things that failed and why, and every identifier that will be needed again — written verbatim, never paraphrased. Summarise the narrative freely; carry the ledger forward untouched. It stays small because it only holds facts that a sentence cannot lose, and it is the difference between an agent that remembers what it learned and one that only remembers what it concluded.

Also, compact on a boundary rather than a token count where you can. Summarising in the middle of a multi-step operation — between calling a tool and reading its result — produces a history that describes an action whose outcome has vanished. Waiting for a natural seam costs a few hundred tokens and avoids a class of confusion that is very hard to diagnose after the fact.

Context is also a trust boundary

Everything in the window arrives as text, and the model has no built-in sense of where each part came from. Your system prompt, a passage retrieved from a document someone uploaded, the body of a web page a tool fetched, and the user's question all look the same once they are concatenated. If a retrieved document contains a sentence addressed to the model — instructions phrased as if they came from you — nothing in the format distinguishes it from an instruction that actually did.

This is why context assembly deserves treating as a security concern and not only an efficiency one. Three habits cover most of it. Label provenance explicitly: wrap retrieved content and tool output in tags that name their source, and state in the system prompt that content inside those tags is data to be analysed, never instructions to be followed. Keep the trusted instructions in the stable prefix, where they are not competing with untrusted text for the model's attention at the end of the window. And give the model's tools the narrowest permissions the task allows, because the meaningful damage comes not from the model being convinced of something false but from what it is then able to do about it.

None of this is airtight, and it should not be sold as such — a sufficiently well-crafted passage can still steer a model that is reading it attentively. What labelling and least privilege buy you is that the failure is contained and visible: a model that gets confused by a poisoned document produces a wrong answer rather than a deleted table, and the logged context shows exactly which source the instruction came from.

Measure the window instead of guessing at it

The reason context problems persist is that the context is invisible. Nobody reads what actually went into the call; they read the prompt template and assume. Logging the fully assembled window — not the template, the final string — is the cheapest diagnostic available, and it reliably surprises people the first time they look. The duplicated instruction block, the retrieved passage that is the same paragraph three times, the tool result that is 40KB of whitespace, the history that still contains a failed attempt from an hour ago: none of these are visible from the code that builds the context, and all of them are obvious in one glance at the output.

Beyond reading it, two measurements pay for themselves. Track token count by section over time — if history is 70% of the window, no amount of retrieval tuning will help, and you have been optimising the wrong thing. And run occasional ablations: take a set of cases where the system got the answer right, remove one part of the context, and see whether it still does. Parts that make no difference when removed are pure cost, and there are usually more of them than anyone expects. The few-shot examples added early on, still there long after the instructions were rewritten around them, are the classic case.

Do both against a fixed set of real cases rather than ad hoc, and the vague question "is our context any good?" turns into something you can answer on a Tuesday afternoon with a number.

Instructions decay as the conversation grows

A system prompt that governs the first exchange cleanly can lose its grip by the fortieth. Nothing removed it — it is still there, at the top, unchanged. But it is now a small fraction of a very large window, competing with thousands of tokens of more recent, more specific material, and the practical effect is that a formatting rule or a tone constraint gets quietly abandoned somewhere in the middle of a long session.

The fix is unglamorous: restate the two or three constraints that genuinely matter close to the end of the context, just before the current query. Not the whole system prompt — that wastes tokens and breaks nothing usefully — only the rules whose violation you would actually care about. A single line is usually enough.

It is worth finding out where your own decay point is rather than assuming one. Run the same instruction-following check at turn 5, turn 25 and turn 50 of a realistic conversation; the turn at which compliance starts slipping is a property of your prompt and your task, and knowing it tells you how often the reminder needs to appear.

Common mistakes

Next step

Good context management is essential for agents and RAG. Go deeper on those, or move to running in production.