LLMOps — LLM in production
Building a prototype is easy. Running a reliable, cheap and safe LLM product for thousands of users is real engineering. Here's how to do it right.
What LLMOps is
LLMOps (Large Language Model Operations) is the set of practices for running and maintaining LLM products in production — the equivalent of DevOps/MLOps, but adapted to the unique challenges of language models: non-deterministic output, variable cost, hallucinations, and dependence on a third-party API.
The difference from a prototype: a prototype needs to work once, on your machine. A production product needs to work a million times — reliably, quickly, cheaply and safely — even when the provider's API goes down, even when a user tries to break it, and even when volume spikes 10x.
LLMOps = everything needed to turn a clever prototype into a product you can rely on: versioning, monitoring, reliability, cost and security.
The lifecycle of an LLM product
LLMOps is a loop, not a straight line:
- Development: prompt, RAG, agent — with evals from day one.
- Testing: the evals run in CI. A change doesn't pass if quality drops.
- Deployment: a controlled rollout (staging → canary → production).
- Monitoring: Observability on real traffic — quality, cost, latency, errors.
- Improvement: production failures become new eval cases, and back to step 1.
Versioning — prompts and models
In the LLM world, the prompt is code — it affects behavior just like logic. So:
- Prompts in git, not pasted into code or a DB without history. Every change goes through review and evals.
- Pin the model version. Don't use an alias that shifts under you — specify the dated or numbered id from your provider's list rather than a floating alias like
-latest, because a model upgrade can change behaviour without you knowing. - Centralized config. Model, temperature, prompt and parameters in one place you can change without a deploy.
- A/B and rollback. Infrastructure to run two prompt/model versions in parallel and roll back immediately if something breaks.
Reliability and model routing
You depend on a third-party API. Plan for failures:
- Retries with backoff: transient errors and rate limits happen. Retry with increasing delay.
- Fallback model: if the primary provider goes down or is slow — automatically switch to an alternative model/provider.
- Timeouts: don't let a request hang the system. Set a timeout and handle it.
- Model routing: route by complexity — a cheap model (Haiku/mini/Flash) for most calls, premium only for complexity. Saves money and load.
- Circuit breaker: if a provider keeps failing, "cut it off" temporarily instead of continuing to retry.
Cost, latency and caching
LLM cost can quietly explode. Control it:
- Prompt caching: providers let you cache the fixed part of the prompt (system, fixed context) — significant savings on repeated calls.
- Semantic caching: if a similar question was already answered — return the cached answer instead of a new call.
- Streaming: stream the response to the user to improve perceived latency.
- Budgets and limits: set cost ceilings and alerts. See the cost-reduction guide.
- Measure cost per request/user — to know what's really expensive and optimize correctly.
What makes this different from DevOps
Most LLMOps advice is DevOps practice with the word "prompt" substituted in. The useful part is the places where the analogy breaks, because those are where teams with strong engineering habits still get caught out.
You cannot assert correctness. Ordinary software has a right answer you can compare against. An LLM system has a distribution of acceptable answers, so tests become statistical: not "does this equal that" but "does quality on a hundred cases stay above a threshold". Every downstream practice changes as a result.
The dependency changes without a version bump. A library you pinned behaves identically forever. A hosted model behind the same id can be updated, and even where it is not, infrastructure changes shift behaviour at the margins. Your system can degrade with no deploy on your side, which is a failure mode ordinary services do not have.
Cost is per request and unbounded. A traditional service has capacity limits; an LLM system has a bill that scales with how enthusiastically people use it and with how long the outputs happen to be.
The input space is adversarial and infinite. Users type anything, and some of them are trying to break it. There is no schema to validate against at the front door.
Evals in CI, practically
"Run evals in CI" is the right instruction and it hides three problems worth solving before you set it up.
- They cost money and time. A few hundred cases against a frontier model on every commit is real spend and several minutes. The workable pattern: a small fast subset on every pull request, the full set nightly and before release.
- They are not deterministic. Even at temperature zero, a provider-side change can move a score. So gate on a threshold with a margin, not on an exact figure, or the build goes red for reasons nobody caused.
- Judge models drift too. If an LLM grades the outputs, the grader is itself a dependency that can change. Pin it, and keep a small human-labelled set to check the judge against periodically.
What to gate on: a quality score that must not drop, a cost-per-case ceiling, and a latency ceiling. Three numbers, each with a threshold, each blocking. See evals for building the set itself.
Releasing a change you cannot diff
You can read a code diff and reason about what it does. A prompt diff tells you what changed textually and nothing about what it will do — which is why LLM releases need a different shape.
- Offline evals on the fixed set. Necessary, and it only covers cases you thought of.
- Shadow mode. Run the new version alongside the old on real traffic without showing users the output. Compare quality, cost and latency on inputs nobody invented. This is the highest-value step and the one most teams skip.
- Canary. A small share of real traffic, watched closely, with an automatic rollback if quality or error rates move.
- Full rollout, with the previous version still one config change away.
Two details that make rollback actually work. Prompt, model and parameters are one unit — a prompt tuned for one model is not tuned for another, so version them together and roll them back together. And keep the rollback in configuration rather than in a deploy, so recovering from a bad prompt takes a minute rather than a release cycle.
Model deprecation is a scheduled outage
The risk nobody plans for until it happens: providers retire models. You get notice, the notice arrives in an email somebody may not read, and on the stated date your pinned version stops answering.
This is the strongest argument for the discipline above, because a migration is exactly the same problem as a release — new behaviour, same prompts — and a team with evals and shadow mode handles it in an afternoon while a team without them discovers the differences in production.
- Know what you are pinned to, everywhere. A register of every model id in the system, with where it is used, is a ten-minute document that saves a bad week.
- Watch the deprecation notices deliberately — a shared inbox rather than one person's.
- Re-run your evals on the successor as soon as it exists, not when the old one dies. The gap is usually months and it is free to use.
- Expect prompt adjustments. Behaviour differs between model generations in ways that break carefully tuned instructions, and the fix is rarely large but it is never zero.
- Keep a second provider integrated, even if unused, so that "switch" is a config change rather than a project.
Logging, and the problem with logging prompts
Debugging an LLM system requires seeing what went in and what came out. That conflicts directly with not storing more personal data than you should.
A workable balance:
- Always log the metadata — model, prompt version, token counts, latency, cost, retrieval ids, which guardrails fired. This is most of what you need for operations and it contains no user content.
- Sample the content. Full prompts and outputs on a small percentage of traffic, plus all errors and all flagged responses. You rarely need every request to diagnose a pattern.
- Redact before storing where you can, and set a short retention on anything containing user text.
- Keep a trace id through the whole chain — retrieval, generation, validation, action — so a support complaint can be reconstructed end to end.
- Treat logs as sensitive. A debug log of prompts is a database of whatever your users typed, with the access controls to match.
The incidents that are specific to this stack
An on-call rotation for an LLM product meets failure modes that do not appear in ordinary services. Worth writing runbooks for these before they happen at two in the morning.
- Quality degraded, nothing deployed. The hardest one, because every instinct says look at the last change and there was not one. Check provider status, check whether your retrieval corpus changed, and compare a sample of today's outputs against last week's on the same inputs. This is precisely what shadow-mode infrastructure lets you do reactively.
- Quota exhausted mid-day. A usage spike or a runaway loop burns the month's budget by lunchtime. Needs per-feature rate limits and a hard ceiling that degrades gracefully rather than erroring — a queue, a cheaper model, or an honest "try again shortly".
- A loop. An agent calls a tool that triggers the agent. Bounded step counts and a per-conversation spend cap turn a catastrophe into a logged incident.
- Prompt injection reaching production. Content in a retrieved document or an email steers the model. The response is containment — revoke the affected credentials, check what actions were taken, then fix the boundary. See prompt injection.
- The provider is down. Which is why the fallback exists, and why it should be exercised occasionally rather than assumed to work.
One habit that helps across all of them: a dashboard showing quality, cost and latency together over time, not three separate graphs in three tools. Most of these incidents show up first as one of the three moving while the others stay flat, and that pattern is the diagnosis.
The cost number that matters
Teams optimise cost per token, which is the wrong denominator and leads to the wrong decisions.
The number that matters is cost per completed outcome — per resolved support ticket, per document processed, per qualified lead. Measured that way, a more expensive model that succeeds first time frequently beats a cheap one that needs two retries and a human, and the analysis reverses.
It also makes the business case legible. "Our AI costs go up when usage goes up" alarms a finance team; "each resolved ticket costs a fraction of what the human-handled equivalent costs, and here is the trend" is a conversation about scaling.
Track it per feature, not globally. Most systems have one expensive path that dominates the bill, and finding it is usually worth more than any amount of prompt trimming.
Turning production failures into tests
The improvement loop is the step that decides whether the system gets better over time or just gets older. It works when it is mechanical.
- Capture the failure. A thumbs-down, an escalation to a human, a validation rejection, a support complaint. Each needs to land somewhere with its trace id.
- Triage weekly. Someone looks at the week's failures and sorts them: retrieval problem, prompt problem, model limitation, genuinely out of scope.
- Add the interesting ones to the eval set, with the correct answer written by a person. This is the part that compounds.
- Fix, and confirm the new case passes without the old ones regressing.
- Watch it in production, because the eval set proves the fix works on that case and not that it works.
An eval set that grows from real failures becomes, after a few months, the most valuable artefact your team owns — and the one that makes changing models a measured decision rather than a leap.
Who owns the prompt
An organisational question that causes more production incidents than any technical one.
Prompts sit awkwardly: they read like content, so product and support people reasonably want to edit them, and they behave like code, so a careless edit can break the system silently. The two failure modes are a prompt only engineers can change — so it never improves, because engineers do not read the support queue — and a prompt anyone can change in a dashboard, which is an unversioned production deploy by someone who does not know it is one.
What works: prompts in version control, edited by whoever knows the domain, shipped through the same review and evals as code. The domain expert writes the change, the evals decide whether it goes out, and the history shows who changed what and why. Give non-engineers a real path to contribute, and put the gate on the evals rather than on the person.
Write the post-mortem into the eval set
Ordinary incident reviews end with an action item and a fix. An LLM system gives you somewhere better to put the lesson.
Every incident involving a wrong answer should end with that exact input added to the eval set, with the output you wanted, and a test that would have caught it. The fix goes in the code; the knowledge goes in the set, where it stays and gets checked on every future change — including the model migration two years from now, when nobody involved in the original incident still works here.
That is the difference between a team that accumulates operational knowledge and one that keeps rediscovering the same edge cases. It costs ten minutes at the end of a review and it is the single highest-return habit on this page.
Where to start, in order
The checklist below is the destination. If you are at a working prototype, this is the order that gets you there without stalling.
- Twenty eval cases in a file, run by hand. Not a framework — a file. This alone changes how you make decisions.
- Pin the model and put the prompt in git. An afternoon, and it makes everything after it possible.
- Log metadata and sample content, with a trace id.
- Timeouts, retries and a fallback. The cheapest reliability you will ever buy.
- Output validation before anything downstream acts on it.
- Evals in CI, once the set is worth gating on.
- Shadow mode and canary, when a bad release would actually hurt.
- Cost per outcome, tracked per feature.
Teams that do the first two get most of the benefit. Teams that start at step seven build impressive infrastructure around a system nobody has measured.
Pre-production checklist
- ✅ Evals run in CI and block regressions
- ✅ Prompts in git with review; model version pinned
- ✅ Guardrails and prompt-injection defense
- ✅ Retries, timeouts and a fallback model
- ✅ Observability: logging, metrics and tracing
- ✅ Cost control: caching, budget and alerts
- ✅ Validation of every output before downstream use
- ✅ A rollback plan and incident response
Treat the LLM system like any critical production service: measurable, reproducible, fault-tolerant and secure. The AI "magic" doesn't exempt you from good engineering — it demands it more.
Next step
Go deeper on the critical production components: monitoring, evals and safety.