Skip to main content
Level: Expert Updated: August 2026

LLMOps — LLM in production

Building a prototype is easy. Running a reliable, cheap and safe LLM product for thousands of users is real engineering. Here's how to do it right.

What LLMOps is

LLMOps (Large Language Model Operations) is the set of practices for running and maintaining LLM products in production — the equivalent of DevOps/MLOps, but adapted to the unique challenges of language models: non-deterministic output, variable cost, hallucinations, and dependence on a third-party API.

The difference from a prototype: a prototype needs to work once, on your machine. A production product needs to work a million times — reliably, quickly, cheaply and safely — even when the provider's API goes down, even when a user tries to break it, and even when volume spikes 10x.

In short

LLMOps = everything needed to turn a clever prototype into a product you can rely on: versioning, monitoring, reliability, cost and security.

The lifecycle of an LLM product

LLMOps is a loop, not a straight line:

  1. Development: prompt, RAG, agent — with evals from day one.
  2. Testing: the evals run in CI. A change doesn't pass if quality drops.
  3. Deployment: a controlled rollout (staging → canary → production).
  4. Monitoring: Observability on real traffic — quality, cost, latency, errors.
  5. Improvement: production failures become new eval cases, and back to step 1.

Versioning — prompts and models

In the LLM world, the prompt is code — it affects behavior just like logic. So:

Reliability and model routing

You depend on a third-party API. Plan for failures:

Cost, latency and caching

LLM cost can quietly explode. Control it:

What makes this different from DevOps

Most LLMOps advice is DevOps practice with the word "prompt" substituted in. The useful part is the places where the analogy breaks, because those are where teams with strong engineering habits still get caught out.

You cannot assert correctness. Ordinary software has a right answer you can compare against. An LLM system has a distribution of acceptable answers, so tests become statistical: not "does this equal that" but "does quality on a hundred cases stay above a threshold". Every downstream practice changes as a result.

The dependency changes without a version bump. A library you pinned behaves identically forever. A hosted model behind the same id can be updated, and even where it is not, infrastructure changes shift behaviour at the margins. Your system can degrade with no deploy on your side, which is a failure mode ordinary services do not have.

Cost is per request and unbounded. A traditional service has capacity limits; an LLM system has a bill that scales with how enthusiastically people use it and with how long the outputs happen to be.

The input space is adversarial and infinite. Users type anything, and some of them are trying to break it. There is no schema to validate against at the front door.

Evals in CI, practically

"Run evals in CI" is the right instruction and it hides three problems worth solving before you set it up.

What to gate on: a quality score that must not drop, a cost-per-case ceiling, and a latency ceiling. Three numbers, each with a threshold, each blocking. See evals for building the set itself.

Releasing a change you cannot diff

You can read a code diff and reason about what it does. A prompt diff tells you what changed textually and nothing about what it will do — which is why LLM releases need a different shape.

  1. Offline evals on the fixed set. Necessary, and it only covers cases you thought of.
  2. Shadow mode. Run the new version alongside the old on real traffic without showing users the output. Compare quality, cost and latency on inputs nobody invented. This is the highest-value step and the one most teams skip.
  3. Canary. A small share of real traffic, watched closely, with an automatic rollback if quality or error rates move.
  4. Full rollout, with the previous version still one config change away.

Two details that make rollback actually work. Prompt, model and parameters are one unit — a prompt tuned for one model is not tuned for another, so version them together and roll them back together. And keep the rollback in configuration rather than in a deploy, so recovering from a bad prompt takes a minute rather than a release cycle.

Model deprecation is a scheduled outage

The risk nobody plans for until it happens: providers retire models. You get notice, the notice arrives in an email somebody may not read, and on the stated date your pinned version stops answering.

This is the strongest argument for the discipline above, because a migration is exactly the same problem as a release — new behaviour, same prompts — and a team with evals and shadow mode handles it in an afternoon while a team without them discovers the differences in production.

Logging, and the problem with logging prompts

Debugging an LLM system requires seeing what went in and what came out. That conflicts directly with not storing more personal data than you should.

A workable balance:

The incidents that are specific to this stack

An on-call rotation for an LLM product meets failure modes that do not appear in ordinary services. Worth writing runbooks for these before they happen at two in the morning.

One habit that helps across all of them: a dashboard showing quality, cost and latency together over time, not three separate graphs in three tools. Most of these incidents show up first as one of the three moving while the others stay flat, and that pattern is the diagnosis.

The cost number that matters

Teams optimise cost per token, which is the wrong denominator and leads to the wrong decisions.

The number that matters is cost per completed outcome — per resolved support ticket, per document processed, per qualified lead. Measured that way, a more expensive model that succeeds first time frequently beats a cheap one that needs two retries and a human, and the analysis reverses.

It also makes the business case legible. "Our AI costs go up when usage goes up" alarms a finance team; "each resolved ticket costs a fraction of what the human-handled equivalent costs, and here is the trend" is a conversation about scaling.

Track it per feature, not globally. Most systems have one expensive path that dominates the bill, and finding it is usually worth more than any amount of prompt trimming.

Turning production failures into tests

The improvement loop is the step that decides whether the system gets better over time or just gets older. It works when it is mechanical.

  1. Capture the failure. A thumbs-down, an escalation to a human, a validation rejection, a support complaint. Each needs to land somewhere with its trace id.
  2. Triage weekly. Someone looks at the week's failures and sorts them: retrieval problem, prompt problem, model limitation, genuinely out of scope.
  3. Add the interesting ones to the eval set, with the correct answer written by a person. This is the part that compounds.
  4. Fix, and confirm the new case passes without the old ones regressing.
  5. Watch it in production, because the eval set proves the fix works on that case and not that it works.

An eval set that grows from real failures becomes, after a few months, the most valuable artefact your team owns — and the one that makes changing models a measured decision rather than a leap.

Who owns the prompt

An organisational question that causes more production incidents than any technical one.

Prompts sit awkwardly: they read like content, so product and support people reasonably want to edit them, and they behave like code, so a careless edit can break the system silently. The two failure modes are a prompt only engineers can change — so it never improves, because engineers do not read the support queue — and a prompt anyone can change in a dashboard, which is an unversioned production deploy by someone who does not know it is one.

What works: prompts in version control, edited by whoever knows the domain, shipped through the same review and evals as code. The domain expert writes the change, the evals decide whether it goes out, and the history shows who changed what and why. Give non-engineers a real path to contribute, and put the gate on the evals rather than on the person.

Write the post-mortem into the eval set

Ordinary incident reviews end with an action item and a fix. An LLM system gives you somewhere better to put the lesson.

Every incident involving a wrong answer should end with that exact input added to the eval set, with the output you wanted, and a test that would have caught it. The fix goes in the code; the knowledge goes in the set, where it stays and gets checked on every future change — including the model migration two years from now, when nobody involved in the original incident still works here.

That is the difference between a team that accumulates operational knowledge and one that keeps rediscovering the same edge cases. It costs ten minutes at the end of a review and it is the single highest-return habit on this page.

Where to start, in order

The checklist below is the destination. If you are at a working prototype, this is the order that gets you there without stalling.

  1. Twenty eval cases in a file, run by hand. Not a framework — a file. This alone changes how you make decisions.
  2. Pin the model and put the prompt in git. An afternoon, and it makes everything after it possible.
  3. Log metadata and sample content, with a trace id.
  4. Timeouts, retries and a fallback. The cheapest reliability you will ever buy.
  5. Output validation before anything downstream acts on it.
  6. Evals in CI, once the set is worth gating on.
  7. Shadow mode and canary, when a bad release would actually hurt.
  8. Cost per outcome, tracked per feature.

Teams that do the first two get most of the benefit. Teams that start at step seven build impressive infrastructure around a system nobody has measured.

Pre-production checklist

The guiding principle

Treat the LLM system like any critical production service: measurable, reproducible, fault-tolerant and secure. The AI "magic" doesn't exempt you from good engineering — it demands it more.

Next step

Go deeper on the critical production components: monitoring, evals and safety.