Skip to main content
Guides AI Observability
Updated July 2026 13 min read Advanced

AI Observability
— seeing what the agent does

An AI agent that "just does not work" is a debugging nightmare — non-deterministic input, a chain of LLM and tool calls, and climbing costs. Observability turns the black box transparent: every trace, every prompt, every token and cost. In this guide: the four pillars, tracing and evals, and LangSmith vs Langfuse.

On model versions: the code on this page pins gpt-4o. As of September 2026 providers retire versions and model names change — check the current model list in OpenAI's documentation and swap the string before running. What the code demonstrates does not depend on the version.

Tracing
What happened
Evals
How good
Cost
How much
Latency
How fast

Why observability is critical for AI systems

In regular software, the same input always returns the same output. In an LLM-based system it does not: the same question can return different answers, an agent may pick a different tool path on each run, and the cost varies from request to request. Without a dedicated monitoring tool, you are "blind" — you do not know why the agent decided what it did, where it got stuck, and how much it cost.

AI Observability is the equivalent of monitoring in the AI world: it records the entire chain of calls to anagent and to an LLM, measures quality with evals, and tracks costs and latency — so you can debug, improve and trust the system in production.

The four pillars

Tracing
A full record of every step: the prompt, the response, tool calls, retrieval — a complete tree of what happened in the request.
Evaluation
Measuring the quality of the answers — automatically (LLM-as-judge, tests) or with human rating.
Cost & Tokens
How many tokens and how much money each request, user or feature consumes. Prevents billing surprises.
Latency
Response times for each step — to identify which model or tool is slowing the system down.

Tracing — the heart of the system

A trace is a record of a whole request, broken intospans — each step in the chain. In a typical agent, one trace contains the initial prompt, the LLM decision, a call to a search tool, a retrieval from avector DB, and the final answer — each with its input, output, time and cost.

When something goes wrong, you open the trace and see exactly where: maybe the retrieval returned irrelevant documents, maybe the tool failed, maybe the prompt was confusing. That is the difference between minutes of debugging and hours of guessing.

OpenTelemetry — the emerging standard

In 2026 there is convergence around OpenTelemetry (OTel) as an open standard for LLM tracing. Most tools support it, so you are not locked to a single vendor — the same instrumentation can send data to Langfuse, LangSmith or any other backend.

Evals — how you know it actually works

"Looks good" is not a metric. Evals are systematic tests of output quality, and without them every "improvement" to a prompt is a gamble. Three main approaches:

The real power: evals alongside CI/CD. You run a set of examples (a dataset) on every change to a prompt or model, and see whether quality rose or fell — before it reaches users. And yes — security testing is part of evals.

LangSmith vs Langfuse — what to choose

Tool Model Best for
LangSmithManaged (closed)Those already on LangChain/LangGraph
LangfuseOpen source + cloudSelf-host, framework-agnostic
Arize / PhoenixOpen source + cloudAdvanced evals, ML too
HeliconeLightweight proxyCost tracking & caching, fast

What to log, and what never to log

Tracing an LLM system means recording the prompt, and the prompt is usually the most sensitive thing in the request. A support agent's trace contains the customer's message. A document assistant's trace contains the document. This is a real difference from ordinary application monitoring, where you log a request id and a status code and almost never the payload — here the payload is the diagnostic information, which is exactly why it is tempting to keep all of it forever.

Decide the policy before you instrument, because retrofitting redaction onto a year of stored traces is miserable work. The questions are the ordinary data-protection ones, they just arrive somewhere new: what categories of personal data can appear in a prompt, where the traces are stored and under whose control, how long they are kept, and who can read them.

A few things that help in practice. Redact at the point of instrumentation rather than at the point of display — a value that never leaves your process cannot leak from the dashboard, and masking in the UI is not the same protection. Prefer structure over free text where you can: if the prompt is assembled from fields, log the fields separately so the sensitive ones can be dropped individually instead of having to pattern-match a blob. Set a retention period short enough that the archive is not a liability and long enough to debug a report from last week; a fortnight is a reasonable starting point for full payloads, with aggregates kept much longer. And treat access to traces as access to production data, because that is what it is.

Two things that specifically catch people out. Tool results go into the trace as well, and a tool that queries your database will happily put a row of customer records into a span. And an error path often logs more than the happy path — the stack trace with the full request attached is how sensitive data most often ends up somewhere it was not meant to be.

You cannot keep every trace

At a hundred requests a day, store everything. At a million, storing everything is a cost centre of its own, and the hard part becomes deciding what to throw away.

The instinct is to sample randomly — keep one in a hundred — which is the worst option available, because the traces you actually need are by definition rare. A one-percent random sample of a system with a one-percent failure rate gives you roughly nothing to look at.

What works is deciding after the fact instead of before. Buffer the trace, let the request finish, and then keep it if anything about it was interesting: an error, a latency outlier, a retry, an unusually high cost, a thumbs-down from the user, a low eval score, a guardrail that fired. Sample the boring successful majority at a low rate purely to keep a baseline for comparison. This costs a little more memory during the request and gives you a store where almost everything in it is worth reading.

Keep the aggregates at full fidelity regardless. Counts, cost, token totals and latency percentiles are small, and you want them computed over every request rather than over the sample — otherwise your cost dashboard is an estimate, and an estimate with an unhelpfully wide error bar at exactly the moment something goes wrong.

One more thing worth building early: a way to force full tracing for a specific user or session on demand. When someone reports that the assistant gave them nonsense yesterday, being able to turn on complete capture for their next attempt is worth more than any amount of retrospective sampling policy.

Alerts that are worth being woken by

A dashboard nobody opens is not observability. The part that changes outcomes is the small set of alerts that fire before a user tells you, and the trick is picking things that are both measurable and actionable.

The ones that consistently earn their place: spend rate, measured per hour rather than per month, because a runaway loop can burn a monthly budget in an afternoon and a monthly total tells you about it far too late. Error rate by type, split so that rate limits, timeouts and parse failures are visible separately — they have completely different causes and blurring them together hides all three. Latency at the 95th percentile, not the mean, since the mean of a distribution with a long tail mostly tells you about the requests nobody was complaining about. Tool failure rate, per tool, which is usually where an agent's problems actually originate. And volume in both directions — a sudden drop is as diagnostic as a spike, and is the classic signature of something upstream having quietly broken.

Two that are worth the extra effort to set up. A fallback-rate alert, if your system degrades to a cheaper model or a canned response under load: silent degradation is the failure mode that survives longest, because every individual response still looks fine. And a regression alert on eval scores from your scheduled runs, which is the closest thing to an alert on quality itself.

What not to do is alert on individual bad outputs. Quality is a distribution, and a single poor answer is normal; paging someone about it trains them to ignore the channel within a week. Watch the rate, and route individual cases to a queue that gets reviewed rather than to a notification.

Reading a trace when something went wrong

Opening a long agent trace for the first time is disorienting — dozens of spans, most of them fine. A fixed order to look in saves a great deal of time.

Start at the end and work backwards. The final answer is wrong; the span that produced it had some input; was that input already wrong? Following the badness upstream converges much faster than reading forward from the top, because it skips everything that happened to be irrelevant.

At each step the question is narrow: was this step given what it needed? Most agent failures are not the model reasoning badly, they are the model reasoning correctly about the wrong material — a retrieval that returned three plausible but unrelated passages, a tool that returned an empty list and was read as "nothing matches" rather than "the query was malformed", a summary from an earlier step that dropped the one constraint that mattered.

Then check the shape of the run rather than its content. An agent that took eleven steps to do a three-step job was lost, even if it arrived. Repeated identical tool calls mean it was not learning from the results. A step whose output is enormous is usually the one that poisoned everything after it.

And compare against a trace of the same request that went well, if you have one. The difference between a good run and a bad run is far easier to see than the flaw in a bad run read on its own — which is, incidentally, the strongest practical argument for keeping that low-rate sample of ordinary successful requests.

Cost you can attribute

A total monthly spend figure tells you whether to worry. It does not tell you what to do, and the gap between those two is filled entirely by tagging.

Attach metadata to every call at the moment you make it: which feature, which user or account, which environment, which prompt version, which model. It costs nothing and it is nearly impossible to reconstruct later — a trace without a feature tag cannot be assigned to a feature retrospectively, no matter how good the tool is. The tags are what turn "we spent more this month" into "the document summariser is responsible for 70% of it, and 90% of that comes from eleven accounts."

That distribution is the normal shape, and it is worth looking for deliberately. Usage is almost always concentrated, and the interesting question is whether the heaviest users are the ones paying you the most. In a per-seat product with usage-based costs underneath, a small number of accounts can be individually unprofitable while the aggregate looks healthy, and nothing but per-account attribution will show you that.

Tag the prompt version too. It is the only way to answer, three weeks after a change, whether the cost per request went up because of the edit or because of something else entirely — and that question comes up more often than any other.

Turning production into your test set

The most valuable thing observability produces is not the dashboard. It is a supply of real failures, and the loop that converts them into tests is what separates a system that improves from one that merely gets watched.

The mechanism is deliberately dull. Every time a genuine failure is found — from a complaint, a thumbs-down, a review of the low-scoring queue — the case gets added to the eval dataset with the output that would have been correct. It runs on every subsequent change. That is the whole discipline, and its effect compounds: after six months the dataset is a precise map of the ways this particular system fails, which no generic benchmark can give you.

Make the path from trace to dataset one click if you possibly can. Any friction here and it stops happening within a fortnight — the failure gets fixed, everyone moves on, and the same class of failure returns two months later with nothing to catch it.

Two habits keep the dataset honest. Keep the cases that are fixed, rather than retiring them; they are now regression tests, and the whole point is noticing when a fix stops holding. And record why each case was added, in a sentence — six months on, a bare input-output pair with no explanation is a case nobody dares to change and nobody fully trusts.

How to start — the first step

Do not wait until it is "tidy". Add tracing right now — it is a handful of lines. With Langfuse it looks like this:

from langfuse.openai import openai  # transparent wrapper

resp = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role":"user","content":"Summarize the document"}],
)
# every call is logged automatically: prompt, response, tokens, cost, latency
The recommended order of operations

1. Add tracing on day one — to see what happens. 2. Set up cost tracking with an alert on overruns. 3. Build a dataset of examples and run evals on every change. 4. Connect feedback from real users into that same dataset. Observability is not a one-off project — it is a continuous improvement loop.