Structured Outputs & Tool Use
The step that turns a "chat" into a "product": getting an LLM to return valid JSON and call tools. Without it, you can't build software on top of a language model.
Why it's critical
Free text is great for humans, but code can't work with it. If the model answers "the lead looks solid, about 85, worth following up" — your software doesn't know what the score is. But if it returns {"score": 85, "priority": "high"} — you can feed that straight into a database, a condition, or the next step in a flow.
Structured Outputs are the ability to force the model to return a fixed, valid data structure. Tool Use (or Function Calling) is the next step — the model doesn't just return data, it chooses which function to call and with which parameters. These two are the foundation of any serious automation, agent or integration.
If the LLM's output feeds the next step in the system (rather than going straight to a user's eyes) — it must be structured and validated.
JSON & Structured Output
The basic way: ask for JSON in the prompt and define a schema. Modern models have a dedicated mode that guarantees valid output. Example with OpenAI (Python):
from openai import OpenAI
client = OpenAI()
resp = client.chat.completions.create(
model=MODEL,
response_format={"type": "json_schema", "json_schema": {
"name": "lead_score",
"schema": {
"type": "object",
"properties": {
"score": {"type": "integer", "minimum": 0, "maximum": 100},
"priority": {"type": "string", "enum": ["high", "medium", "low"]},
"reason": {"type": "string"}
},
"required": ["score", "priority", "reason"],
"additionalProperties": False
}
}},
messages=[
{"role": "system", "content": "Score the lead quality from the details."},
{"role": "user", "content": "A SaaS company, 50 employees, defined budget, high urgency."}
],
)
import json
data = json.loads(resp.choices[0].message.content)
print(data["score"], data["priority"])
The schema guarantees score is a number between 0 and 100 and priority is one of the allowed values. That's far more reliable than "ask for JSON" in the prompt alone.
Tool Use / Function Calling
Here the model decides on its own which tools to call. You define the available tools, and the model returns which one to call and with which parameters. Example with Anthropic (Claude):
import anthropic
client = anthropic.Anthropic()
tools = [{
"name": "get_weather",
"description": "Returns the current weather in a given city",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}]
msg = client.messages.create(
model=MODEL,
max_tokens=1024,
tools=tools,
messages=[{"role": "user", "content": "What's the weather in Tel Aviv?"}],
)
for block in msg.content:
if block.type == "tool_use":
print(block.name, block.input) # get_weather {'city': 'Tel Aviv'}
# here you run the real function and return the result to the model
The full flow: (1) the model asks to call a tool; (2) your code runs the real function; (3) you return the result to the model; (4) the model composes a final answer. This is exactly the mechanism behind AI agents and MCP.
Validation & Retry — don't trust, verify
Even with json_schema, it's worth validating on your side before feeding the system. In Python people use Pydantic; in TypeScript — Zod.
from pydantic import BaseModel, Field, ValidationError
class LeadScore(BaseModel):
score: int = Field(ge=0, le=100)
priority: str
reason: str
def parse_with_retry(client, messages, tries=2):
for attempt in range(tries):
raw = call_model(client, messages) # call the model
try:
return LeadScore.model_validate_json(raw) # validate
except ValidationError as e:
# return the error to the model and ask it to fix
messages.append({"role": "user",
"content": f"The output was invalid: {e}. Return valid JSON only."})
raise RuntimeError("failed to get valid output")
This pattern — try, validate, and if it fails return the error to the model and ask for a fix — is an industry standard. It turns a fragile system into a stable one.
Schema enforcement is not the model trying harder
There are two quite different things that both get called structured output, and knowing which one you have determines what can still go wrong.
Asking nicely — "return JSON only" in the prompt — leaves the model free to produce anything. It usually complies and occasionally wraps the JSON in an explanation, adds a trailing comma, or opens with "Here is the JSON:". Your parser then fails on a Tuesday.
Constrained decoding — the provider's schema mode — is a different mechanism. The generation itself is restricted so that only tokens which keep the output valid against your schema can be chosen. The result is not "more likely to be valid"; within the schema's rules, it cannot be invalid.
That distinction matters because it tells you which problems are solved and which are not:
- Solved: malformed JSON, missing required fields, wrong types, values outside an enum, extra fields where you forbade them.
- Not solved: whether the values are correct. A schema guarantees you get a date in the right format, not the right date. It will happily give you a well-typed, confidently wrong answer.
- Not solved: semantic consistency between fields — a start date after an end date, a total that does not match the line items.
So schema mode removes an entire class of parsing bug and none of the accuracy work. Validation on your side is still required, for different reasons than before.
Designing a schema the model can fill correctly
A schema is an instruction, not only a contract. How you write it changes how often the values are right.
- Describe every field. Schema formats allow a description per property, and models read them. "due_date — the date the invoice must be paid, not the date it was issued" prevents a whole category of confusion that a field name alone cannot.
- Enums over free strings wherever the set is closed. Free text means you will be normalising "High", "high" and "HIGH" forever.
- Flat over deeply nested. Nested structures are constructed incorrectly more often, and they are harder to validate usefully.
- Make optional things genuinely optional — or better, use an explicit null. A required field the model cannot determine forces it to invent a value, which is the worst outcome available.
- Add a confidence or a not-found path. A schema with no way to express "this was not in the document" guarantees fabrication, because the only valid outputs are values.
- Order fields so reasoning comes first. Generation is sequential, so a "reason" field placed before "score" lets the model work through the problem before committing to the answer. The same fields in the opposite order produce a number and then a justification of it.
Retrying well, and when to stop
The try-validate-return-the-error loop is the right pattern and it has failure modes worth building against.
- Bound the attempts — two or three. A model that failed the schema twice is usually failing for a reason another attempt will not fix, and an unbounded loop is how a small bug becomes a large bill.
- Return the specific error, not a generic complaint. "score must be between 0 and 100; received 150" is actionable. "Invalid output" gives the model nothing to correct.
- Do not resend the whole conversation on every retry. Send the original request plus the error; a growing transcript of failed attempts is expensive and pushes the model towards repeating itself.
- Distinguish a schema failure from a semantic one. A malformed value is worth retrying. A value that parsed but is wrong will not improve by asking again, because nothing told the model it was wrong.
- Decide what happens on final failure — a null result, a queue for human review, an exception. Silently returning the last invalid attempt is the default in hand-rolled code and the worst option.
- Log the failures. A rising rate of schema violations is a signal: the model changed, your inputs changed, or your schema is asking for something the data does not contain.
Validation the schema cannot do
Once the shape is guaranteed, your validation layer should be checking meaning rather than syntax — which is a different set of rules and usually more valuable.
- Cross-field arithmetic. Line items sum to the subtotal; subtotal plus tax equals the total. This catches misread digits that no schema can.
- Ranges that reflect your business, not just the type. An invoice a thousand times your usual amount is technically a valid number.
- Referential checks. Does this customer id exist? Is this supplier one you have? Matching against your own records beats trusting an extracted name.
- Internal consistency. Dates in a sensible order, a status compatible with the other fields.
- Provenance. For extraction, ask for a verbatim quote alongside each value and confirm the quote appears in the source. This is the single strongest check available and it costs a string search.
Then route by what failed: a syntax failure is a retry, a semantic failure is a human. Treating them the same means either bothering people with things the model could fix, or silently accepting things it could not.
What structure costs you
Worth knowing, because the trade-offs are real and rarely mentioned.
Tokens. A schema is sent with every request, and JSON syntax — braces, quotes, repeated field names — is charged like any other output. For a high-volume extraction job this is a meaningful share of the bill. Shorter field names and a flatter structure genuinely save money at scale.
Quality, sometimes. Constraining generation can reduce the model's ability to reason freely on the way to an answer. The usual mitigation is the ordering point above — put a free-text reasoning field first — which recovers most of it at the cost of tokens you then discard.
Portability. Schema modes differ between providers in what they support: nesting depth, unions, optional handling, whether the schema must be strict. A schema tuned for one provider may need adjusting for another, which matters if your fallback model is from a different vendor.
Latency. Usually negligible, occasionally not for large schemas, and worth measuring rather than assuming if you are latency-sensitive.
When not to force structure
- When a human reads the output. Prose is the right format for prose. Forcing a chat answer through a schema makes it worse for the reader and harder for you.
- When the shape is genuinely unknown — exploratory analysis, open-ended summarising. Defining a schema before you know what is in the data forces the answer into a shape you guessed.
- For a single boolean. If you need a yes or a no, ask for a yes or a no and check the first word. A schema is more machinery than the job needs.
- When it would hide the uncertainty. A structured answer looks authoritative; if the honest output is "it depends", a schema with no field for that turns nuance into a false precision.
Schemas change, and the data outlives them
The schema you ship in week one will not be the schema you are running in month six. A field gets added, an enum gains a value, something you modelled as a string turns out to need a nested object. That is normal. What catches people out is that every record you have already extracted and stored was produced under an older schema, and nothing in the model or the provider knows or cares about that.
So write the schema version into the record itself, as a plain field, from the very first run. It costs one integer and it is the difference between "which of these 40,000 rows were extracted before we fixed the date handling?" being a one-line query and being an afternoon of archaeology. Bump it whenever the meaning of a field changes — not when you only add an optional one, which old readers can ignore safely.
Treat additions and removals differently. Adding an optional field is cheap: old records simply lack it, and anything reading them has to handle its absence anyway. Removing or renaming a field is a migration, because code downstream is already reading it. Changing what a field means while keeping its name is the worst of the three, because nothing breaks loudly — the pipeline keeps running and quietly mixes two different definitions in one column. If you must redefine a field, give it a new name and deprecate the old one.
Enums deserve their own note. A closed enum is what makes downstream code safe to write, but it also means the day reality contains a value you did not anticipate, the model has to pick a wrong one — it has no way to say "none of these." Include an explicit other member, paired with a free-text field that captures what it actually saw. Then read that field periodically: it is the cheapest signal you will ever get about where your schema has drifted from the world it is supposed to describe.
And keep old schema definitions in version control rather than only in the running code. When a record from six months ago looks wrong, the question is almost always "was it wrong, or was it right under the rules we had then?" — and you can only answer that if those rules still exist somewhere.
Streaming a structured answer
Streaming and schemas pull in opposite directions. Streaming exists so the user sees something immediately; a schema is only meaningful once it is complete. A half-arrived JSON object is not a small JSON object — it is not JSON at all, and a parser will reject it right up until the final brace lands.
There are partial-parsing libraries that will close the open braces for you and hand back a usable object mid-stream, and for a visible UI that is often the right call: fields appear as they fill, and the interface feels fast. Just be clear about what you are showing. A field that has arrived is not necessarily a field the model has finished reasoning about, and one that is still empty may be about to be filled — so do not let anything irreversible fire off a partial read, and do not render a half-arrived number as if it were final.
For everything that is not a UI — a pipeline step, a tool call, anything that writes to a database — wait for the complete object and validate it as a whole. You gain nothing from starting early on work you may have to undo, and cross-field checks cannot run on a fragment by definition.
The useful middle ground: stream for the human, buffer for the machine. If one call feeds both, let the interface show progress from the partial parse while the actual decision waits for the validated object. They are two different consumers of the same stream, and only one of them is in a hurry.
Tips & common mistakes
- Low temperature (0–0.2) for structured outputs — less "creativity," more consistency.
- additionalProperties: false in the schema — prevents the model from adding unexpected fields.
- Clear field names help the model —
due_datebeatsd. - Don't ask for free text + JSON in the same answer. It breaks parsing. Separate the calls.
- Always wrap in try/except. Even the best model can occasionally return invalid output.
- Limit the number of tools. Too many tools confuse the model — group by context.
Next step
Now that the output is structured — the next step is to connect knowledge (RAG) and measure quality (Evals).