Skip to main content
AI Engineering Structured Outputs & Tool Use
Level: Beginner–Intermediate Updated: September 2026

Structured Outputs & Tool Use

The step that turns a "chat" into a "product": getting an LLM to return valid JSON and call tools. Without it, you can't build software on top of a language model.

Why it's critical

Free text is great for humans, but code can't work with it. If the model answers "the lead looks solid, about 85, worth following up" — your software doesn't know what the score is. But if it returns {"score": 85, "priority": "high"} — you can feed that straight into a database, a condition, or the next step in a flow.

Structured Outputs are the ability to force the model to return a fixed, valid data structure. Tool Use (or Function Calling) is the next step — the model doesn't just return data, it chooses which function to call and with which parameters. These two are the foundation of any serious automation, agent or integration.

Rule of thumb

If the LLM's output feeds the next step in the system (rather than going straight to a user's eyes) — it must be structured and validated.

JSON & Structured Output

The basic way: ask for JSON in the prompt and define a schema. Modern models have a dedicated mode that guarantees valid output. Example with OpenAI (Python):

from openai import OpenAI
client = OpenAI()

resp = client.chat.completions.create(
    model=MODEL,
    response_format={"type": "json_schema", "json_schema": {
        "name": "lead_score",
        "schema": {
            "type": "object",
            "properties": {
                "score": {"type": "integer", "minimum": 0, "maximum": 100},
                "priority": {"type": "string", "enum": ["high", "medium", "low"]},
                "reason": {"type": "string"}
            },
            "required": ["score", "priority", "reason"],
            "additionalProperties": False
        }
    }},
    messages=[
        {"role": "system", "content": "Score the lead quality from the details."},
        {"role": "user", "content": "A SaaS company, 50 employees, defined budget, high urgency."}
    ],
)
import json
data = json.loads(resp.choices[0].message.content)
print(data["score"], data["priority"])

The schema guarantees score is a number between 0 and 100 and priority is one of the allowed values. That's far more reliable than "ask for JSON" in the prompt alone.

Tool Use / Function Calling

Here the model decides on its own which tools to call. You define the available tools, and the model returns which one to call and with which parameters. Example with Anthropic (Claude):

import anthropic
client = anthropic.Anthropic()

tools = [{
    "name": "get_weather",
    "description": "Returns the current weather in a given city",
    "input_schema": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
    }
}]

msg = client.messages.create(
    model=MODEL,
    max_tokens=1024,
    tools=tools,
    messages=[{"role": "user", "content": "What's the weather in Tel Aviv?"}],
)

for block in msg.content:
    if block.type == "tool_use":
        print(block.name, block.input)   # get_weather {'city': 'Tel Aviv'}
        # here you run the real function and return the result to the model

The full flow: (1) the model asks to call a tool; (2) your code runs the real function; (3) you return the result to the model; (4) the model composes a final answer. This is exactly the mechanism behind AI agents and MCP.

Validation & Retry — don't trust, verify

Even with json_schema, it's worth validating on your side before feeding the system. In Python people use Pydantic; in TypeScript — Zod.

from pydantic import BaseModel, Field, ValidationError

class LeadScore(BaseModel):
    score: int = Field(ge=0, le=100)
    priority: str
    reason: str

def parse_with_retry(client, messages, tries=2):
    for attempt in range(tries):
        raw = call_model(client, messages)          # call the model
        try:
            return LeadScore.model_validate_json(raw)  # validate
        except ValidationError as e:
            # return the error to the model and ask it to fix
            messages.append({"role": "user",
                "content": f"The output was invalid: {e}. Return valid JSON only."})
    raise RuntimeError("failed to get valid output")

This pattern — try, validate, and if it fails return the error to the model and ask for a fix — is an industry standard. It turns a fragile system into a stable one.

Schema enforcement is not the model trying harder

There are two quite different things that both get called structured output, and knowing which one you have determines what can still go wrong.

Asking nicely — "return JSON only" in the prompt — leaves the model free to produce anything. It usually complies and occasionally wraps the JSON in an explanation, adds a trailing comma, or opens with "Here is the JSON:". Your parser then fails on a Tuesday.

Constrained decoding — the provider's schema mode — is a different mechanism. The generation itself is restricted so that only tokens which keep the output valid against your schema can be chosen. The result is not "more likely to be valid"; within the schema's rules, it cannot be invalid.

That distinction matters because it tells you which problems are solved and which are not:

So schema mode removes an entire class of parsing bug and none of the accuracy work. Validation on your side is still required, for different reasons than before.

Designing a schema the model can fill correctly

A schema is an instruction, not only a contract. How you write it changes how often the values are right.

Retrying well, and when to stop

The try-validate-return-the-error loop is the right pattern and it has failure modes worth building against.

Validation the schema cannot do

Once the shape is guaranteed, your validation layer should be checking meaning rather than syntax — which is a different set of rules and usually more valuable.

Then route by what failed: a syntax failure is a retry, a semantic failure is a human. Treating them the same means either bothering people with things the model could fix, or silently accepting things it could not.

What structure costs you

Worth knowing, because the trade-offs are real and rarely mentioned.

Tokens. A schema is sent with every request, and JSON syntax — braces, quotes, repeated field names — is charged like any other output. For a high-volume extraction job this is a meaningful share of the bill. Shorter field names and a flatter structure genuinely save money at scale.

Quality, sometimes. Constraining generation can reduce the model's ability to reason freely on the way to an answer. The usual mitigation is the ordering point above — put a free-text reasoning field first — which recovers most of it at the cost of tokens you then discard.

Portability. Schema modes differ between providers in what they support: nesting depth, unions, optional handling, whether the schema must be strict. A schema tuned for one provider may need adjusting for another, which matters if your fallback model is from a different vendor.

Latency. Usually negligible, occasionally not for large schemas, and worth measuring rather than assuming if you are latency-sensitive.

When not to force structure

Schemas change, and the data outlives them

The schema you ship in week one will not be the schema you are running in month six. A field gets added, an enum gains a value, something you modelled as a string turns out to need a nested object. That is normal. What catches people out is that every record you have already extracted and stored was produced under an older schema, and nothing in the model or the provider knows or cares about that.

So write the schema version into the record itself, as a plain field, from the very first run. It costs one integer and it is the difference between "which of these 40,000 rows were extracted before we fixed the date handling?" being a one-line query and being an afternoon of archaeology. Bump it whenever the meaning of a field changes — not when you only add an optional one, which old readers can ignore safely.

Treat additions and removals differently. Adding an optional field is cheap: old records simply lack it, and anything reading them has to handle its absence anyway. Removing or renaming a field is a migration, because code downstream is already reading it. Changing what a field means while keeping its name is the worst of the three, because nothing breaks loudly — the pipeline keeps running and quietly mixes two different definitions in one column. If you must redefine a field, give it a new name and deprecate the old one.

Enums deserve their own note. A closed enum is what makes downstream code safe to write, but it also means the day reality contains a value you did not anticipate, the model has to pick a wrong one — it has no way to say "none of these." Include an explicit other member, paired with a free-text field that captures what it actually saw. Then read that field periodically: it is the cheapest signal you will ever get about where your schema has drifted from the world it is supposed to describe.

And keep old schema definitions in version control rather than only in the running code. When a record from six months ago looks wrong, the question is almost always "was it wrong, or was it right under the rules we had then?" — and you can only answer that if those rules still exist somewhere.

Streaming a structured answer

Streaming and schemas pull in opposite directions. Streaming exists so the user sees something immediately; a schema is only meaningful once it is complete. A half-arrived JSON object is not a small JSON object — it is not JSON at all, and a parser will reject it right up until the final brace lands.

There are partial-parsing libraries that will close the open braces for you and hand back a usable object mid-stream, and for a visible UI that is often the right call: fields appear as they fill, and the interface feels fast. Just be clear about what you are showing. A field that has arrived is not necessarily a field the model has finished reasoning about, and one that is still empty may be about to be filled — so do not let anything irreversible fire off a partial read, and do not render a half-arrived number as if it were final.

For everything that is not a UI — a pipeline step, a tool call, anything that writes to a database — wait for the complete object and validate it as a whole. You gain nothing from starting early on work you may have to undo, and cross-field checks cannot run on a fragment by definition.

The useful middle ground: stream for the human, buffer for the machine. If one call feeds both, let the interface show progress from the partial parse while the actual decision waits for the validated object. They are two different consumers of the same stream, and only one of them is in a hurry.

Tips & common mistakes

Next step

Now that the output is structured — the next step is to connect knowledge (RAG) and measure quality (Evals).