Skip to main content
Level: Advanced Updated: September 2026

Tool Use — tools for agents

A good agent is only as good as the tools you gave it. How to design tools the model understands, calls correctly, and doesn't break. The practical guide.

In the Structured Outputs & Tool Use guide we saw the basic mechanics of function calling. Here we go deeper into design — how to build tools an AI agent can actually use successfully, which is the art that separates an agent that works from one that gets stuck.

The agent loop — a reminder

An agent runs in a loop: the model gets a goal + a list of tools, decides which to call, the code runs them and returns results, and repeats until the task is done.

while not done:
    response = model.generate(messages, tools=TOOLS)
    if response.tool_calls:
        for call in response.tool_calls:
            result = run_tool(call.name, call.input)   # your code runs it
            messages.append(tool_result(call.id, result))
    else:
        done = True   # the model finished and returned a final answer

The tools are the model's "hands." If they're poorly designed — the agent will choose wrong, send wrong parameters, or get into loops.

Designing a good tool — 6 principles

  1. Clear name and description. The description is the tool's "prompt." get_order_status with the description "returns an order's status by order number" is far better than lookup with no explanation. Write when to use it and when not to.
  2. Few, well-defined parameters. Each parameter with a description, a type and an enum where possible. Fewer parameters = fewer errors.
  3. One tool = one task. Don't build a mega-tool that does ten things by a flag. Split into focused tools.
  4. Concise results. Return only what the model needs. Don't return 5,000 lines of JSON — filter/summarize (see Context Engineering).
  5. Limit the number of tools. Too many tools confuse it. If there are many, group by context or use dynamic retrieval.
  6. Consistent naming. The same convention for all tools — it's easy for the model to learn the pattern.
The key insight

Most "agent failures" are actually tool-design failures — a vague description or a bloated result. Improve the tools before you blame the model.

Error handling — don't let a tool crash the agent

A tool that fails should return a clear error message to the model, not crash. The model can read the error and try differently:

def run_tool(name, args):
    try:
        return TOOLS[name](**args)
    except Exception as e:
        # return the error to the model as a result, don't raise
        return {"error": str(e), "hint": "check the parameters and try again"}

Design tools for a capable colleague with no memory

The most useful mental model when designing a tool: the caller is a competent new colleague who reads only your documentation, has never seen your system, cannot ask anyone, and will forget everything between calls.

That framing answers most design questions immediately. Would a new colleague know that status takes the values 1 through 7? No — so use an enum with names. Would they know they must call one endpoint before another? Only if you wrote it down. Would they guess your date format? They would guess wrong.

It also explains the commonest class of failure. Agents rarely fail because the model is weak; they fail because the tool assumed knowledge that exists only in your team's heads. Before blaming the model, read your tool descriptions as though you had never seen the system.

The description is the most important code you will write

A tool description is not documentation — it is the prompt that decides whether the tool gets chosen, and chosen correctly. It deserves the attention you would give a function's implementation.

A good one covers four things:

Compare: "Searches orders." against "Finds a customer's orders by email address. Use when the customer cannot supply an order number. Do not use to check the status of a known order — use get_order_status for that. Returns up to ten recent orders, newest first, or an empty list." Same function, and the second one gets called correctly.

One habit worth adopting: when an agent misuses a tool, fix the description rather than the prompt. The instruction lives next to the thing it describes, applies everywhere the tool is used, and survives the next prompt rewrite.

Parameters, and making wrong ones impossible

Every parameter is a chance for the model to be wrong. The design goal is to make the wrong value unrepresentable rather than to validate it afterwards.

Error messages are instructions

Returning the error to the model rather than crashing is right, and the wording of that error determines whether the next attempt succeeds or repeats the mistake.

An error the model can act on says three things: what went wrong, why, and what to do instead.

That last pattern matters enormously: say explicitly when not to retry. Without it, a model faced with a permanent failure will try again, and again, burning tokens and time on something that cannot succeed. Distinguish retryable errors from terminal ones in the message itself.

And keep internal details out. A stack trace tells the model nothing useful and may leak your schema into a conversation.

What a tool returns is context you are spending

Every token a tool returns sits in the context for the rest of the conversation, is re-read on every subsequent step, and is paid for each time. Tool results are the largest source of context bloat in most agents.

A useful check: look at the full context after a five-step agent run. If tool output dominates it, that is where your cost and your degradation are coming from, and trimming it is usually the highest-return change available.

Stopping conditions and runaway agents

An agent loop can fail to terminate in ways ordinary code does not, and the defences belong in your loop rather than in the prompt.

Designing the approval step so it is not a rubber stamp

Marking a tool as requiring approval is the easy half. The hard half is that an approval gate a person clicks through without reading is worse than no gate, because it creates a record of authorisation for something nobody examined.

What makes it real:

The underlying principle is that people become blind to any confirmation they see often. Reserve the gate for actions that genuinely warrant it, and design everything else to be safely reversible so it does not need one.

Testing tools without the model

Agent bugs are hard to reproduce because the model's choices vary. Splitting the testing makes them tractable.

Test the tools as ordinary functions. They are ordinary functions. Unit tests for the happy path, the empty result, the malformed input, the permission failure — none of which needs a model and all of which catch real bugs.

Test selection separately. Take twenty realistic user requests and check which tool the model picks and with what arguments, without executing anything. This isolates description quality from everything else, and it is where most improvement comes from.

Then test end to end on a handful of full scenarios, including ones designed to fail — a tool erroring mid-task, an ambiguous request, an action requiring approval.

Keep all three as a set and re-run them when you change a description, add a tool or change models. Adding a tool changes selection behaviour for the tools already there, which is the regression people do not expect.

Make every action safe to repeat

An agent retries. It retries after a timeout, after an ambiguous error, and sometimes because it lost track of whether the last call succeeded. Ordinary software retries too, but an agent decides to retry based on a judgement rather than a rule, which makes it less predictable and the safeguard more necessary.

So every tool that changes something should be safe to call twice with the same arguments:

Then test it directly: call each writing tool twice with identical arguments and confirm the world changed once. This takes minutes and catches the failure that is otherwise discovered by a customer receiving the same email three times.

What happens as the toolset grows

Tool design that works with six tools stops working at forty, and the symptoms are specific: the model picks plausible-but-wrong tools, mixes up similar ones, and spends steps exploring.

What helps, in order:

Human-in-the-loop — for dangerous tools

Not every action should run automatically. Tools that perform an irreversible action (sending an email to a customer, making a payment, deleting data) need human approval before running. Design the flow so the agent proposes the action, and a human approves.

Orchestrating multiple tools

When a task requires a sequence of tools (search → filter → act), the model manages the order. For this to work:

Security — tools are an attack surface

A tool gives the model the ability to act in the real world — and that's dangerous if the model is "influenced" by prompt injection:

Next step

Build a full agent, connect tools in a standard way with MCP, and secure it.