Tool Use — tools for agents
A good agent is only as good as the tools you gave it. How to design tools the model understands, calls correctly, and doesn't break. The practical guide.
In the Structured Outputs & Tool Use guide we saw the basic mechanics of function calling. Here we go deeper into design — how to build tools an AI agent can actually use successfully, which is the art that separates an agent that works from one that gets stuck.
The agent loop — a reminder
An agent runs in a loop: the model gets a goal + a list of tools, decides which to call, the code runs them and returns results, and repeats until the task is done.
while not done:
response = model.generate(messages, tools=TOOLS)
if response.tool_calls:
for call in response.tool_calls:
result = run_tool(call.name, call.input) # your code runs it
messages.append(tool_result(call.id, result))
else:
done = True # the model finished and returned a final answer
The tools are the model's "hands." If they're poorly designed — the agent will choose wrong, send wrong parameters, or get into loops.
Designing a good tool — 6 principles
- Clear name and description. The description is the tool's "prompt."
get_order_statuswith the description "returns an order's status by order number" is far better thanlookupwith no explanation. Write when to use it and when not to. - Few, well-defined parameters. Each parameter with a description, a type and an enum where possible. Fewer parameters = fewer errors.
- One tool = one task. Don't build a mega-tool that does ten things by a flag. Split into focused tools.
- Concise results. Return only what the model needs. Don't return 5,000 lines of JSON — filter/summarize (see Context Engineering).
- Limit the number of tools. Too many tools confuse it. If there are many, group by context or use dynamic retrieval.
- Consistent naming. The same convention for all tools — it's easy for the model to learn the pattern.
Most "agent failures" are actually tool-design failures — a vague description or a bloated result. Improve the tools before you blame the model.
Error handling — don't let a tool crash the agent
A tool that fails should return a clear error message to the model, not crash. The model can read the error and try differently:
def run_tool(name, args):
try:
return TOOLS[name](**args)
except Exception as e:
# return the error to the model as a result, don't raise
return {"error": str(e), "hint": "check the parameters and try again"}
- Validation errors: if a parameter is invalid, tell the model exactly what's wrong.
- Limit attempts. If the agent fails the same tool 3 times — stop and escalate, don't enter an infinite loop.
- Timeouts: a slow tool (an external API) needs a timeout so it doesn't hang the agent.
Design tools for a capable colleague with no memory
The most useful mental model when designing a tool: the caller is a competent new colleague who reads only your documentation, has never seen your system, cannot ask anyone, and will forget everything between calls.
That framing answers most design questions immediately. Would a new colleague know that status takes the values 1 through 7? No — so use an enum with names. Would they know they must call one endpoint before another? Only if you wrote it down. Would they guess your date format? They would guess wrong.
It also explains the commonest class of failure. Agents rarely fail because the model is weak; they fail because the tool assumed knowledge that exists only in your team's heads. Before blaming the model, read your tool descriptions as though you had never seen the system.
The description is the most important code you will write
A tool description is not documentation — it is the prompt that decides whether the tool gets chosen, and chosen correctly. It deserves the attention you would give a function's implementation.
A good one covers four things:
- What it does, in one sentence, in plain language.
- When to use it — the situation, not the mechanism.
- When not to use it, especially against the tool it is most likely to be confused with. This line does more than any other to prevent wrong selection.
- What it returns, and what it does when there is nothing to return.
Compare: "Searches orders." against "Finds a customer's orders by email address. Use when the customer cannot supply an order number. Do not use to check the status of a known order — use get_order_status for that. Returns up to ten recent orders, newest first, or an empty list." Same function, and the second one gets called correctly.
One habit worth adopting: when an agent misuses a tool, fix the description rather than the prompt. The instruction lives next to the thing it describes, applies everywhere the tool is used, and survives the next prompt rewrite.
Parameters, and making wrong ones impossible
Every parameter is a chance for the model to be wrong. The design goal is to make the wrong value unrepresentable rather than to validate it afterwards.
- Enums over free text wherever the set is closed. A status parameter that accepts any string will receive any string.
- Flat over nested. Deeply nested objects are constructed incorrectly far more often than flat ones. If you can flatten, flatten.
- Explicit formats in the description, with an example. "An ISO 8601 date, for example 2026-03-14" removes an entire category of error.
- Avoid free-form dates entirely where you can. "Last quarter" means something different depending on when it is read; prefer explicit ranges, and resolve relative expressions in your own code.
- Sensible defaults so optional parameters can genuinely be omitted, rather than the model inventing a plausible value.
- Never expose an internal identifier the model cannot know. If a call needs a database id, the tool should accept what the user actually said and resolve it itself.
Error messages are instructions
Returning the error to the model rather than crashing is right, and the wording of that error determines whether the next attempt succeeds or repeats the mistake.
An error the model can act on says three things: what went wrong, why, and what to do instead.
- Not "Error: invalid input" — but "date must be YYYY-MM-DD; received '14 March'. Convert the date and retry."
- Not "Not found" — but "No order matches that number. Ask the customer to confirm it, or use find_orders_by_email."
- Not "Permission denied" — but "This user cannot view other customers' orders. Do not retry; tell the user this requires support."
That last pattern matters enormously: say explicitly when not to retry. Without it, a model faced with a permanent failure will try again, and again, burning tokens and time on something that cannot succeed. Distinguish retryable errors from terminal ones in the message itself.
And keep internal details out. A stack trace tells the model nothing useful and may leak your schema into a conversation.
What a tool returns is context you are spending
Every token a tool returns sits in the context for the rest of the conversation, is re-read on every subsequent step, and is paid for each time. Tool results are the largest source of context bloat in most agents.
- Return fields, not records. If the agent needs a status and a date, return those two. A full row of forty columns is thirty-eight columns of noise.
- Cap list lengths, and say what you capped. "Showing 10 of 247 matches; narrow the search to see others" is both shorter and more useful than 247 rows.
- Summarise large payloads in your own code before returning them. Deterministic filtering beats hoping the model ignores the irrelevant parts.
- Return stable identifiers the next tool can take, so the model passes a token rather than reconstructing a value.
- Prefer plain text or compact structure over deeply nested JSON. Models read both; one costs three times the tokens.
A useful check: look at the full context after a five-step agent run. If tool output dominates it, that is where your cost and your degradation are coming from, and trimming it is usually the highest-return change available.
Stopping conditions and runaway agents
An agent loop can fail to terminate in ways ordinary code does not, and the defences belong in your loop rather than in the prompt.
- A maximum step count, always, with a clear message when it is hit. Not optional.
- A spend cap per conversation. The failure mode that produces a memorable invoice is a loop nobody bounded.
- Repeat detection. If the same tool is called with the same arguments twice, something is stuck — return a message saying so rather than the same result again.
- A wall-clock timeout for the whole task, separate from individual call timeouts.
- A defined outcome when a limit is hit — hand to a human, return partial results, or fail cleanly. Silent truncation is the worst option and the default in hand-rolled loops.
Designing the approval step so it is not a rubber stamp
Marking a tool as requiring approval is the easy half. The hard half is that an approval gate a person clicks through without reading is worse than no gate, because it creates a record of authorisation for something nobody examined.
What makes it real:
- Show the action, not the agent's explanation of it. The recipient address, the amount, the row count, the destination — rendered plainly. A confident paragraph describing what it is about to do is the thing you are supposed to be checking, not evidence.
- Approve one action, not a plan. "Do these six things" gets a yes; each individually does not. The multi-step approval is where the unintended step hides.
- Make the irreversible ones look different. Sending to a customer and deleting records should not use the same button as fetching a report.
- Show what it will affect, before it happens. "This will email 412 people" is a number that changes decisions; "send the campaign" is not.
- Log the approval with the exact parameters that were approved, and execute exactly those — not a regenerated version that might differ.
The underlying principle is that people become blind to any confirmation they see often. Reserve the gate for actions that genuinely warrant it, and design everything else to be safely reversible so it does not need one.
Testing tools without the model
Agent bugs are hard to reproduce because the model's choices vary. Splitting the testing makes them tractable.
Test the tools as ordinary functions. They are ordinary functions. Unit tests for the happy path, the empty result, the malformed input, the permission failure — none of which needs a model and all of which catch real bugs.
Test selection separately. Take twenty realistic user requests and check which tool the model picks and with what arguments, without executing anything. This isolates description quality from everything else, and it is where most improvement comes from.
Then test end to end on a handful of full scenarios, including ones designed to fail — a tool erroring mid-task, an ambiguous request, an action requiring approval.
Keep all three as a set and re-run them when you change a description, add a tool or change models. Adding a tool changes selection behaviour for the tools already there, which is the regression people do not expect.
Make every action safe to repeat
An agent retries. It retries after a timeout, after an ambiguous error, and sometimes because it lost track of whether the last call succeeded. Ordinary software retries too, but an agent decides to retry based on a judgement rather than a rule, which makes it less predictable and the safeguard more necessary.
So every tool that changes something should be safe to call twice with the same arguments:
- Search before you create. Look for an existing record matching a stable key and update it rather than inserting a second one. Ten minutes of work that prevents duplicate customers, duplicate invoices and duplicate emails.
- Accept an idempotency key where the operation genuinely cannot be checked first — a payment, an outbound message. Same key, same outcome, executed once.
- Return the same result for a repeat rather than erroring. An error on the second call reads to the model as a failure and prompts a third attempt.
- Make reads free of side effects, so an exploring agent cannot change anything by looking.
Then test it directly: call each writing tool twice with identical arguments and confirm the world changed once. This takes minutes and catches the failure that is otherwise discovered by a customer receiving the same email three times.
What happens as the toolset grows
Tool design that works with six tools stops working at forty, and the symptoms are specific: the model picks plausible-but-wrong tools, mixes up similar ones, and spends steps exploring.
What helps, in order:
- Merge near-duplicates. Three search tools that differ slightly are three chances to pick wrong. One with a well-described parameter is usually better.
- Make names systematically distinct.
get_order,get_order_statusandget_orderswill be confused. Name by intent, not by resource shape. - Expose tools by phase. An agent does not need the refund tools while it is still identifying the customer. Presenting a relevant subset per stage sharpens selection and shortens the prompt.
- Sub-agents with their own toolsets, each with a clean context, once a single agent's list becomes unmanageable.
Human-in-the-loop — for dangerous tools
Not every action should run automatically. Tools that perform an irreversible action (sending an email to a customer, making a payment, deleting data) need human approval before running. Design the flow so the agent proposes the action, and a human approves.
- Mark "sensitive" tools that require approval.
- Show the user exactly what's about to happen (which tool, which parameters).
- Only after approval — run it.
Orchestrating multiple tools
When a task requires a sequence of tools (search → filter → act), the model manages the order. For this to work:
- Descriptions that hint at a sequence ("use search_products before add_to_cart").
- Return identifiers the model can pass to the next tool (e.g.
order_id). - Sub-agents for complex tasks — a sub-agent with its own set of tools and a clean context.
- For connecting external tools in a standard way — MCP.
Security — tools are an attack surface
A tool gives the model the ability to act in the real world — and that's dangerous if the model is "influenced" by prompt injection:
- Least privilege. Each tool gets only the access it truly needs.
- Validation on the tool side. Don't trust the parameters the model sent — validate them (e.g. that the user is allowed to access this order_id).
- Sandboxing for tools that run code.
- Human approval for risky actions, as above.
- See AI agent security for a deeper dive.
Next step
Build a full agent, connect tools in a standard way with MCP, and secure it.