An agent that can call tools is far more useful than one that only talks — and far easier to get subtly wrong. The failures are rarely the model "being dumb". They are almost always interface design problems.
What makes tool use unreliable
Four recurring causes:
Ambiguous tool boundaries. search, find_user and lookup all plausibly answer "find the customer". The model picks inconsistently, because the instructions were genuinely ambiguous.
Under-specified parameters. A date argument with no format, no timezone and no example gets "next Tuesday", "2026-03-04", or "04/03/2026" depending on the phase of the moon.
Silent failure. A tool returns {} or null on error. The model interprets emptiness as "no results" and reports confidently that the customer has no orders.
No stopping condition. Nothing tells the agent when it is finished, so it keeps calling tools, or calls the same one repeatedly with slight variations.
Writing tool definitions the model understands
Treat the definition as documentation for a capable new colleague who cannot ask questions.
{
"name": "search_orders",
"description": "Search a customer's orders by date range and status. Use this when the user asks about past purchases, refunds or order history. Requires a customer_id — call lookup_customer first if you only have an email.",
"input_schema": {
"type": "object",
"properties": {
"customer_id": { "type": "string", "description": "Internal id, e.g. cus_8f3c2a. Not an email." },
"from": { "type": "string", "description": "Inclusive start date, ISO 8601 (YYYY-MM-DD), UTC." },
"status": { "type": "string", "enum": ["paid", "pending", "refunded"], "description": "Omit for all statuses." }
},
"required": ["customer_id"]
}
}
What is doing the work here:
When to use it, and when not — the description states the trigger condition
Prerequisites stated explicitly — "call lookup_customer first" prevents a whole class of error
Formats with examples — ISO 8601, UTC, a sample id
Enums instead of free text — the model cannot invent
"complete"Minimal required set — optional parameters with documented defaults
Fewer, well-separated tools beat many overlapping ones. If two tools could plausibly answer the same request, merge them or sharpen the descriptions until they cannot.
Validating arguments
Never pass model output to a system of record unvalidated.
const Schema = z.object({
customer_id: z.string().regex(/^cus_[a-z0-9]+$/),
from: z.string().date().optional(),
status: z.enum(["paid", "pending", "refunded"]).optional(),
});
const parsed = Schema.safeParse(args);
if (!parsed.success) {
return { error: "invalid_arguments", detail: parsed.error.issues.map(i => `${i.path}: ${i.message}`).join("; ") };
}
Return the validation error to the model rather than throwing. Given a specific message, models correct themselves reliably on the next turn. A generic "something went wrong" gives it nothing to act on.
Authorization is separate and non-negotiable: check that this user may access this record, in your code, every time. The model is an untrusted caller. An agent that can read any customer_id it invents is a data breach waiting to happen.
Error handling and retries
Errors are part of the interface. Make them informative and actionable:
{ "error": "customer_not_found", "detail": "No customer with id cus_999. Use lookup_customer with an email to find the correct id." }
Distinguish three classes:
Retryable (rate limit, timeout) — retry with backoff in your code, not by asking the model again
Correctable (bad arguments) — return to the model with specifics
Terminal (not found, forbidden) — return clearly; the model should stop or take a different route
Cap retries per tool and per run. Without a cap, a correctable error the model cannot actually correct becomes an expensive infinite loop.
Orchestration patterns
Tool loop. The default: model calls a tool, receives the result, decides the next step, repeats until it answers. Simple and effective for most work. Always bound the iterations.
Plan then execute. For multi-step tasks, have the model produce a plan first, then execute steps. Easier to review and to show the user, and it prevents wandering.
Restrict tools by phase. Expose only the tools relevant to the current stage. A smaller menu makes selection more reliable.
Human approval for consequential actions. Anything that spends money, emails a customer, or deletes data should require confirmation. Split prepare_refund (safe, returns a preview) from execute_refund (requires an explicit approval token).
Idempotency. Give mutating tools an idempotency key so a retry cannot double-charge.
Observability
You cannot debug what you cannot see. For every run, log the input, each tool call with arguments and result, timing, tokens, and the final output.
Then watch for the signals that matter: tool error rate by tool, average iterations per run (rising means confusion), loops repeating the same call, and outcome rate — did the run actually accomplish the task?
Most agent improvement comes from reading traces of real failures and finding, over and over, that the tool description was ambiguous or the error message was unhelpful. The fix is nearly always in the interface, not the prompt.