An agent is only as capable as the tools it can call. Teams spend weeks tuning the system prompt and then expose their backend to the model as thirty thin wrappers around REST endpoints — and wonder why the agent picks the wrong one, passes malformed IDs and drowns in 50-kilobyte JSON responses. Tools for agents are a new kind of API whose consumer is a language model. They need different design choices than APIs for programmers. This guide collects those choices, based on vendor guidance, research and what we have seen work.

Why agent tools are different from APIs

A developer reading your API documentation can spend an hour understanding it, write code once, and call the API millions of times deterministically. A model reads the tool definitions on every request, must decide in the moment which tool to call and with what arguments, and pays for every token of input and output in its limited context window.

Research on SWE-agent showed that the design of the agent-computer interface — what commands exist, how output is formatted, how errors are reported — changes agent performance as much as the choice of model (Yang et al., 2024). Anthropic's engineering guide on writing tools for agents reaches the same conclusion from production experience: tool design is one of the highest-leverage things you can improve.

Principle 1: Design for tasks, not endpoints

A one-to-one mapping from REST endpoints to tools is the most common mistake. If answering "when is my order arriving?" requires get_customer, list_orders, get_order, get_shipment and get_carrier_status, the agent must chain five calls correctly, and each one costs latency and context.

Instead, design tools around what the agent needs to accomplish:

Endpoint-shaped tools Task-shaped tool
list_orders, get_order, get_shipment, get_tracking get_order_status(order_ref) returning status, items, carrier and ETA
list_users, list_events, create_event schedule_meeting(attendees, duration, window) that finds free slots and books
read_logs(service, from, to) returning everything search_logs(service, query, time_range, max_results)

Task-shaped tools consolidate common multi-step operations, reduce the number of decisions the model must make and keep the context clean. Keep a small number of lower-level tools for the unusual cases.

Principle 2: Names and descriptions are prompts

The model chooses tools based on their names and descriptions. Write them as you would explain the tool to a new colleague:

{
  "name": "search_customer_tickets",
  "description": "Search support tickets for ONE customer. Use this when the user asks about their previous requests or a problem they reported before. Returns at most 10 tickets, newest first, each with id, subject, status and a 200-character summary. To read the full conversation of a ticket, call get_ticket_thread with its id. Does not search other customers' tickets.",
  "input_schema": {
    "type": "object",
    "properties": {
      "customer_id": {"type": "string", "description": "Internal customer id like CUS-48213. Get it from the session context, never ask the user for it."},
      "query": {"type": "string", "description": "Keywords in Ukrainian or English. Leave empty to list recent tickets."},
      "status": {"type": "string", "enum": ["open", "closed", "any"], "default": "any"}
    },
    "required": ["customer_id"]
  }
}

Good descriptions say when to use the tool, what it returns, what it does not do, and how it relates to other tools. Parameter descriptions specify formats and where values come from. Namespacing related tools with a common prefix (crm_search_contacts, crm_get_deal) helps when an agent has many tools from different systems.

Principle 3: Return what the agent needs, not everything you have

Tool output goes straight into the model's context. A raw database row with 60 columns, internal UUIDs and nested metadata wastes tokens and distracts the model.

  • Return high-signal fields. Human-readable names over opaque IDs where possible; include IDs only when the agent needs them for follow-up calls.
  • Offer a verbosity option (response_format: "concise" | "detailed") so the agent can ask for more when needed.
  • Paginate and truncate with clear signals: "Showing 10 of 248 results. Narrow the query or request page 2."
  • Format for reading. Short structured text or compact JSON both work; consistency matters more than the format.
  • Cap sizes in code. A tool that can occasionally return 2 MB will eventually break a session.

Context is the scarcest resource an agent has, as explained in context engineering.

Principle 4: Make errors actionable

When a tool fails, the error message is the model's only clue for recovery. Compare:

  • Error 400 — the agent will retry the same call or give up.
  • Invalid date "15.09.2026": use ISO format YYYY-MM-DD, e.g. 2026-09-15. — the agent fixes the argument and succeeds.
  • Customer CUS-482 not found. Customer ids have 5 digits after "CUS-"; use search_customers by name or email to find the right id. — the agent knows exactly what to do next.

Return errors as tool results rather than raising exceptions that end the loop. Distinguish errors the model can fix (bad arguments) from errors it cannot (service down), and say so: "Billing service is unavailable. Tell the user to try again later; do not retry."

Principle 5: Make arguments hard to get wrong

  • Use enums for closed sets.
  • Accept natural formats when it is cheap to normalise them in code (dates, phone numbers, amounts with or without currency).
  • Avoid parameters the model must compute, such as offsets, Unix timestamps or internal codes. Let the tool translate "last 7 days" or a human-readable name.
  • Validate strictly and explain failures, as above.
  • Use strict schema modes where the provider supports them, so arguments always match the schema; see structured outputs.

Principle 6: Permissions live in the tool, not in the prompt

The model will eventually call a tool in a way you did not intend — because of a misunderstanding or because of prompt injection in content it read. Safety must be enforced in tool code:

  • Execute every call with the end user's permissions, checked by the backend.
  • Separate read tools from write tools; give write tools narrow scopes and require confirmation for consequential actions.
  • Never accept identity from the model. Take customer_id from the authenticated session, not from tool arguments, when it determines what data is visible.
  • Rate-limit and log every call with user, arguments and result.

The full security model is described in AI agent security and prompt injection.

Principle 7: Keep the toolset small and distinct

Every additional tool adds tokens to every request and another option the model can confuse. Overlapping tools — search_docs and find_documents and query_kb — are a reliable source of wrong choices. Aim for a small set of distinct tools per agent; if you need dozens, split responsibilities across specialised agents or load tool definitions dynamically based on the task. Some platforms support tool search, where the model discovers rarely used tools on demand instead of seeing all definitions upfront.

Exposing tools via MCP

If several agents, IDEs or chat clients need the same capabilities, implement tools once as a Model Context Protocol server. MCP standardises tool discovery and invocation, so the same CRM or ERP tools work in Claude, IDE assistants and your own agents. All the design principles above still apply — MCP is a transport, not a design. Our business-oriented introduction is in MCP for business, and an ERP example is in Odoo AI integration.

Testing tools with evals

Tool quality is measurable. Build an evaluation set of realistic tasks that require the tools, run the agent, and measure:

Metric What it reveals
Task success rate Overall usefulness
Correct tool selection Confusing names or overlapping tools
Argument error rate Unclear schemas or formats
Calls per task Missing task-shaped tools, poor outputs
Tokens per task Bloated responses
Recovery after errors Quality of error messages

Then iterate: read transcripts where the agent failed, change descriptions, outputs or tool boundaries, and re-run. Anthropic's guide recommends letting a model analyse failed transcripts and propose tool improvements — a surprisingly effective loop. The general eval methodology is in LLM evals.

A checklist for each tool

  • Designed around a task the agent actually performs
  • Name is specific; description says when to use, what it returns and what it does not do
  • Every parameter has a description with format and source
  • Output contains only high-signal fields, with size limits and pagination
  • Errors explain what went wrong and what to do next
  • Permissions enforced in code, identity taken from the session
  • Write actions require confirmation where consequences matter
  • Covered by eval tasks; tool-selection and argument errors are tracked

FAQ

How many tools can an agent handle? Modern models handle dozens of well-designed tools, but accuracy and cost both degrade as the set grows. Start small and add tools when evals show a need.

Should tools return JSON or text? Either works. Choose based on what makes the content easiest to read for the model, and stay consistent within a toolset.

Can we auto-generate tools from our OpenAPI spec? As a starting point, yes. Then consolidate endpoints into task-shaped tools, rewrite descriptions and trim responses — the generated version is rarely good enough.

How do we version tools? Like APIs: additive changes are safe, breaking changes need a new name or version, and evals should run on every change.

Sources

  1. Anthropic (2025). Writing effective tools for agents — with agents.
  2. Yang et al. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.
  3. Anthropic. Tool use with Claude.
  4. OpenAI. Function calling guide.
  5. Model Context Protocol.
  6. Patil et al. (2023). Gorilla: Large Language Model Connected with Massive APIs.