The first invoice from an LLM provider is usually a pleasant surprise. The invoice three months after launch, when traffic has grown and the agent makes six tool calls per question, often is not. The good news: most AI applications we audit can cut their model spend by 50–80% without any loss in quality, mostly through a handful of mechanical changes. This guide covers the levers in the order we apply them — free wins first, quality trade-offs last — with real price math.
Where LLM costs actually come from
You pay per token, separately for input and output. For an AI agent, input usually dominates: every request resends the system prompt, tool definitions, retrieved documents and the conversation so far. Output tokens cost more per token, but there are fewer of them.
Current list prices for Anthropic's models illustrate the spread (Anthropic pricing):
| Model | Input, $ per million tokens | Output, $ per million tokens | Typical use |
|---|---|---|---|
| Claude Fable 5.1 | 10.00 | 50.00 | Hardest reasoning and long-horizon agent work |
| Claude Opus 5.5 | 4.00 | 20.00 | Complex agents, coding, analysis |
| Claude Sonnet 5.5 | 2.00 | 10.00 | Everyday agents, support, enterprise workloads |
| Claude Haiku 4.5 | 1.00 | 5.00 | Classification, extraction, high-volume routing |
Prices as of September 2026; check the pricing page before budgeting.
Before optimizing, measure. Log input_tokens, output_tokens and cache fields from every response, grouped by feature and route. You cannot cut what you cannot see, and the expensive route is rarely the one the team suspects.
Lever 1: Prompt caching (the biggest free win)
Most requests in an AI application share a long, identical prefix: the system prompt, tool definitions, policy documents, few-shot examples. Prompt caching lets the provider store that prefix and charge a fraction of the price when it is reused.
With Anthropic's API, a cache write costs 1.25× the base input price for the default 5-minute lifetime (2× for the optional 1-hour lifetime), and a cache read costs about 0.1× — less on some newer models (Anthropic: prompt caching). With a 5-minute cache, the second request already breaks even.
Example. A support assistant sends a 20,000-token prefix (instructions, tools, policy excerpts) with each of 100,000 monthly requests on Claude Sonnet 5.5:
- Without caching: 2 billion input tokens × $2 = $4,000 for the prefix alone.
- With caching and a high hit rate: about 2 billion tokens read at $0.20 per million = $400, plus a small amount of cache writes.
To make caching work:
- Put stable content first, volatile content last. Tools, system prompt and documents go before the user's question. A timestamp or user ID at the top of the system prompt silently breaks every cache hit.
- Keep the order deterministic. Serialize tool lists and JSON in a fixed order.
- Verify. If
cache_read_input_tokensstays at zero across repeated requests, something in your prefix changes on every call.
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2000,
system=[{
"type": "text",
"text": STABLE_INSTRUCTIONS_AND_POLICIES, # identical on every request
"cache_control": {"type": "ephemeral"}, # cache everything up to here
}],
messages=history + [{"role": "user", "content": question}],
)
print(response.usage.cache_read_input_tokens, response.usage.input_tokens)
Caching pairs naturally with retrieval: keep the stable knowledge in the cached prefix and send only the retrieved passages per request, as described in RAG for business.
Lever 2: Batch processing for anything that can wait
Not every request needs an answer in two seconds. Nightly report generation, document classification, evaluation runs, data enrichment and translation of catalogs can wait minutes or hours. Batch APIs process such requests asynchronously at a large discount — 50% on Anthropic's Message Batches API (Anthropic: batch processing).
Rule of thumb: if a human is not waiting for the result, it belongs in a batch. Evals are a perfect candidate; see How to evaluate LLM applications.
Lever 3: Input token hygiene
- Trim tool definitions. Twenty tools with long descriptions can cost thousands of tokens per request. Expose only the tools relevant to the route, and keep descriptions precise rather than long.
- Return compact tool results. An API call that returns a 200-field JSON record when the model needs five fields multiplies the cost of every following turn, because the result stays in the conversation.
- Retrieve less, rank better. Sending 40 document chunks "just in case" is expensive and often lowers quality. Five to eight well-ranked chunks usually beat 40.
- Manage long conversations. Summarize or clear old tool results instead of resending a 100-turn history forever.
Lever 4: Output token hygiene
Output tokens cost five times as much as input on most models. Ask for what you need:
- Set response length expectations in the prompt ("answer in at most three sentences").
- Use structured outputs for machine-consumed responses instead of prose that your code then parses.
- Don't ask the model to repeat the question, restate context or add generic disclaimers.
Lever 5: Tune effort and reasoning
Modern models think before answering, and thinking tokens are billed as output. Many models expose an effort setting that controls how much they reason. Hard tasks — multi-step agents, coding, analysis — benefit from high effort; classification, extraction and simple chat often perform just as well at low effort for a fraction of the cost.
Tune effort per route, measure quality on your eval set, and judge cost per completed task, not cost per request. A cheaper setting that needs two extra turns to finish the job isn't cheaper.
Lever 6: Choose and route models deliberately
The most expensive model is not always the best value, and the cheapest is not always the cheapest per outcome.
- Match the model to the route. A high-volume classifier can run on a small model while the complex agent uses a stronger one.
- Try the strong model at lower effort before building a cascade. Newer models at low effort often match older ones at high effort, and one model keeps one cache.
- Use a router only when the data supports it. A small model that classifies requests and sends easy ones to a cheaper model can save money, but each extra model adds latency, a separate cache and another thing to evaluate.
Every model change needs an eval run first. A 30% cheaper model that resolves 10% fewer support cases may cost more once human escalations are counted, as we discuss in AI support agents: where they work and where they don't.
Lever 7: Guardrails against runaway spend
- Per-user and per-session limits on tokens and tool calls.
- A maximum number of agent iterations per task.
- Alerts when daily spend deviates from the trend.
- Separate API keys or workspaces per product, so one feature can't eat another's budget.
The order we apply these levers
| Step | Lever | Typical saving | Quality risk |
|---|---|---|---|
| 1 | Measure tokens by route | — | None |
| 2 | Prompt caching | 30–80% of input cost | None |
| 3 | Batch for non-interactive work | 50% on those requests | None |
| 4 | Input and output hygiene | 10–40% | Low |
| 5 | Effort tuning per route | 10–50% | Measure |
| 6 | Model selection and routing | 20–70% | Measure carefully |
A monthly cost review routine
Cost control works best as a habit rather than a one-off project:
- Pull usage by route and model for the last month and compare with the previous one.
- Check cache hit rates on every route with a long shared prefix; a sudden drop usually means someone added a timestamp or reordered tools.
- Look at the top 1% most expensive requests. Runaway agent loops and oversized documents hide there.
- Review cost per completed task, not just total spend, together with quality metrics from your evals.
- Re-test model and effort choices when providers release new models; prices and capabilities change several times a year.
FAQ
Does caching change the model's answers? No. Cached and uncached requests produce equivalent results; only the billing and latency differ.
How do I estimate costs before launch? Count tokens for realistic requests with the provider's token-counting endpoint, multiply by expected volume, and add 30–50% for retries, longer conversations and growth.
Is self-hosting an open-source model cheaper? Sometimes, at very high and steady volume. Include GPU costs, engineering time, and the quality gap on your own eval set before deciding.
What is a realistic cost per support conversation? It ranges from fractions of a cent for simple FAQ answers to tens of cents for multi-step agents with tools. Measure your own, by route.
Should we cache tool definitions too? Yes. Tools are part of the prompt prefix, so a stable, ordered tool list placed before the cache breakpoint is cached together with the system prompt.
Sources
- Anthropic. Pricing.
- Anthropic. Prompt caching.
- Anthropic. Batch processing.
- Anthropic. Create strong empirical evaluations.
- OWASP GenAI Security Project. LLM10: Unbounded Consumption.