LLM APIs are remote services with unusual characteristics: requests take seconds to minutes, capacity is shared across thousands of customers, rate limits are measured in tokens as well as requests, and outages of popular providers make headlines. A feature that works perfectly in a demo can fail daily in production. The good news: the reliability patterns are well known from distributed systems. This guide adapts them to LLM integrations.
Know your failure modes
| Failure | Typical signal | Retry? |
|---|---|---|
| Rate limit exceeded | HTTP 429, often with retry-after header |
Yes, after the indicated delay |
| Provider overloaded | HTTP 529 / 503 or provider-specific "overloaded" error | Yes, with backoff |
| Server error | HTTP 500, 502, 504 | Yes, with backoff, limited attempts |
| Timeout / connection reset | Client-side timeout, broken stream | Yes, if the operation is safe to repeat |
| Invalid request | HTTP 400 (bad parameters, context too long) | No — fix the request |
| Authentication / permission | HTTP 401, 403 | No — alert |
| Content refusal or policy block | Successful response with refusal or stop reason | No — handle in application logic |
| Truncated output | Finish reason max_tokens / length |
Maybe — continue or increase limit |
| Invalid structured output | Parse or validation error | Yes, once, with the error message |
Provider documentation lists the exact codes and headers (Anthropic errors, Anthropic rate limits, OpenAI rate limits). Classify errors explicitly in your client; treating everything as "retry three times" either hammers a broken request or gives up on a transient one.
Timeouts: set them deliberately
Default HTTP client timeouts are rarely right for LLMs:
- Connect timeout: short (5–10 s). Failing to connect is not an LLM problem.
- Time to first token: for streaming, a separate deadline (for example 20–30 s) catches stuck requests early.
- Idle timeout between chunks: a stream that stops sending data for 30–60 s is likely dead.
- Total deadline: derived from the use case — a chat answer might get 60–90 s, a background report several minutes.
Propagate deadlines: if your HTTP endpoint has a 30-second budget, the model call cannot have 120 seconds. And always cancel the upstream request when the client disconnects — otherwise you pay for tokens nobody reads; see building an AI chat in Next.js.
Retries with exponential backoff and jitter
Retry transient errors with exponentially growing delays and randomness, so that thousands of clients do not retry in lockstep. The "full jitter" approach — a random delay between zero and the exponential cap — performs well under contention (AWS Architecture Blog).
import random, time
RETRYABLE = {429, 500, 502, 503, 504, 529}
def call_with_retries(fn, max_attempts=4, base=1.0, cap=30.0):
for attempt in range(1, max_attempts + 1):
try:
return fn()
except ApiError as e:
if e.status not in RETRYABLE or attempt == max_attempts:
raise
retry_after = e.headers.get("retry-after")
delay = float(retry_after) if retry_after else random.uniform(0, min(cap, base * 2 ** attempt))
time.sleep(delay)
Rules that prevent retries from making things worse:
- Honour
retry-afterwhen the provider sends it. - Cap attempts (3–5) and total retry time within the request deadline.
- Retry at one layer only. If the SDK retries, your code wrapping it, and the queue around that, one failure becomes dozens of calls. Most official SDKs retry automatically; configure them rather than stacking your own.
- Use a retry budget — for example, retries may not exceed 10% of requests over a minute — so a provider outage does not multiply your traffic. The Google SRE book covers this in depth.
Rate limits: manage them proactively
Providers limit requests per minute, input tokens per minute and output tokens per minute, per model and per organisation tier. Hitting them is normal at scale. Better than reacting to 429s:
- Client-side rate limiting with a token bucket per model, sized slightly below your limits.
- Priority classes: interactive requests first; batch and background work fill the remaining capacity.
- Queues for background work with concurrency limits; see the section below.
- Batch APIs for anything that can wait hours — much cheaper and with separate limits; see LLM cost optimization.
- Prompt caching reduces token pressure for repeated prefixes; some providers count cached tokens differently against limits.
- Monitor headroom: response headers often report remaining requests and tokens; export them as metrics.
Fallbacks: models and providers
When the primary model is unavailable or rate-limited, a fallback keeps the feature alive:
- Same provider, different model — often a smaller model with separate capacity.
- Same model via another platform — many models are available through several clouds (for example, the provider's own API and major cloud marketplaces), which gives provider diversity with identical behaviour.
- Different provider's model — maximum independence, but different behaviour; prompts and output parsing must be tested with each fallback.
const chain = [
{ provider: "primary", model: process.env.CHAT_MODEL! },
{ provider: "cloud-b", model: process.env.CHAT_MODEL_CLOUD_B! },
{ provider: "fallback", model: process.env.FALLBACK_MODEL! },
]
export async function generate(req: Req) {
let lastError: unknown
for (const target of chain) {
if (breaker.isOpen(target.provider)) continue
try {
const res = await callModel(target, req, { timeoutMs: 45_000 })
breaker.success(target.provider)
return { ...res, servedBy: target }
} catch (e) {
lastError = e
if (!isRetryableAcrossProviders(e)) throw e // e.g. 400 invalid request
breaker.failure(target.provider)
}
}
throw lastError
}
Log which target served each request (servedBy) and include it in traces. Run your eval suite against every fallback model so you know what quality users get during an incident; see LLM evals. Gateways such as LiteLLM-style proxies or cloud AI gateways can centralise this logic for all services.
Circuit breakers
A circuit breaker stops sending requests to a dependency that is failing, fails fast instead, and periodically probes for recovery. For LLM providers, open the circuit when error rates or latencies exceed a threshold over a short window, and route traffic to the fallback immediately rather than waiting for each request to time out. Resilience4j for JVM services, Polly for .NET and libraries like opossum for Node implement the pattern.
Idempotency for agents and side effects
Retries are dangerous when the operation has side effects. An agent step that creates an invoice and then times out may be retried and create a second invoice. Protect side effects:
- Idempotency keys on write operations, derived from the run ID and step number. APIs such as Stripe's show the pattern.
- Check before act: tools that create records should first search for an existing record from the same run.
- Checkpoint agent state after each completed step so a resumed run does not repeat finished work.
- Separate decision from execution: the model proposes an action; deterministic code executes it exactly once.
These matter even more for long-running and multi-agent systems.
Queues for background and bulk work
For non-interactive LLM work — document processing, enrichment, nightly summaries — put requests on a queue:
- Workers pull jobs with a concurrency limit that respects rate limits.
- Failed jobs retry with backoff and go to a dead-letter queue after N attempts.
- Jobs are idempotent and carry all inputs needed for a retry.
- Throughput adapts automatically: during provider slowdowns, the queue grows instead of requests failing.
This architecture also makes cost predictable and pairs well with batch APIs. Our extraction pipelines use it; see LLM document processing.
Graceful degradation
Decide in advance what users see when AI is unavailable:
- Chat assistant: a clear message with a link to human support or search, not a spinner forever.
- AI-powered search: fall back to keyword search.
- Classification and routing: fall back to rules or a default queue for human triage.
- Content generation: let users continue without it; queue the generation for later.
Feature flags allow you to turn AI features off quickly during an incident. The best AI features are additive: the product still works without them.
Observability for reliability
Measure what you need to manage the above:
- Error rate by type (429, 5xx, timeouts) and by provider and model.
- Retry count and fallback rate; time spent in retries.
- Latency percentiles including time to first token.
- Rate-limit headroom.
- Circuit breaker state changes.
Alert on fallback activation and sustained error rates. The tracing setup is in LLM observability.
A reliability checklist
- Errors are classified; only transient ones are retried
- Timeouts: connect, first token, idle and total deadline
- Exponential backoff with jitter, honouring
retry-after, retries at one layer only - Client-side rate limiting and priority between interactive and background work
- At least one tested fallback model or provider for critical features
- Circuit breaker per provider
- Idempotency keys for every side effect triggered by AI
- Queue with dead-letter handling for background jobs
- Defined degraded experience and feature flag to disable AI features
- Dashboards and alerts for errors, fallbacks and latency
FAQ
Should we always have a second provider? For customer-facing critical features, yes, or at least the same model through a second platform. For internal tools, a clear degraded mode may be enough.
Do official SDKs handle retries? Most retry a few times on connection errors, 429 and 5xx by default. Configure their limits and do not add another retry layer on top.
How do we test reliability? Inject faults in staging: mock 429s, slow streams and outages, and confirm that fallbacks, breakers and degraded UX work.
Is streaming harder to make reliable? Yes. A stream can fail mid-way; decide whether to show the partial answer, resume or regenerate, and never retry a half-streamed response silently.
Sources
- AWS Architecture Blog. Exponential Backoff And Jitter.
- Google SRE Book. Addressing Cascading Failures and Handling Overload.
- Martin Fowler. CircuitBreaker.
- Anthropic. API errors and rate limits.
- OpenAI. Rate limits guide.
- Stripe. Idempotent requests.