Traditional monitoring tells you that a request took 4 seconds and returned HTTP 200. For an LLM application, that is almost useless. The answer might be wrong, the retriever might have found nothing relevant, the agent might have called the same tool eleven times, and the request might have cost forty cents. LLM observability is about seeing what happened inside: prompts, retrieved context, tool calls, tokens, latency per step and, ultimately, quality. This guide shows what to capture, how to structure it with OpenTelemetry, and how to use it day to day.
Why LLM apps need different observability
LLM systems have properties that classic APM does not handle well:
- Non-determinism. The same input can produce different outputs; you need the actual prompt and response to debug.
- Multi-step pipelines. A RAG answer involves query rewriting, embedding, search, reranking and generation. An agent may take dozens of steps. Failures hide in the middle.
- Quality is not binary. A 200 response with a hallucinated answer is a failure that no error rate shows.
- Cost is per request. Token usage varies by orders of magnitude between requests and is a first-class metric.
- Sensitive content. Prompts contain user data, documents and sometimes secrets — logging them needs care.
What to capture
For every LLM call, at minimum:
| Field | Why |
|---|---|
| Model and provider, parameters (temperature, max tokens) | Reproduce and compare |
| Prompt template name and version | Correlate quality with prompt changes; see prompt engineering |
| Input and output tokens, cached tokens | Cost and caching efficiency |
| Latency: time to first token, total | User experience |
| Finish reason (stop, length, tool use, refusal) | Truncation and refusal tracking |
| Tool calls with arguments and results | Agent debugging |
| Retrieved document IDs and scores | RAG debugging |
| User/session/tenant IDs (pseudonymised) | Grouping, per-tenant cost |
| Feedback and eval scores attached later | Quality over time |
Store the full prompt and completion where policy allows, with retention limits and access control. Without them, most debugging is guesswork.
Structure it as traces
A trace for one user request should mirror the pipeline. For a RAG chat:
trace: POST /api/chat (2.9 s, $0.011)
├─ span: rewrite_query llm claude-… in 420 / out 35 tok 310 ms
├─ span: embed_query embedding 45 ms
├─ span: hybrid_search db top_k=50 80 ms
├─ span: rerank reranker 50 → 8 190 ms
└─ span: generate_answer llm in 6,200 (cached 4,800) / out 310 tok 2.2 s
attributes: prompt=support_answer@v14, finish_reason=stop
For an agent, each loop iteration becomes a span with the model's decision and child spans for tool executions. For multi-agent systems, each subagent is a child trace of the orchestrator's span. This structure lets you answer questions such as "which step got slower after last week's release?" or "how often does the agent call search_orders twice in a row?".
OpenTelemetry GenAI semantic conventions
OpenTelemetry has semantic conventions for generative AI that standardise span names and attributes for model calls, agents and tools: attributes like gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons and operation names for chat, embeddings and tool execution. Using them means your LLM telemetry flows through the same Collector and backends as the rest of your OpenTelemetry setup, and tools that understand the conventions can render it properly.
A manual span in Python:
from opentelemetry import trace
tracer = trace.get_tracer("support-bot")
def generate_answer(messages, prompt_version: str):
with tracer.start_as_current_span("chat claude") as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.request.model", MODEL)
span.set_attribute("app.prompt.version", prompt_version)
resp = client.messages.create(model=MODEL, max_tokens=800, messages=messages)
span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)
span.set_attribute("gen_ai.response.finish_reasons", [resp.stop_reason])
return resp
In practice, auto-instrumentation libraries (for example OpenLLMetry and OpenInference) patch common SDKs and frameworks and emit these spans for you. Some conventions are still marked experimental; pin library versions and expect attribute names to evolve.
Tools
| Tool | Type | Notes |
|---|---|---|
| Langfuse | Open source, self-hostable + cloud | Traces, prompt management, evals, cost tracking; OpenTelemetry ingestion |
| Arize Phoenix | Open source | Tracing via OpenInference, evals, dataset tools |
| LangSmith | SaaS (self-host on enterprise) | Deep LangChain/LangGraph integration |
| Your existing APM (Grafana, Datadog, Honeycomb…) | OTel backend | Good for latency/cost; weaker for prompt-level review |
A common setup: emit OpenTelemetry once, send infrastructure metrics and spans to your existing backend, and send GenAI spans (with content) to an LLM-specific tool for prompt-level analysis. Self-hosting matters when prompts contain personal or confidential data; see privacy for LLM apps.
Metrics and dashboards
Build one dashboard per LLM feature with:
- Volume: requests, conversations, agent runs.
- Latency: p50/p95 time to first token and total; per-step latency for pipelines.
- Cost: tokens and money per request, per user, per tenant, per feature; cache hit rate. Use these to drive LLM cost optimization.
- Reliability: provider errors, rate limits (429), timeouts, retries, fallbacks; see LLM API reliability.
- Behaviour: finish reasons (truncation rate), refusal rate, tool error rate, steps per agent run.
- Quality: user feedback rate and score, online eval scores, escalation-to-human rate.
Alert on what users feel or what costs money: p95 latency, error rate, sudden cost spikes per tenant, and agent runs exceeding step limits.
Connecting observability to quality
Traces become far more valuable when connected to evaluation:
- Online evals. Run cheap automated checks on a sample of production traces — groundedness, format validity, policy violations — and attach scores to traces.
- Feedback. Attach thumbs up/down and comments to the trace that produced the answer.
- Datasets from production. Filter traces with low scores or negative feedback, review them, and add representative cases to your golden dataset.
- Release comparison. Tag traces with app and prompt versions; compare scores and costs before and after each release.
This loop — production traces feeding offline evals — is the core of the process described in how to evaluate LLM applications.
Debugging with traces: three typical cases
"The bot answered confidently but wrong." Open the trace and check the retrieval span first. In most cases the right document was not among the retrieved chunks, or it was ranked below the cut-off. The fix belongs in chunking, hybrid search or reranking, not in the prompt; see embeddings and chunking.
"Answers became slower this week." Compare per-span latency before and after the release. A typical culprit is a prompt change that moved dynamic content to the beginning of the prompt, breaking the cache prefix: cached tokens drop to zero, input cost and time to first token jump.
"The agent sometimes loops." Filter agent traces by step count above the 95th percentile and read them. Repeated identical tool calls usually point to an unhelpful error message or a tool output that does not tell the model the task is complete; see designing tools for agents.
Privacy and security of telemetry
Prompts and completions in your observability stack are a data store with personal data. Treat them that way:
- Redact or pseudonymise PII before export where possible (emails, phone numbers, tax IDs); tools like Microsoft Presidio help.
- Separate content from metadata. Keep token counts and latency for long periods; keep content for a short retention window.
- Restrict access to content to the people who need it for debugging.
- Never log secrets. Scrub API keys and tokens from tool arguments.
- Mind data residency when using SaaS tools; EU data may need EU hosting.
A rollout plan
- Day 1: log model, tokens, latency, finish reason and prompt version for every call. Even a structured log line is a big step.
- Week 1: add OpenTelemetry tracing with GenAI conventions across the pipeline; export to an LLM observability tool.
- Week 2–3: dashboards for cost, latency, errors and behaviour; alerts on spikes.
- Month 2: feedback capture, online evals on a sample, weekly review of low-scoring traces.
FAQ
Should we log full prompts in production? Usually yes, with redaction, short retention and access control. Without content, debugging quality problems is nearly impossible.
Does tracing add latency? Negligible if exports are asynchronous and batched, as OpenTelemetry does by default.
How do we trace streaming responses? Start the span at request, record time to first token as an event or attribute, and end it when the stream completes, including token usage from the final chunk.
Can we use our existing Grafana or Datadog? Yes for metrics and latency. For reading prompts, comparing versions and labelling traces, LLM-specific tools are more productive.
Sources
- OpenTelemetry. Semantic conventions for generative AI systems.
- Langfuse documentation.
- Arize Phoenix.
- Traceloop. OpenLLMetry.
- Microsoft. Presidio — data protection and de-identification SDK.
- Hamel Husain. Your AI product needs evals.