The most common question we hear from teams with an AI feature in production is: "We changed the prompt — is it better or worse now?" Without evaluations, nobody knows. Someone tries five questions, the answers look fine, the change ships, and a week later support tickets reveal a regression in a case nobody tested. This guide shows how to build LLM evals that catch those regressions before users do: what to measure, where test data comes from, when to trust an LLM as a judge, and how to run it all on every change.
Why LLM apps need their own kind of testing
Traditional tests assert exact outputs: add(2, 2) == 4. LLM outputs are non-deterministic, open-ended and often have many correct answers. That does not make testing impossible; it changes its shape.
An eval is a repeatable experiment: a fixed set of inputs, the system under test, and a grading method that turns outputs into scores. Run the same eval before and after a change, and you can compare versions with numbers instead of impressions. Anthropic's documentation puts it simply: define success criteria first, then build tests that measure them (Anthropic: define success criteria; develop tests).
Evals matter for three decisions you will make repeatedly:
- Prompt and pipeline changes — did the new chunking, retriever or prompt help?
- Model changes — can we move to a newer or cheaper model without losing quality? See LLM cost optimization.
- Release gates — is this version good enough to ship?
Step 1: Define what "good" means
Vague goals produce useless evals. "The assistant should be helpful" cannot be measured; these can:
| Quality dimension | Example criterion | How to grade |
|---|---|---|
| Correctness | Answer matches the reference for factual questions | Exact or fuzzy match, LLM judge |
| Groundedness | Every claim is supported by the retrieved sources | LLM judge with sources |
| Refusal | Says "I don't know" when the answer is not in the docs | Code check on a labeled subset |
| Format | Returns valid JSON matching the schema | Schema validation |
| Safety | Never reveals another customer's data | Code check, red-team set |
| Tone and length | Under 120 words, no marketing language | Code check, LLM judge |
| Task success (agents) | The ticket was created with correct fields | Check the resulting system state |
Pick three to five dimensions that matter most for your product. Each one needs a clear pass/fail definition that two people would apply the same way.
Step 2: Build a golden dataset from reality
The golden dataset is the heart of every eval. Its quality matters more than its size.
- Start from real traffic. Pull 100–200 questions from support tickets, chat logs, search queries and sales emails. Synthetic questions written by the team are cleaner than reality and miss the typos, ambiguity and mixed languages users actually produce.
- Cover the distribution, then the edges. Most items should look like everyday traffic; add deliberate edge cases: questions with no answer in the docs, ambiguous requests, adversarial inputs and multi-step tasks.
- Store references, not just inputs. For each item, record the expected answer or the key facts it must contain, the source document if relevant, and tags such as
billing,refusalormulti-turn. - Version it. Keep the dataset in Git next to the code. When a production failure appears, add it to the set so it can never silently come back.
{"id": "billing-017",
"input": "Can I get a refund if I cancel my annual plan after 3 months?",
"must_include": ["pro-rated refund", "within 30 days"],
"source": "policies/refunds.md#annual-plans",
"tags": ["billing", "policy"]}
Step 3: Choose grading methods, cheapest first
Use the simplest grader that measures the criterion reliably.
Code-based checks are fast, free and deterministic: exact match, regular expressions, JSON schema validation, "contains the required fact", length limits, or verifying that an agent's API call produced the right record. Use them wherever possible.
LLM-as-a-judge handles what code cannot: groundedness, completeness, tone. Strong models agree with human preferences most of the time on many tasks (Zheng et al., 2023), but judges have known biases. They tend to prefer the first of two answers and longer answers, so randomize order and control for length (Wang et al., 2023).
Rules for a judge you can trust:
- Grade one criterion per call with a clear rubric and a binary or 1–3 scale. "Rate quality 1–10" produces noise.
- Ask for a short reason before the verdict, and log both.
- Calibrate against humans. Have a person label 50–100 outputs and measure agreement. Research on evaluator alignment shows that criteria drift as people see more outputs, so iterate the rubric with real examples (Shankar et al., 2024).
- Use a capable model for judging and keep it fixed between runs, or you are comparing two moving targets.
JUDGE_PROMPT = """You grade answers from a customer support assistant.
Criterion: GROUNDEDNESS. Every factual claim in the answer must be
supported by the sources. Ignore style.
<sources>{sources}</sources>
<answer>{answer}</answer>
First write one sentence of reasoning, then output exactly
PASS or FAIL on the last line."""
Human review remains the ground truth. Budget a few hours a week for someone who knows the domain to read a sample of production conversations and low-scoring eval items. It is the fastest way to discover failure modes your metrics don't capture yet.
For retrieval-augmented systems, measure retrieval and generation separately — recall of the right document versus faithfulness of the answer — as described in RAG for business. Frameworks such as RAGAS automate parts of this (Es et al., 2023).
Step 4: Run evals on every change
An eval that someone runs "when they remember" is not a safety net. Wire it into the development loop:
- Pull requests: a fast subset (30–50 items, code checks plus a few judge calls) runs in CI and posts scores to the PR.
- Nightly or pre-release: the full set with all judges, compared against the last release.
- Thresholds, not perfection: fail the build if a key metric drops by more than an agreed margin, for example groundedness by more than 3 points.
- Run each item more than once when outputs vary a lot; average across two or three samples to separate real changes from noise.
Open-source tools like promptfoo let you describe test cases and assertions in config files and run them from CI; a small custom script is often enough as well. The tool matters less than the habit. Our CI setup for small teams is in CI/CD without pain.
Step 5: Close the loop with production
Offline evals predict quality; production data confirms it. Track:
- Explicit feedback — thumbs up/down, but expect low response rates.
- Implicit signals — rephrased questions, escalations to a human, abandoned sessions, copy-paste of answers.
- Business outcomes — resolution rate, time to resolution, conversion.
Every week, sample failures from production, label them, and add representative ones to the golden set. Over a few months the dataset becomes the most valuable asset in the project: it encodes what "good" means for your users. In customer support, these are the same metrics we use to decide where AI agents work and where they don't.
Common mistakes
- Testing only happy paths. Most production incidents come from refusals, ambiguity and adversarial inputs.
- One aggregate score. A single "quality: 82%" hides that billing questions dropped from 90% to 60%. Report by tag.
- Judging with the model under test. Use a fixed, capable judge model and calibrate it.
- Never refreshing the dataset. Products, docs and users change; the eval must change with them.
- No cost or latency tracking. A prompt that is 2% better but 3x slower may be a regression for users.
FAQ
How many test cases do we need? Start with 50–100 real ones and grow to a few hundred. Coverage of real scenarios beats size.
Can we rely only on LLM-as-a-judge? Not without calibration. Measure agreement with human labels first, and keep code checks for everything they can verify.
How do we evaluate agents? Grade the final state — was the right record created, was the right email drafted — and the trajectory: number of steps, tool errors and unnecessary actions.
How much does running evals cost? Usually a small fraction of production spend. Batch processing and caching cut judge costs further.
Who should own the eval set? A product owner or domain expert, together with engineering. Engineers keep it running; domain experts decide what a correct answer is.
Sources
- Anthropic. Define your success criteria and Create strong empirical evaluations.
- Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.
- Wang et al. (2023). Large Language Models are not Fair Evaluators.
- Shankar et al. (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.
- Es et al. (2023). RAGAS: Automated Evaluation of Retrieval Augmented Generation.
- Hamel Husain. Your AI product needs evals.
- promptfoo documentation.