A new "state-of-the-art" model is announced almost every week, each with a chart showing it beating the competition on a dozen benchmarks. For a team building a product, the practical question is different: which model gives our users the best results, at a cost and latency we can afford, without locking us in? Public benchmarks are a starting point for that decision, not the answer. This guide explains how to read benchmarks, how to build a lightweight evaluation for your own task, and how to weigh the trade-offs.
What public benchmarks measure
| Benchmark | What it measures | Useful for | Caveats |
|---|---|---|---|
| SWE-bench / SWE-bench Verified | Resolving real GitHub issues in Python repositories | Coding agents | Python-only; agent scaffolding affects results as much as the model |
| LMArena (Chatbot Arena) | Human pairwise preferences in open-ended chat | General assistant quality, style | Rewards style and length; prompts differ from your domain |
| HELM | Many scenarios and metrics, including robustness and calibration | Holistic comparison | Coverage lags newest models |
| MMLU / MMLU-Pro, GPQA | Multiple-choice knowledge and reasoning | Rough capability tiers | Saturated at the top; contamination risk |
| MTEB | Embedding models | Retrieval and RAG | See embeddings guide |
| Long-context tests (RULER and similar) | Retrieval and reasoning over long inputs | Document-heavy apps | Synthetic tasks |
SWE-bench, built from 2,294 real issues and pull requests across 12 popular Python repositories, became the standard for coding agents (Jimenez et al., 2023); its human-validated subset, SWE-bench Verified, removed problematic tasks (OpenAI, 2024). Chatbot Arena ranks models by crowdsourced pairwise votes using a Bradley-Terry model (Chiang et al., 2024).
Why benchmarks mislead
- Contamination. Test questions leak into training data. When researchers built GSM1k, a fresh set mirroring a popular math benchmark, several models dropped substantially, suggesting overfitting to the original (Zhang et al., 2024).
- Saturation. When top models all score above 90%, differences are noise.
- Scaffolding. Agent benchmarks measure model + harness. A vendor's number with their custom agent does not transfer to your setup.
- Distribution mismatch. Your inputs — Ukrainian customer emails, Odoo data, legal contracts — look nothing like benchmark questions.
- Style bias. Preference-based leaderboards reward longer, more formatted answers that may not suit your product.
- Missing dimensions. Benchmarks rarely measure refusal behaviour, instruction following in long system prompts, tool-calling reliability or consistency across runs.
Use benchmarks to build a shortlist of 3–5 candidates. Then test on your own task.
Build a task-specific eval in two days
You do not need a research team. A practical eval:
- 50–200 real inputs from your domain, including edge cases and the languages your users write in.
- Grading per item: code checks where possible (correct JSON, right category, contains required facts), LLM-as-a-judge with a clear rubric where not, plus human spot checks.
- Same prompt and settings for all candidates first; then a quick prompt adaptation per model, because models respond differently to the same instructions.
- Multiple runs for variance.
- Record cost and latency per item, not just quality.
The full methodology is in how to evaluate LLM applications. For agents, include multi-step tasks with tools and grade final state and trajectory.
Model Quality (judge) JSON valid p95 latency Cost / 1k req
candidate-A 0.91 100% 6.8 s $14.20
candidate-B 0.89 100% 3.1 s $4.10
candidate-C 0.82 97% 1.4 s $0.90
Results like these often lead to a routing decision rather than a single winner.
Dimensions to weigh
Quality on your task
The eval result is the primary signal. Look at failures by category: a model that is 2 points better overall but fails all refund questions may be the wrong choice.
Cost
Price per token is only part of the story. Compare cost per successful task: a cheaper model that needs more retries, longer prompts or more agent steps can cost more. Factor in prompt caching discounts, batch pricing and output-token-heavy workloads; see LLM cost optimization.
Latency
Measure time to first token and total time on your prompts. Reasoning modes improve quality on hard tasks and increase latency; many products use them selectively. For chat, sub-second first token matters more than total time; for background jobs, throughput matters more.
Capabilities
Tool calling reliability, structured output support, context length, vision and PDF input, multilingual quality (test Ukrainian explicitly), and available reasoning or effort controls.
Operational factors
Rate limits at your tier, regional availability, data residency, retention options and DPAs (privacy guide), uptime history, and availability through multiple clouds for fallback (reliability guide).
Open vs closed
| Closed (API) models | Open-weight models | |
|---|---|---|
| Quality ceiling | Highest for complex reasoning, coding, agents | Strong and improving; gap varies by task |
| Data control | Vendor terms, DPAs | Full, if self-hosted |
| Cost model | Per token | Infrastructure + engineering |
| Customisation | Prompting, some fine-tuning | Full fine-tuning, quantization, decoding control |
| Lifecycle | Models deprecated on vendor schedule | You decide when to upgrade |
Many teams combine both: frontier API models for hard tasks, open models for sensitive or high-volume ones; see self-hosting LLMs.
Routing: you rarely need one model
The best production setups route requests:
- By task: a small fast model for classification and extraction; a frontier model for complex reasoning, coding and agent planning.
- By difficulty: try the cheaper model first; escalate when validation fails or confidence signals are weak.
- By user tier or latency budget.
Routing turns model selection from a one-time bet into a configuration you can tune. It is one of the levers in AI agent architecture as well.
Avoiding lock-in
- Abstract the call site. A thin internal interface (or a library such as the Vercel AI SDK, Spring AI or LiteLLM) keeps provider-specific code in one place; see AI chat in Next.js and Spring AI.
- Keep prompts and evals in your repository. They are the real asset; with them, switching is a measured experiment.
- Avoid deep dependence on proprietary features unless they deliver clear value — and when you use them (server-side conversation state, proprietary file stores), know how you would migrate.
- Re-evaluate on a schedule. Every quarter or with major releases, run your eval on new candidates. Model deprecations will force migrations anyway; being ready makes them routine.
Model upgrades are migrations
Switching to a newer version of the same model family is not free. Newer models may follow instructions more literally, change output length or formatting, or handle tools differently. Treat every upgrade as a change: run evals, compare by category, read a sample of outputs, roll out gradually, and keep the previous model available as a fallback until you are confident. Vendors publish migration notes; read them before upgrading.
A selection process that works
- Define the task, success criteria, latency and budget constraints.
- Shortlist 3–5 models using public benchmarks, capabilities and operational requirements.
- Run your eval with the same prompt, then with light per-model prompt tuning.
- Compare quality by category, cost per successful task and latency.
- Decide on a primary model, a fallback, and possibly a router.
- Ship behind a feature flag; monitor production quality and cost.
- Re-run the eval quarterly and on major releases.
FAQ
Should we just pick the top model on LMArena? It is a reasonable default for a general chat assistant, but not for extraction, classification, coding agents or domain-specific tasks. Test on your data.
How many eval items do we need to compare models? Fifty well-chosen items reveal large differences; 200+ are needed to trust differences of a few points. Report confidence intervals or at least variance across runs.
Are open models good enough now? For many narrow tasks, yes. For complex agents and coding, frontier closed models usually still lead. Evaluate for your case.
How often do we need to switch? Only when evals show meaningful gains or a model is deprecated. Chasing every release is expensive; scheduled re-evaluation is enough.
Sources
- Jimenez et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?; OpenAI (2024). Introducing SWE-bench Verified.
- Chiang et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference; LMArena.
- Liang et al. (2022). Holistic Evaluation of Language Models (HELM).
- Zhang et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic.
- Artificial Analysis — independent model price, speed and quality comparisons.
- Anthropic. Models overview.