A new "state-of-the-art" model is announced almost every week, each with a chart showing it beating the competition on a dozen benchmarks. For a team building a product, the practical question is different: which model gives our users the best results, at a cost and latency we can afford, without locking us in? Public benchmarks are a starting point for that decision, not the answer. This guide explains how to read benchmarks, how to build a lightweight evaluation for your own task, and how to weigh the trade-offs.

What public benchmarks measure

Benchmark What it measures Useful for Caveats
SWE-bench / SWE-bench Verified Resolving real GitHub issues in Python repositories Coding agents Python-only; agent scaffolding affects results as much as the model
LMArena (Chatbot Arena) Human pairwise preferences in open-ended chat General assistant quality, style Rewards style and length; prompts differ from your domain
HELM Many scenarios and metrics, including robustness and calibration Holistic comparison Coverage lags newest models
MMLU / MMLU-Pro, GPQA Multiple-choice knowledge and reasoning Rough capability tiers Saturated at the top; contamination risk
MTEB Embedding models Retrieval and RAG See embeddings guide
Long-context tests (RULER and similar) Retrieval and reasoning over long inputs Document-heavy apps Synthetic tasks

SWE-bench, built from 2,294 real issues and pull requests across 12 popular Python repositories, became the standard for coding agents (Jimenez et al., 2023); its human-validated subset, SWE-bench Verified, removed problematic tasks (OpenAI, 2024). Chatbot Arena ranks models by crowdsourced pairwise votes using a Bradley-Terry model (Chiang et al., 2024).

Why benchmarks mislead

  • Contamination. Test questions leak into training data. When researchers built GSM1k, a fresh set mirroring a popular math benchmark, several models dropped substantially, suggesting overfitting to the original (Zhang et al., 2024).
  • Saturation. When top models all score above 90%, differences are noise.
  • Scaffolding. Agent benchmarks measure model + harness. A vendor's number with their custom agent does not transfer to your setup.
  • Distribution mismatch. Your inputs — Ukrainian customer emails, Odoo data, legal contracts — look nothing like benchmark questions.
  • Style bias. Preference-based leaderboards reward longer, more formatted answers that may not suit your product.
  • Missing dimensions. Benchmarks rarely measure refusal behaviour, instruction following in long system prompts, tool-calling reliability or consistency across runs.

Use benchmarks to build a shortlist of 3–5 candidates. Then test on your own task.

Build a task-specific eval in two days

You do not need a research team. A practical eval:

  1. 50–200 real inputs from your domain, including edge cases and the languages your users write in.
  2. Grading per item: code checks where possible (correct JSON, right category, contains required facts), LLM-as-a-judge with a clear rubric where not, plus human spot checks.
  3. Same prompt and settings for all candidates first; then a quick prompt adaptation per model, because models respond differently to the same instructions.
  4. Multiple runs for variance.
  5. Record cost and latency per item, not just quality.

The full methodology is in how to evaluate LLM applications. For agents, include multi-step tasks with tools and grade final state and trajectory.

Model         Quality (judge)  JSON valid  p95 latency  Cost / 1k req
candidate-A        0.91          100%         6.8 s        $14.20
candidate-B        0.89          100%         3.1 s         $4.10
candidate-C        0.82           97%         1.4 s         $0.90

Results like these often lead to a routing decision rather than a single winner.

Dimensions to weigh

Quality on your task

The eval result is the primary signal. Look at failures by category: a model that is 2 points better overall but fails all refund questions may be the wrong choice.

Cost

Price per token is only part of the story. Compare cost per successful task: a cheaper model that needs more retries, longer prompts or more agent steps can cost more. Factor in prompt caching discounts, batch pricing and output-token-heavy workloads; see LLM cost optimization.

Latency

Measure time to first token and total time on your prompts. Reasoning modes improve quality on hard tasks and increase latency; many products use them selectively. For chat, sub-second first token matters more than total time; for background jobs, throughput matters more.

Capabilities

Tool calling reliability, structured output support, context length, vision and PDF input, multilingual quality (test Ukrainian explicitly), and available reasoning or effort controls.

Operational factors

Rate limits at your tier, regional availability, data residency, retention options and DPAs (privacy guide), uptime history, and availability through multiple clouds for fallback (reliability guide).

Open vs closed

Closed (API) models Open-weight models
Quality ceiling Highest for complex reasoning, coding, agents Strong and improving; gap varies by task
Data control Vendor terms, DPAs Full, if self-hosted
Cost model Per token Infrastructure + engineering
Customisation Prompting, some fine-tuning Full fine-tuning, quantization, decoding control
Lifecycle Models deprecated on vendor schedule You decide when to upgrade

Many teams combine both: frontier API models for hard tasks, open models for sensitive or high-volume ones; see self-hosting LLMs.

Routing: you rarely need one model

The best production setups route requests:

  • By task: a small fast model for classification and extraction; a frontier model for complex reasoning, coding and agent planning.
  • By difficulty: try the cheaper model first; escalate when validation fails or confidence signals are weak.
  • By user tier or latency budget.

Routing turns model selection from a one-time bet into a configuration you can tune. It is one of the levers in AI agent architecture as well.

Avoiding lock-in

  • Abstract the call site. A thin internal interface (or a library such as the Vercel AI SDK, Spring AI or LiteLLM) keeps provider-specific code in one place; see AI chat in Next.js and Spring AI.
  • Keep prompts and evals in your repository. They are the real asset; with them, switching is a measured experiment.
  • Avoid deep dependence on proprietary features unless they deliver clear value — and when you use them (server-side conversation state, proprietary file stores), know how you would migrate.
  • Re-evaluate on a schedule. Every quarter or with major releases, run your eval on new candidates. Model deprecations will force migrations anyway; being ready makes them routine.

Model upgrades are migrations

Switching to a newer version of the same model family is not free. Newer models may follow instructions more literally, change output length or formatting, or handle tools differently. Treat every upgrade as a change: run evals, compare by category, read a sample of outputs, roll out gradually, and keep the previous model available as a fallback until you are confident. Vendors publish migration notes; read them before upgrading.

A selection process that works

  1. Define the task, success criteria, latency and budget constraints.
  2. Shortlist 3–5 models using public benchmarks, capabilities and operational requirements.
  3. Run your eval with the same prompt, then with light per-model prompt tuning.
  4. Compare quality by category, cost per successful task and latency.
  5. Decide on a primary model, a fallback, and possibly a router.
  6. Ship behind a feature flag; monitor production quality and cost.
  7. Re-run the eval quarterly and on major releases.

FAQ

Should we just pick the top model on LMArena? It is a reasonable default for a general chat assistant, but not for extraction, classification, coding agents or domain-specific tasks. Test on your data.

How many eval items do we need to compare models? Fifty well-chosen items reveal large differences; 200+ are needed to trust differences of a few points. Report confidence intervals or at least variance across runs.

Are open models good enough now? For many narrow tasks, yes. For complex agents and coding, frontier closed models usually still lead. Evaluate for your case.

How often do we need to switch? Only when evals show meaningful gains or a model is deprecated. Chasing every release is expensive; scheduled re-evaluation is enough.

Sources

  1. Jimenez et al. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?; OpenAI (2024). Introducing SWE-bench Verified.
  2. Chiang et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference; LMArena.
  3. Liang et al. (2022). Holistic Evaluation of Language Models (HELM).
  4. Zhang et al. (2024). A Careful Examination of Large Language Model Performance on Grade School Arithmetic.
  5. Artificial Analysis — independent model price, speed and quality comparisons.
  6. Anthropic. Models overview.