Multi-agent systems are the most hyped corner of AI engineering. Diagrams with a "planner agent", a "researcher agent", a "critic agent" and a "writer agent" talking to each other look impressive in slides. In production, the picture is more nuanced. Some multi-agent designs deliver results no single agent can match. Many others are slower, more expensive and less reliable than one well-built agent. This article summarises what the teams who run these systems at scale have published, and gives you a framework to decide.
What "multi-agent" actually means
A multi-agent system is any architecture where more than one LLM-driven loop works on a task, each with its own context, instructions and possibly tools. Common shapes:
- Orchestrator-subagents: a lead agent plans, spawns subagents for subtasks, and synthesises their results. Subagents usually do not talk to each other.
- Pipeline of specialists: agents hand work to each other in a fixed order (research → draft → review).
- Peer collaboration / debate: several agents discuss, critique or vote.
- Hierarchies: orchestrators that manage other orchestrators.
The first shape dominates successful production deployments. The others appear more often in research and demos.
The case for multiple agents
Anthropic described how they built the multi-agent research feature in Claude: a lead agent breaks a research question into parts and spawns subagents that search in parallel, each with its own context window. On their internal research evaluation, the multi-agent system outperformed a single agent by 90.2% (Anthropic, 2025).
Why it worked:
- Parallelism. Subagents explore different directions at the same time, cutting wall-clock time on broad tasks.
- Separate contexts. Each subagent can read tens of thousands of tokens of sources and return a condensed summary, so the lead agent's context stays clean. This is context engineering applied at the architecture level.
- More total reasoning. Their analysis found that token usage alone explained most of the performance variance on their browsing evaluation — multi-agent systems work partly because they let you spend more tokens productively.
The case against
The same report is candid about cost: agents used about 4x more tokens than chat interactions, and multi-agent systems about 15x more. That is only worth it when the value of the task is high.
Cognition, the team behind the Devin coding agent, argued in Don't build multi-agents that for tasks like coding, splitting work across agents that do not share full context leads to conflicting decisions: one subagent styles a component one way, another builds a matching piece differently, and the combination fails. Their principles: share full context and full agent traces, and remember that actions carry implicit decisions that conflict if made in isolation.
Research backs up the fragility. A study analysing traces from several popular multi-agent frameworks built a taxonomy of 14 failure modes grouped into specification and system design issues, inter-agent misalignment, and task verification and termination problems (Cemri et al., 2025). Many failures came not from the models but from the system design: unclear roles, lost information in hand-offs, no one responsible for verifying the final result.
Reconciling the two views
The two positions are less contradictory than they appear:
| Task property | Favours multi-agent | Favours single agent |
|---|---|---|
| Subtasks independent? | Yes — research, search, data gathering | No — tightly coupled edits, design |
| Output combined how? | Summaries, lists, findings | A single coherent artefact (code, document) |
| Context per subtask | Large (many sources) | Small or shared |
| Value per task | High; tokens are cheap relative to value | Low or high-volume |
| Latency tolerance | Parallelism helps wall-clock | Single loop is simpler |
Read-heavy, breadth-first tasks — research, due diligence, competitive analysis, log investigation across many services — benefit from subagents. Write-heavy, tightly coupled tasks — implementing a feature, drafting a contract — usually work better with one agent holding the full context, possibly using subagents only for read-only investigation.
Designing an orchestrator-subagent system
If your task fits, these practices, drawn from Anthropic's write-up and our own projects, improve reliability:
Make delegation explicit
The lead agent must give each subagent a complete brief: objective, output format, tools to use, boundaries ("only sources from 2024 onwards", "do not edit files"), and how much effort to spend. Vague delegation ("research the market") produces duplicated work and gaps.
Subagent brief:
Objective: Find pricing for the top 5 Ukrainian e-commerce platforms for
SMB merchants (monthly fee, transaction fee, setup cost).
Scope: Official pricing pages and their 2026 updates only.
Output: JSON array [{platform, monthly_fee_uah, transaction_fee_pct,
setup_fee_uah, source_url, retrieved_on}]. Use null if not published.
Budget: max 10 searches. Stop when all 5 are covered.
Do not: speculate about unpublished prices.
Scale effort to the task
Teach the orchestrator how many subagents to use: one for a simple fact, a few for a comparison, more for broad research. Without guidance, agents tend to over-spawn on easy questions.
Return condensed, structured results
Subagents should return structured findings with sources, not raw transcripts. Large artefacts go to shared storage (files, a database) with references passed back, so the lead agent is not flooded.
Verify at the end
Someone must check the final output against the original question: a dedicated verification step, a citation checker, or a human. Missing verification is one of the main failure categories in the Cemri taxonomy.
Engineer for long-running state
Multi-agent runs take minutes to hours. Plan for failures mid-run: checkpoint progress, make tool calls idempotent, resume instead of restarting. The same reliability patterns as for any distributed system apply; see LLM API reliability.
Observability is not optional
Debugging a multi-agent system without traces is nearly impossible: the failure may be in the lead agent's plan, a subagent's search, or the synthesis. Trace every agent as a span with its brief, tool calls, tokens and output, linked to the parent. OpenTelemetry's GenAI semantic conventions and LLM-specific tools make this manageable; see LLM observability and tracing. Track per-run cost — a single badly prompted orchestrator can spawn dozens of subagents.
Evaluating multi-agent systems
Evaluate end results first: is the final answer correct, complete, well-sourced? Use LLM-as-a-judge with rubrics, calibrated against human review, as described in LLM evals. Then evaluate the process: number of subagents, duplicated work, tokens per task, time to completion. Start with small eval sets — Anthropic notes that early on, a couple of dozen realistic queries reveal large effects; you do not need hundreds to see whether a change helps.
Cost model before you build
A rough calculation saves surprises:
Single agent: ~30k tokens/task × $X per 1M tokens
Multi-agent: lead 40k + 5 subagents × 60k = ~340k tokens/task (~11x)
Worth it if: value_of_better_answer × tasks/month > extra_cost × tasks/month
For a due-diligence report worth hours of analyst time, 11x tokens is trivial. For a support chatbot answering thousands of simple questions a day, it is ruinous. Model routing — a strong model for the lead, cheaper ones for subagents — narrows the gap; see LLM cost optimization.
Business use cases that fit
- Research and analysis: market scans, vendor comparisons, regulatory research across many documents.
- Incident investigation: subagents examine logs, metrics and recent deploys of different services in parallel.
- Large document review: contracts or tender documents split into sections, each analysed for specific risks, then consolidated.
- Data enrichment: many independent records enriched in parallel, each by a subagent with search tools.
Use cases that usually do not fit: customer chat, form processing, code changes in one module, anything latency-sensitive. For those, a workflow or single agent is the better architecture.
FAQ
Do subagents need different models? Not necessarily. A strong lead model with cheaper subagents is a common cost optimisation, but measure quality: weak subagents produce weak findings.
Should agents talk to each other directly? Rarely. Hub-and-spoke through the orchestrator is easier to debug and control than free-form agent conversations.
Is "agent debate" useful? Research shows debate and self-critique can improve some reasoning tasks, but the gains are inconsistent and the costs high. A single evaluator step with clear criteria is usually enough.
What about frameworks for multi-agent systems? They help with plumbing (state, hand-offs, tracing). The hard parts — task decomposition, briefs, verification — remain your design work.
Sources
- Anthropic (2025). How we built our multi-agent research system.
- Cognition (2025). Don't build multi-agents.
- Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail?
- Anthropic (2024). Building effective agents.
- Du et al. (2023). Improving Factuality and Reasoning in Language Models through Multiagent Debate.
- OpenTelemetry. Semantic conventions for generative AI.