"We want an AI agent" is one of the most common requests we hear. Often what the business actually needs is a reliable workflow with one or two LLM calls in the right places. Sometimes it really is an agent. Choosing wrong is expensive in both directions: an over-engineered agent is slow, costly and unpredictable; an under-powered workflow breaks on every case it did not anticipate. This guide explains the building blocks, the patterns built from them and how to choose.

Workflows versus agents

Anthropic's widely cited essay Building effective agents draws a useful line:

  • Workflows are systems where LLMs and tools are orchestrated through predefined code paths. You decide the steps; the model fills in the intelligent parts.
  • Agents are systems where the LLM dynamically directs its own process and tool usage, deciding what to do next based on results.

The essay's main advice is to start with the simplest solution and add complexity only when it demonstrably improves outcomes. That matches our experience: most production systems we build are workflows with an agent inside one well-bounded step.

The building block: an augmented LLM

Every pattern starts with one LLM call enhanced with:

  • Retrieval — relevant documents from your knowledge base; see RAG for business.
  • Tools — functions the model can call: search, database queries, APIs; see designing tools for agents.
  • Memory — relevant facts from previous interactions.

Many problems are solved by a single well-built augmented call. Try that first and measure it with evals before building anything bigger.

Workflow patterns

Prompt chaining

Break a task into sequential steps, each an LLM call that processes the previous output. Add programmatic checks ("gates") between steps.

Example: generate a product description → check length and banned claims in code → translate to Ukrainian → check terminology against a glossary.

Use when the task decomposes cleanly into fixed subtasks. You trade latency for accuracy, since each call has an easier job.

Routing

Classify the input and send it to a specialised prompt, model or pipeline.

Example: support messages routed to billing, technical or account flows; simple questions to a small, cheap model and complex ones to a strong model — a major cost lever described in LLM cost optimization.

Use when inputs fall into distinct categories that are better handled separately.

Parallelization

Run several LLM calls at once and aggregate results in code. Two variants:

  • Sectioning: independent subtasks in parallel, for example checking a contract for five types of risk at once.
  • Voting: the same task several times, taking the majority or the most cautious result, for example flagging content for review if any of three checks flags it.

Use when subtasks are independent or when you need higher confidence.

Orchestrator-workers

A central LLM breaks the task into subtasks dynamically, delegates them to worker calls and synthesises the results. Unlike parallelization, the subtasks are not known in advance.

Example: a coding change that touches an unknown number of files; a research question that requires several searches.

Evaluator-optimizer

One call generates, another evaluates against criteria and returns feedback, and the loop repeats until the evaluator is satisfied or a limit is reached.

Example: translation with nuance, writing a summary that must include specific facts, code that must pass tests.

Use when you have clear evaluation criteria and iteration measurably improves results.

Autonomous agents

An agent is a loop: the model chooses an action, the environment returns a result, the model decides what to do next — the pattern formalised in ReAct. Agents fit open-ended problems where you cannot predict the number of steps or hard-code the path: resolving a support case that may need several lookups, fixing a bug in an unfamiliar codebase, researching a question across many sources.

def run_agent(task: str, tools: dict, max_steps: int = 20) -> str:
    messages = [{"role": "user", "content": task}]
    for step in range(max_steps):
        response = llm(messages=messages, tools=list(tools.values()))
        messages.append({"role": "assistant", "content": response.content})
        if response.stop_reason != "tool_use":
            return response.text  # model decided it is done
        results = []
        for call in response.tool_calls:
            try:
                output = tools[call.name].run(**call.input)
            except Exception as e:
                output = f"Error: {e}"  # let the model see and recover
            results.append(tool_result(call.id, output))
        messages.append({"role": "user", "content": results})
    raise StepLimitExceeded(task)

The loop itself is simple. Everything that makes agents work in production surrounds it: tool design, context management, stopping conditions, permissions, observability and human checkpoints.

How to choose

Question If yes If no
Can you write down the steps in advance? Workflow Consider an agent
Is a wrong action costly or irreversible? Workflow, or agent with human approval Agent may be fine
Is latency critical (under 2–3 s)? Single call or short chain Agents are acceptable
Is the volume high and margin thin? Workflow with routing to small models Agent cost may be fine
Can success be checked automatically? Agent or evaluator-optimizer can self-correct Keep humans in the loop

A practical heuristic: build the workflow first. Where it repeatedly fails because the path cannot be predicted, replace that step with an agent. You end up with a hybrid — deterministic code for the predictable parts, an agent for the open-ended part — that is easier to test and cheaper to run than an end-to-end agent.

Production concerns for agents

Stopping conditions

Agents need explicit limits: maximum steps, maximum tokens, maximum wall time, and maximum spend per task. When a limit is hit, return a useful partial result or escalate to a human — do not just fail silently.

Human checkpoints

For actions with real-world consequences — sending emails, issuing refunds, changing records — require confirmation. This is also the most effective defence against prompt injection; see AI agent security.

Context growth

Every step adds tool results to the context. Long-running agents need compaction, summarisation, external memory or subagents with fresh contexts. This is the subject of context engineering.

Observability

You cannot debug an agent from its final answer. Trace every step: the model's decision, tool calls with arguments, results, tokens and latency. See LLM observability.

Evaluation

Evaluate both the outcome (was the task done correctly?) and the trajectory (how many steps, how many tool errors, any dangerous actions attempted?). An agent that gets the right answer in 40 steps instead of 6 is a cost and latency problem.

Frameworks: use them or not?

Frameworks such as LangGraph, the OpenAI Agents SDK, the Claude Agent SDK, Spring AI or LangChain4j speed up prototyping and provide useful plumbing: tool calling, state management, tracing. They also add abstraction layers that can hide what prompts and calls are actually sent. Our advice:

  • Understand the raw API first — build one agent loop by hand.
  • Adopt a framework when you need its features (persistence, human-in-the-loop, tracing), not by default.
  • Make sure you can see the exact prompts and tool calls the framework sends.

For Java teams, Spring AI is a natural fit; for TypeScript, the Vercel AI SDK covers most needs. Exposing your business systems as tools through the Model Context Protocol makes them reusable across frameworks and clients.

A worked example: invoice processing

Requirement: process incoming supplier invoices into the ERP.

  • Version 1 (workflow): extract fields with structured output → validate totals and tax IDs in code → match the supplier in the ERP by tax ID → create a draft bill. Exceptions go to a human queue. This handles the large majority of invoices.
  • Version 2 (hybrid): for exceptions — unknown supplier, mismatched PO, unusual currency — an agent with read-only tools (search suppliers, search purchase orders, read contract terms) investigates and proposes a resolution for a human to approve.

The agent only runs on the hard cases, with read-only tools and human approval. That is the shape most successful business agents take. More on the extraction part in LLM document processing.

FAQ

Is a chatbot with RAG an agent? Usually not; it is an augmented LLM in a fixed retrieve-then-answer workflow. It becomes an agent when the model decides when and how often to search and which tools to use.

Do we need multiple agents? Rarely at first. Multi-agent systems help with broad, parallelisable tasks and cost significantly more tokens; see multi-agent systems.

How do we test an agent before production? Build a scenario set with expected end states, run it in a sandbox with mocked or staging tools, and grade both outcomes and trajectories.

Which model should an agent use? Agents benefit from the strongest model you can afford for planning and tool selection; subtasks can be routed to cheaper models.

Sources

  1. Anthropic (2024). Building effective agents.
  2. Yao et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models.
  3. Schick et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools.
  4. Anthropic. Tool use with Claude.
  5. OpenAI. A practical guide to building agents.