Tag Archives: Agent Design Patterns

Multi-Agent Design Patterns: When Do They Actually Help? Three small diagrams: one agent with tools; a coordinator with three sub-agents; a generator, reviewer and correction flow.

Multi-Agent Design Patterns: When Do They Actually Help?

Objective

The goal of this post is to show, with real numbers from a working system, how to decide whether an agentic system should be a single agent or multiple agents working together.

I’ve read about cases where multi-agent systems helped, for example:

  • Anthropic’s research system: a lead agent sends several research agents to investigate different parts of a question in parallel, then combines what they find.
  • Cognition’s reviewer pattern: a second agent reviews the first agent’s work and sends back what needs fixing.
  • LangChain’s language comparison: to compare Python, JavaScript and Rust for web development, a separate agent researches each language and a coordinator combines the results.
  • OpenAI’s Navier–Stokes proof: thousands of agents explored different approaches to a long-open math problem in parallel, sharing their best findings across groups.

I wanted to know whether any of these patterns would help my own marketplace support agent: an AI agent that answers support staff’s questions about orders, payouts, refunds and policies in a small simulated marketplace, using read-only SQL and policy search. More details are in the use case below. This post walks through the patterns, how I looked for one that fits a support agent, the two designs I tested, and what I found.

On my agent:

  • Splitting the work across seller agents did not help. A coordinator agent with three seller agents gave the same answers as one agent with parallel tools, but took longer and cost about 2.5× more.
  • A reviewer agent did help, but only when it was a stronger model. GPT-5.4 reviewing GPT-5 nano’s answers removed every incorrect answer at about half the cost of running GPT-5.4 alone. When nano reviewed nano, nothing improved.

The lesson: As per my understanding, multi-agent helps in two situations below, there can be more.

  • The work splits into independent pieces that can run in parallel, each with its own context. When the steps are mostly sequential, or one query already covers the work, extra agents don’t help.
  • A second agent can see what the first one can’t.

Agent Design Patterns

Before comparing design patterns, it helps to ask one question: who controls the next step, the code or the model?

  • Workflow: in a workflow, code defines the path.
  • Agent: an agent’s model decides the next step, such as which tool to call.
  • Combinations: real systems often combine code-controlled and model-controlled steps. For example, code fixes the order of the steps, while agents decide how to carry out each one.

More model calls, different models or parallel tool calls don’t by themselves make a system multi-agent.

These are the patterns I considered:

PatternTypeHow it flowsExample
Workflow (chaining, routing, parallel steps)WorkflowCode defines each stepMy model-routing posts: classify the question, then send it to a cheap or strong model
Single agent with toolsAgentOne model loop decides which tool to call nextMy support agent: catalog → SQL → policy search → answer
Coordinator with sub-agentsMulti-agentEach agent works autonomously: a coordinator agent splits the task, sub-agents work in parallel in their own contexts, and the coordinator combines the resultsAnthropic’s research system; LangChain’s language comparison
Generator–evaluatorWorkflow + agentsCode fixes the order (draft → review → revise if needed); a generator agent drafts and an evaluator agent checksCognition’s reviewer; Anthropic’s verification subagent
HandoffWorkflow + agentsCode defines which specialist agents exist and who can hand off to whom; an agent decides when to hand the conversation over, and the specialist takes overA triage agent hands a refund request to a refunds agent, which continues with the user
Independent agents across systemsMulti-agent across organizationsAgents owned by different organizations exchange tasks and results over a protocol such as A2AA marketplace’s support agent asks a payment provider’s agent to investigate a failed payout

Why add agents at all? Anthropic’s January 2026 post on when to use multi-agent systems gives three reasons:

  • context protection: keep irrelevant detail out of an agent’s context
  • parallelization: investigate independent parts at the same time
  • specialization: focused tools and prompts

It also warns that multi-agent systems typically use 3–10× more tokens, and advises starting with the simplest approach that works.

Three examples show when splitting pays off:

  • Anthropic’s research system: a lead agent breaks a research question into parts, several research agents search in parallel, each in its own context, and the lead agent combines their findings.
  • OpenAI’s Navier–Stokes proof: on the order of 10,000 agents worked in groups, each trying a different variant or approach, and Codex shared the best findings across groups. They reached a proposed proof of the 90-year-old Millennium Prize problem in 88 hours, using about 130 billion output tokens, and a separate Lean step verified it. Mathematicians have not yet accepted it.
  • LangChain’s language comparison: for “Compare Python, JavaScript, and Rust for web development”, a coordinator sends each language to its own agent, which researches it with about 2,000 tokens of language-specific material. The coordinator then combines the three results.
    • Separate agents: 5 model calls and about 9K tokens.
    • Handoffs: passing control from one agent to the next takes 7+ calls and about 14K tokens, because the steps run one after another.
    • One agent holding all three languages’ material: about 15K tokens.

LangChain credits the 67% fewer tokens to “context isolation”.

All three have the same shape: independent pieces of work, each needing its own context, combined at the end. Cognition’s reviewer is a different idea: a second agent with fresh context checks the first agent’s work.

Not everyone finds that more agents help. Google’s research on scaling agent systems finds the benefit depends on how well the task decomposes, and a 2026 paper found single agents outperform multi-agent systems on multi-hop reasoning when both get the same thinking budget.

The Use Case

The agent is an internal customer support agent I built over a small simulated marketplace with buyers, sellers, orders, shipments, seller payouts, returns, refunds and versioned policies. Support staff ask it questions in plain English. It answers by calling read-only tools (a data catalog, SQL over the business data, order history and a policy-document search) and cites the evidence it used. It runs on the OpenAI Agents SDK behind a LiteLLM gateway.

I use two models: GPT-5 nano, the lower-cost model, and GPT-5.4, the stronger one. More background on the agent is in my model routing posts.

Looking for a Multi-Agent Fit

My first idea was a SQL agent and a RAG agent: one for the business data, one for the policy documents. It didn’t make sense.

  • The two are mostly sequential. The records SQL finds decide which policy to look up: which payout, which date, which return.
  • Combining contexts adds complexity. Splitting them means passing one agent’s findings to the other, then combining both into one answer, with nothing gained.
  • SQL and documents are tools, not separate jobs. One agent can use both.

The second idea looked more promising: one agent per seller. When payouts don’t go through, each seller can have different reasons: an item not yet delivered, a hold that hasn’t passed, an active return, a failed provider attempt. Investigating each seller separately looked like independent branches, the parallelization case. That became the first experiment.

The second experiment came from a different observation. In my earlier evaluations, nano’s wrong answers often had the right evidence available but drew the wrong conclusion. A second agent checking the work seemed worth testing.

Experiment 1: A Coordinator With Seller Agents

The task was a single investigation: “Investigate all sellers’ payouts, explain supported reasons, include paid and zero-due controls, and summarize unpaid totals by seller.” A fixed test dataset had 3 sellers and 11 payouts covering seven outcomes:

  • not delivered
  • under the 24-hour hold
  • active return
  • eligible but not yet released
  • failed provider attempt
  • fully refunded (nothing due)
  • paid

I compared two designs, both using GPT-5.4 and the same five read-only tools:

  • One agent with parallel tool calls allowed. This is not multi-agent: it’s a single agent that can run several tool calls at the same time.
  • A coordinator agent with three seller agents. Each seller agent has its own tool loop and context. The coordinator delegates, then combines their findings.
Two designs side by side. Left: one support agent, a single model loop that chooses the next tool, calls read-only tools: catalog, SQL for all payouts, ten order histories at once, then policy search, and returns the answer. Right: a coordinator agent receives the question, delegates to three seller agents that each have their own tools and context and return findings, then combines the findings into the answer.
The same investigation, done by one agent with parallel tools or by a coordinator with three seller agents.

Each design ran three times, in rotated order. An LLM reviewer checked every answer against the expected payouts, sellers and amounts. Here are the averages; one coordinator run did not complete, and its recorded usage is included.

DesignMean timeMean cost (cache-adjusted)Mean model requests
One agent, parallel tools82.0 s$0.1626.0
Coordinator + 3 seller agents108.1 s$0.40119.7

The answers were equally right on the business facts: every payout covered, the right seller and order, correct amounts and totals. The difference was time and cost. The coordinator design used about 3× the model requests and cost about 2.5× more.

The reasons are visible in the traces:

  • One query covered every seller. The single agent found all the payouts with one SQL query, so there was no heavy independent branch to hand off. Each seller’s investigation was small.
  • The single agent already worked in parallel. After finding the order IDs, it fetched 10 order histories at once in two of the three runs. Parallel tool calls gave concurrency without extra agents.
  • Sub-agents add overhead: each one repeats the instructions and catalog lookup, then packages its findings and evidence for the coordinator, which then has to combine them.

This is the opposite of LangChain’s language example. There, each language needs its own research and its own context; here, one query already held everything.

The limits matter. This was one task, three runs and three sellers, all under a common payout policy, and I didn’t give both designs equal budgets. With many more sellers, or seller-specific policies that need real independent investigation, sub-agents might pay off. Not shown here is the honest conclusion, not “sub-agents never help”.

Experiment 2: The Generator–Evaluator Pattern

The second design adds a reviewer. Anthropic calls this pattern evaluator-optimizer; I’ll call it generator–evaluator.

  1. A generator agent (nano) drafts the answer with tools.
  2. A reviewer agent sees the question, the draft and the full evidence the generator used. It has no answer key, and in this experiment no tools. It returns “no changes needed” or specific feedback.
  3. If the reviewer flags problems, a correction agent (nano) revises the answer once, with tools available.

The application controls the order of the steps; the generator and correction agents run their own tool loops.

A left-to-right flow. A generator agent (GPT-5 nano) drafts the answer with tools and passes it to a reviewer agent (nano or GPT-5.4), which checks the draft against the evidence. If no changes are needed, the draft becomes the final answer. If changes are needed, feedback goes to a correction agent (GPT-5 nano), which makes one correction with tools and produces the final answer.
The generator–evaluator flow: draft, review, and at most one correction.

I ran it on the same 18 support cases I use for regression testing: order lookups, current and historical policy, payout eligibility, refund histories, reconciliation, an ambiguous name and a request the agent must refuse. Two cases are short follow-up conversations, giving 20 answers in total. An LLM reviewer graded every final answer against the expected results, separately from the reviewer agent’s own verdict. These grades are not human-adjudicated.

How it unfolded. I first ran it on 8 hard cases, with GPT-5.4 reviewing nano’s answers. Results moved from 3 correct / 3 partial / 2 incorrect to 5 / 3 / 0. It looked as if the pattern worked.

But was it the pattern, or a stronger model doing the reviewing? So I tested nano reviewing nano on the same 8 cases. It asked for no changes, and nothing improved. Then I ran both reviewers across all 18 cases:

DesignCorrect / partial / incorrectEstimated cost
One agent: nano11 / 4 / 3$0.037
Nano generator + nano reviewer11 / 4 / 3$0.053
Nano generator + GPT-5.4 reviewer14 / 4 / 0$0.582
One agent: GPT-5.417 / 1 / 0$1.210

Costs are model-usage estimates at uncached list prices, including nano’s drafts and the review and correction calls.

Scatter chart of cost against fully correct cases for the four designs. One agent nano: 11 of 18 at $0.037. Nano generator plus nano reviewer: 11 of 18 at $0.053, marked as dominated. Nano generator plus GPT-5.4 reviewer: 14 of 18 at $0.58. One agent GPT-5.4: 17 of 18 at $1.21. A line connects the three non-dominated designs.
Same-model review added cost without changing results. A stronger reviewer moved nano most of the way toward GPT-5.4 at about half its cost.

What the results show:

  • A reviewer can only catch what it can see. Across all 18 cases, nano as reviewer asked for three corrections, and none changed a grade. It shares the generator’s blind spots.
  • The stronger reviewer removed all three incorrect answers. Three answers became correct and one improved from incorrect to partial. The total came in about 52% below running GPT-5.4 alone, with three fewer fully correct answers.
  • Detection isn’t correction. GPT-5.4 asked for 13 corrections, and 4 answers improved. Several times nano was told what evidence was missing and still didn’t fetch it.
  • The gain comes from the reviewer’s capability. The pattern is how you spend that capability selectively: GPT-5.4 reads and judges nano’s work instead of doing all of it.

On a cost-versus-accuracy chart, nano reviewing nano is dominated: the same accuracy as nano alone, at higher cost. The useful choices are nano alone, nano with a GPT-5.4 reviewer, and GPT-5.4 alone. Which one is right depends on your accuracy bar and budget.

The limits: one pass over 18 known cases, and LLM-reviewer grades. This doesn’t show that self-review never works; it shows it didn’t work here.

When Multiple Agents Make Sense Across Systems

The experiments above looked at splitting one application’s work across agents. Multiple agents also make sense when the work crosses a boundary between systems, especially systems owned by different companies. The marketplace and a payment provider are two systems run by two different companies. The marketplace’s support agent could ask the provider’s agent to investigate a failed payout. Each side keeps its own tools, data and permissions, and they exchange only the task and the result.

That’s what the A2A (Agent2Agent) protocol is for: delegating work to an agent you don’t control. Google introduced it, and it’s now an open project under the Linux Foundation. Inside your own system, tools exposed through function calls or MCP are usually enough. I haven’t implemented A2A here, so I’m not claiming a measured benefit.

Lessons Learned

  • Look for independent branches, not tools. SQL and document search were mostly sequential, so separate agents would only have added a context-merging step.
  • One good query can beat a team. When a single bulk query covered every seller, sub-agents added time, cost and model calls without better answers.
  • Parallel tools give you concurrency without extra agents. The single agent fetched ten order histories at once, though it still had to find the orders first.
  • A reviewer is only as good as what it can see. Same-model review found nothing. A stronger reviewer removed every incorrect answer, which makes review a way to spend a strong model only where it matters.
  • Detection isn’t correction. Specific feedback helped most, but the correction agent didn’t always fetch the evidence it was told was missing.

A Checklist: Before You Add Multiple Agents

  1. Can one agent with the right tools, and parallel tool calls, do it?
  2. Are each agent’s task and context separate and isolated, or would the agents need each other’s details?
  3. Can the agents’ work run in parallel, or is it mostly sequential?
  4. Is one agent’s context actually overloaded? Measure it.
  5. For a reviewer: is it more capable than the generator on these errors? Test same-model review as a control.
  6. Compare against single-agent baselines on accuracy, cost, latency and failures.
  7. Across system boundaries, such as another team or another company, consider independent agents and A2A.

Limits

  • Both experiments are small: one investigation task with three runs, and 18 known cases with one pass.
  • Grades come from an LLM reviewer against expected answers, not human adjudication.
  • Costs are model-usage estimates, not invoices.

Summary

The lesson is simple: add multiple agents only for a reason you can name and measure. On my support agent, a coordinator with seller agents gave the same answers as one agent with parallel tools, at about 2.5× the cost, because one query already covered every seller. A reviewer agent helped only when it was stronger than the generator: GPT-5.4 reviewing nano removed every incorrect answer at about half the cost of GPT-5.4 alone, while nano reviewing nano changed nothing. Multi-agent pays off when the work splits into independent pieces that can run in parallel, each with its own context, or when a second agent can see what the first one can’t.

Acknowledgements

This post was written with the help of Claude Opus 5.5 in Claude Code, which also helped me research the multi-agent patterns and plan the experiments. The experiments were built and run with GPT-6 Astra, which also graded the answers against the expected results.

References

Patterns and Evidence

Earlier Posts