Multi-Agent System

Multi-Agent System

Definition: A multi-agent system (MAS) is an architecture in which multiple autonomous AI agents — each with a distinct role, toolset, and objective — interact within a shared environment to solve a task that exceeds the capacity, context window, or reliability of any single agent acting alone. Agents communicate through structured messages, a shared memory store, or an orchestrating controller, and coordination ranges from strict hierarchical delegation to decentralized negotiation between peers. The defining property is not “more than one model call” but genuine division of labor: each agent specializes, and the system’s overall behavior emerges from how their outputs are routed, combined, and verified. Modern LLM-based multi-agent systems apply decades-old distributed-AI theory — contract nets, blackboard systems, negotiation protocols — to language models acting as the reasoning unit inside each agent.

How It Works

A MAS is defined by three design choices: what roles the agents play, how they exchange information, and who (if anyone) is in charge of sequencing the work. Getting these three choices right is most of the engineering effort in building a reliable system.

Agent Roles and Specialization

  • A planner decomposes the overall goal into an ordered or dependency-graphed list of subtasks
  • A researcher/retriever gathers external information via search, APIs, or a Retrieval-Augmented Generation (RAG) pipeline
  • An executor (coder, form-filler, API caller) performs concrete actions using Function Calling (Tool Use)
  • A critic/verifier checks a prior agent’s output against explicit constraints before it’s allowed to propagate
  • A synthesizer merges multiple agents’ partial outputs into one coherent final answer
  • A router/classifier inspects an incoming request and decides which downstream specialist should handle it, common at the entry point of larger systems
  • Role-narrowing is not cosmetic: a single agent juggling planning, execution, and verification in one context window suffers from prompt dilution, where instructions for one sub-role compete for attention with instructions for another
  • A narrowly scoped agent gets a tight system prompt, a minimal relevant tool subset, and a context window containing only what that role needs — all of which measurably improve task-specific accuracy over a monolithic “do everything” prompt
  • Roles can be static (fixed at design time, e.g., always “planner then coder then reviewer”) or dynamic (an orchestrator decides at runtime which specialist a given subtask needs)
  • The same underlying model can play every role, differentiated purely by system prompt and tool access, or different roles can be assigned to genuinely different models sized to each job’s difficulty
  • Clear role boundaries also make failure attribution possible after the fact — if the “coder” role never has permission to touch deployment config, a deployment incident can immediately rule that agent out

Communication and Coordination Protocols

  • Direct message passing — structured messages, typically JSON, sent agent-to-agent or agent-to-orchestrator, each carrying a task description, prior context, and an expected output schema
  • Shared memory / blackboard — a common data store (a document, a database row, a vector index) that any agent can read from and write to, avoiding the need for agents to know about each other directly
  • Tool-call relay — one agent’s output becomes the literal input arguments to another agent’s tool call, chaining execution without a natural-language message in between
  • Publish/subscribe channels — agents emit events to a topic and any interested agent consumes them, useful when the set of participating agents is dynamic or unknown in advance
  • Coordination can be synchronous, where the orchestrator waits for one agent to finish before invoking the next, which is simple to reason about and debug
  • Coordination can instead be asynchronous/parallel, where independent agents run concurrently and their results are joined once all have returned, which cuts wall-clock latency but requires a join step to reconcile conflicting results
  • Message schemas matter more than they first appear to: an underspecified schema (free-text instead of structured fields) makes downstream parsing brittle and error-prone as the system scales past two or three agents
  • Versioning message schemas becomes necessary once a system has more than a handful of agents maintained by different people, for the same reason API versioning matters between services

Orchestration Patterns

The most common production pattern is orchestrator-worker: a central controller agent owns the task list, dispatches subtasks to specialized workers, and aggregates their results before deciding the next step or returning a final answer. This mirrors a manager delegating to a team, and it concentrates termination conditions, budget limits, and safety checks in one place that’s easy to audit.

Two alternative topologies show up when a strict hierarchy doesn’t fit the problem. The sequential pipeline is the simplest of all — no branching, no aggregation step, just a fixed relay:

  • In a sequential pipeline, agents form a fixed chain — agent A’s output is agent B’s only input, with no branching or return path, which suits linear workflows like draft → fact-check → edit
  • In a peer-to-peer topology, agents address each other directly with no central router, useful for negotiation-style tasks (two agents bargaining over a resource split, or a debate between an advocate and a skeptic agent) but harder to keep observable and bounded
  • Hybrid topologies are common in practice: a top-level orchestrator delegates to a sub-orchestrator that itself runs a small peer-to-peer negotiation before reporting a single consolidated result upward

State, Memory, and Context Propagation

  • A structured task object — a JSON blob that accumulates fields as it passes through agents — is the simplest way to carry state across handoffs
  • A shared scratchpad that every agent can append to and read works well when agents need visibility into each other’s intermediate reasoning
  • A vector store that agents write intermediate findings into, retrieved later via Retrieval-Augmented Generation (RAG)-style lookups, scales better than a growing text blob once the amount of accumulated state gets large
  • Naively forwarding an entire conversation transcript to every downstream agent causes unbounded context growth as the pipeline gets longer
  • Production systems typically summarize state at each handoff, keeping only what the next agent actually needs rather than everything that came before
  • Termination needs an explicit mechanism — a maximum iteration count, a convergence check where no agent proposes further changes, or an explicit “done” signal from the orchestrator or a designated verifier
  • Without a hard stop condition, agents can loop indefinitely, re-litigating the same decision without ever converging on a final answer
  • Idempotency at the state-write level — repeating the same update twice produces the same result as doing it once — protects against duplicate side effects when a retry follows a handoff whose acknowledgment was lost
  • Persistent state (a database row, a file) survives across sessions and restarts; ephemeral state (an in-memory object passed between calls) is simpler but lost if the process crashes mid-task

Implementation Topologies: In-Process vs Distributed

  • In-process — all agents run as function calls within a single application process, sharing memory directly; simplest to build and debug, but a crash takes down every agent at once and the system can’t scale one role independently of the others
  • Distributed/service-based — each agent runs as its own service behind an API, communicating over the network; scales and fails independently, and different agents can be deployed, versioned, and scaled on different schedules, at the cost of added latency and operational complexity
  • Most production systems start in-process and only split agents into separate services once one of them needs independent scaling, a different deployment cadence, or isolation for security reasons
  • Statelessness at the agent boundary — an agent that takes a well-defined input and returns a well-defined output without hidden dependencies on prior calls — is what makes either topology tractable; agents with hidden mutable state are hard to test, retry, or run in parallel regardless of deployment style
  • Retry semantics differ sharply between the two: an in-process retry is just calling the function again, while a distributed retry must account for the possibility that the previous call partially succeeded before failing (a tool call that fired but whose result never made it back)

When a Multi-Agent System Is (and Isn’t) the Right Choice

Signals a Split Into Multiple Agents Will Help

  • The task naturally decomposes into subtasks requiring genuinely different skills or tool access (research vs. code generation vs. numerical verification)
  • Subtasks are largely independent and can run in parallel, so splitting reduces wall-clock latency, not just conceptual complexity
  • A single-agent prompt has grown so long and multi-purpose that the model starts missing instructions or conflating unrelated parts of the task
  • The task benefits from an explicit adversarial or verification step — a second opinion catches errors a single pass would miss
  • Different subtasks justify different cost/latency tradeoffs, so routing cheap models to easy steps and an expensive model to the hard step meaningfully reduces spend
  • The task has a high enough cost of error that an independent verification pass is worth its extra latency and spend — irreversible actions, financial transactions, or anything touching production systems all raise that bar

Signals a Single Agent Is Still the Better Choice

  • The task is short, well-scoped, and doesn’t require distinct expertise at different stages
  • Added coordination latency (extra round trips between agents) would hurt the user experience more than specialization would help it
  • The team can’t yet observe or debug a single agent reliably — adding more agents multiplies debugging surface before the basics are solid
  • The task’s failure modes are well understood and low-stakes enough that the extra reliability from a verifier agent isn’t worth its cost
  • A simpler alternative — better Prompt Engineering or a single agent with more tools — hasn’t been tried yet and might close the gap without the added complexity
  • The task’s subtasks share so much context that splitting them would mean re-sending most of the same information to every agent anyway, eliminating the context-window benefit that motivated the split in the first place

Why It Matters

  • Task decomposition beats monolithic prompting on long-horizon work. Benchmarks on complex, multi-step tasks consistently show specialized agents with narrow context outperforming a single agent trying to hold an entire task in one context window
  • Context window management. Specialization keeps each agent’s working context focused on its sub-problem instead of accumulating every instruction, tool schema, and intermediate result the whole task might ever need
  • Parallelism cuts latency. Independent subtasks — three research queries, five file edits — can run concurrently instead of serially, which matters for interactive products where users are waiting on a response
  • Fault isolation. A malfunctioning tool call or a bad prompt in the “coder” agent doesn’t require touching the “planner” or “reviewer” agent’s logic; a bug’s blast radius is contained to one role
  • Reusability across pipelines. A well-built verifier or summarizer agent can be dropped into many different workflows, the same way a well-written function gets reused across a codebase
  • Human-comprehensible checkpoints. Discrete handoffs between agents create natural points for a human to review, approve, or intervene before a costly or irreversible action is taken
  • Emergent capability through interaction. Debate-style setups, where one agent argues a position and another critiques it, surface errors and blind spots that neither agent catches reasoning alone
  • Cost and latency become tunable. Cheap, fast models can be routed to easy sub-tasks (classification, formatting) while an expensive frontier model is reserved for the hard reasoning step
  • Industry adoption is concrete, not speculative. Production coding assistants split planning, code generation, and code review across distinct agent invocations; support platforms route tickets to specialist agents by category; research tools spin up parallel sub-agents before synthesizing a report
  • It’s a practical path to more capable systems today. Rather than waiting for a single model to improve at everything simultaneously, MAS lets teams ship reliability gains now by composing several instances of current models around a well-designed protocol
  • Alignment and safety benefit from separation of duties. A dedicated critic agent whose sole job is catching unsafe or out-of-policy actions is easier to audit and harden than one clause buried inside a much longer general-purpose prompt, relevant to AI Alignment
  • It maps naturally onto existing organizational thinking. Engineering teams already think in terms of roles, handoffs, and review gates, which makes multi-agent architectures easier to design, staff, and reason about than an opaque single model expected to do everything
  • It creates a natural audit trail. Because state moves between agents as discrete messages rather than living inside one continuous stream of reasoning, it’s easier after the fact to reconstruct exactly which step produced which decision

Coordination Patterns

Three coordination topologies cover most real deployments, each with different failure characteristics:

PatternControl FlowBest ForTypical Failure Mode
Hierarchical (orchestrator-worker)Central controller assigns and aggregates; workers never talk to each other directlyWell-decomposable tasks with a clear manager role (coding pipelines, research-and-report tools)Orchestrator becomes a bottleneck or single point of failure; over-centralized logic re-creates a monolith
Peer-to-peer negotiationAgents exchange messages directly, no central router; consensus reached via protocol (bidding, voting, debate)Resource allocation, adversarial/debate verification, multi-party bargainingCombinatorial message growth, deadlock or infinite back-and-forth without a forcing function
Blackboard / shared-memoryAgents independently read and write to a common state store, triggered by what’s currently postedLoosely coupled tasks where the set of contributing agents is dynamic or unknown in advanceRace conditions on shared state, agents talking past each other, stale reads
Sequential pipelineFixed linear chain, each agent’s output feeds the next agent’s input directlyLinear content workflows (draft → check → edit) with no need for branching decisionsA single stuck or wrong step blocks everything downstream, with no way to route around it

Peer-to-peer topologies also carry a structural cost. A fully connected mesh of nn agents requires up to (n2)=n(n−1)2\binom{n}{2} = \dfrac{n(n-1)}{2} communication channels:

channels(n)=n(n−1)2\text{channels}(n) = \frac{n(n-1)}{2}

Message volume therefore grows quadratically, O(n2)O(n^2), as agents are added. A hierarchical hub-and-spoke design keeps this linear, O(n)O(n), since every agent only ever talks to the orchestrator — one of the main practical reasons orchestrator-worker dominates in production systems even though peer-to-peer is more flexible in theory.

Agent countFull-mesh channelsHub-and-spoke channels
333
5105
104510
2019020

Practical Limits on Agent Count

  • Beyond roughly six to eight agents in active coordination, most systems become difficult for a human to supervise in real time — the number of possible interaction paths grows faster than any one person’s ability to track them
  • Diminishing returns set in well before that ceiling: each additional agent adds coordination overhead and another multiplicative term to the reliability product shown below, so the marginal benefit of a seventh specialist is rarely as large as the first three
  • Large agent counts are more tractable inside a blackboard or hierarchical pattern, where an individual agent only needs to understand the shared state or its immediate manager, than in peer-to-peer, where every agent potentially needs to reason about every other agent’s intent
  • Systems that appear to need dozens of agents are often better modeled as one agent with many tools, or a hierarchy of a handful of team-lead agents each managing a small set of specialists, rather than one flat pool of peers

Negotiation and Consensus Mechanisms

Where agents must resolve competing goals rather than simply divide labor, classic distributed-AI protocols reappear inside LLM systems:

  • Contract Net Protocol — the orchestrator announces a task, available agents “bid” with an estimate of cost or confidence, and the task is awarded to the best bidder
  • Auction-based allocation — agents bid resources (compute budget, time, priority) against each other, common in simulation and scheduling use cases
  • Voting / majority consensus — several agents independently attempt the same sub-task and the majority (or highest-confidence) answer wins, trading extra compute for higher accuracy
  • Debate — one agent argues for an answer, another argues against or proposes an alternative, and a judge agent decides based on the exchange, surfacing reasoning errors a single pass would miss
  • Weighted arbitration — the orchestrator assigns different agents different trust weights based on past reliability, rather than treating every vote equally
  • Escalation to a human — the fallback consensus mechanism whenever automated agreement can’t be reached within a fixed number of negotiation rounds

Human-in-the-Loop as a Fourth Pattern

  • Many production systems add a human as an explicit participant rather than a pure fallback — approving a plan before execution, reviewing a diff before merge, or confirming an irreversible action before it fires
  • Where the human checkpoint sits matters: gating before an expensive or irreversible step catches problems cheaply, while gating only at the very end means all the upstream work has already been spent by the time a human sees it
  • Human-in-the-loop trades latency and throughput for safety, and is typically reserved for the highest-stakes steps in a pipeline rather than applied uniformly across every single agent handoff

Reliability, Error Propagation, and Observability

Multi-agent systems introduce a reliability problem that single-agent systems don’t have: errors compound across handoffs. If a pipeline has nn agents in sequence and each succeeds independently with probability pip_i, the probability the full pipeline succeeds is the product of the individual probabilities:

P(success)=∏i=1npiP(\text{success}) = \prod_{i=1}^{n} p_i

With five agents each 90% reliable, end-to-end success is only 0.95≈0.590.9^5 \approx 0.59 — a system that looks robust agent-by-agent can fail more than 40% of the time in aggregate. Raising per-agent reliability from 90% to 98% has an outsized effect at the same chain length: the same five-agent pipeline jumps from roughly 59% to roughly 90% end-to-end success, which is why teams invest heavily in tightening individual agent prompts rather than only adding more downstream checks.

Per-agent reliability3-agent chain5-agent chain8-agent chain
99%97.0%95.1%92.3%
95%85.7%77.4%66.3%
90%72.9%59.0%43.0%

This is the mathematical core of why cascading errors is the top pitfall in practice, and it drives most mitigation strategies used in real systems:

  • Dedicated verifier agents check a prior agent’s output against explicit constraints before it’s allowed to propagate downstream, rather than trusting it implicitly
  • Self-consistency / majority voting runs the same sub-task multiple times and keeps the majority answer, raising effective per-step reliability at the cost of extra compute
  • Checkpointing and retries with backoff mean a transient failure — a flaky tool call, a rate-limited API — doesn’t force the entire pipeline to restart from scratch
  • Circuit breakers halt the pipeline and escalate to a human once an agent’s confidence drops below a threshold or a retry budget is exhausted, instead of letting errors silently compound
  • Structured tracing logs every message, tool call, and intermediate state so a failure can be attributed to the specific agent and step that introduced it
  • Shorter chains where possible — every additional agent in a sequential chain is another multiplicative term in the success-probability product, so collapsing two weak steps into one strong one often beats adding a third checker

Evaluating and Testing Multi-Agent Systems

A correct final answer can mask an unreliable process that only happened to get there by luck, so testing a MAS requires more than checking the last message in the transcript.

Per-Agent Evaluation

  • Test each agent’s role in isolation first, feeding it realistic inputs and checking its output against explicit success criteria — the same way a unit test checks one function rather than an entire program
  • Establish a baseline success rate per agent before wiring agents together; an 80%-reliable agent buried inside a five-agent chain is a debugging nightmare months later if nobody measured it going in
  • Track the specific failure modes each agent exhibits — a wrong tool call, a malformed output schema, a missed edge case — rather than collapsing everything into a single pass/fail number

End-to-End Evaluation

  • Integration tests should run the full pipeline against realistic tasks and check the final output, but also assert on the path taken — which agents fired, how many revision loops occurred, whether termination happened cleanly
  • Regression suites built from real production failures (“replay” tests) catch cases where a prompt change to one agent silently breaks a downstream agent’s assumptions
  • Useful production metrics include end-to-end task success rate, average number of agent-to-agent handoffs per task, cost per completed task, and latency from request to final answer
  • Cost and latency should be tracked per agent as well as in aggregate, since a single expensive or slow agent inside an otherwise cheap pipeline is easy to miss when only totals are visible

Adversarial and Edge-Case Testing

  • Feed conflicting instructions to two agents that are supposed to agree, and confirm the system escalates or resolves the conflict rather than silently picking one arbitrarily
  • Test with ambiguous or malicious inputs at the system’s entry point to verify prompt-injection defenses hold at every agent boundary the input can reach, not just the first one
  • Force non-convergence deliberately (an agent that never agrees with another) to confirm the termination logic actually halts the loop rather than running until a timeout or budget cap is hit by accident

Cost and Latency Budgeting

Every agent invocation is a separate model call with its own token cost and latency, so a MAS’s total cost and speed are direct functions of its topology, not just its task’s underlying difficulty.

Cost Scales With Calls, Not With Task Complexity Alone

  • Total cost is roughly the sum, across every agent invocation, of (input tokens + output tokens) × that model’s per-token price — a five-agent pipeline that runs two revision loops can trigger far more calls than the topology diagram suggests
  • Routing cheap, fast models to simple sub-tasks (classification, formatting, extraction) and reserving an expensive frontier model for the one step that genuinely needs deep reasoning is the single highest-leverage cost lever in a MAS
  • Retries, revision loops, and majority-voting consensus mechanisms all multiply call count directly — a voting scheme that samples five candidates and picks the majority costs roughly five times what one confident call would

Latency Depends on Topology, Not Just Agent Count

  • A sequential pipeline’s latency is the sum of every agent’s latency plus handoff overhead between each step — adding an agent to a sequential chain always makes the pipeline slower
  • A parallel fan-out followed by a join step has latency closer to the slowest individual agent rather than the sum, since independent agents run concurrently
  • Mixed topologies benefit from identifying which subtasks are truly independent versus which have a genuine dependency — misclassifying a parallelizable step as sequential is a common and entirely avoidable source of extra latency

Comparison

ConceptNumber of Reasoning LociCoordinationKey Distinction from MAS
Single Intelligent Agent with multiple toolsOneNone needed — one control loop decides everythingDivision of labor happens across tools, not across independently reasoning agents; no inter-agent negotiation or role specialization
Expert SystemsOne (rule engine)None — deterministic rule firingNo autonomy, learning, or negotiation; behavior is fully specified by a static rule/fact base rather than emergent from agent interaction
Ensemble methods (bagging/boosting in classical ML)Many modelsStatistical aggregation (voting/averaging)Ensembles combine predictions from models trained on the same task; MAS agents perform different sub-tasks and pass structured state, not just vote on one output
Plain Function Calling (Tool Use) pipelineOneSequential tool invocations by one controllerThe “agents” are stateless tool functions, not autonomous reasoners with their own goals, memory, or decision-making — no agent-to-agent messaging exists
Multi-agent Reinforcement Learning (MARL)Many, but trained not promptedLearned policies shaped by reward signals over many training episodesMARL agents acquire behavior through gradient-based training in a simulated environment; LLM-based MAS agents use pretrained models directed at inference time purely through prompted roles and tool access
  • The line between a heavily-tooled single agent and a two-agent MAS is genuinely blurry in practice — the question worth asking is whether there are two independent reasoning processes with their own context and objective, or one process simply making several tool calls in a row
  • MARL and LLM-based MAS are sometimes combined in robotics and simulation work: a MARL-trained policy can act as one specialized agent’s decision layer inside an otherwise prompt-driven multi-agent system
  • None of the concepts in this table are mutually exclusive with a MAS — a single specialist agent inside a larger multi-agent pipeline might itself be implemented as an ensemble, or wrap a rule-based expert system for a narrow, deterministic sub-step

Real-World Use Cases

  • Multi-step research assistants that spin up parallel sub-agents to investigate different facets of a question, then a synthesis agent merges findings into a single cited report
  • Coding agents that split planning, code generation, test execution, and code review across distinct sub-agent roles so a reviewer agent catches issues before code is merged
  • Customer support triage systems that route incoming tickets to specialist agents (billing, technical, account security) after a classifier agent determines category and urgency
  • Simulation and game environments where dozens of NPC agents each pursue local goals, producing emergent group behavior — economies, social dynamics, crowd movement — that no single script authored directly
  • Scientific discovery pipelines pairing a hypothesis-generation agent with an experiment-design agent and a results-analysis agent, iterating across a discovery loop
  • Financial and trading systems where cooperating bots handle signal generation, risk assessment, and execution as separate roles with independent checks
  • Supply chain and logistics negotiation between buyer, seller, and carrier agents that bid and counter-bid over price and delivery windows using contract-net-style protocols
  • Content production pipelines where a drafting agent, a fact-checking agent, and an editing agent operate in sequence before a piece is published
  • Enterprise workflow automation where agents cooperate across otherwise siloed systems — CRM, ticketing, email — to complete a task that spans tools no single API call covers
  • Robotics swarms where physically distributed agents coordinate via a shared blackboard to divide territory or tasks without a single point of control
  • Document processing systems that route extraction, classification, and validation of incoming forms or contracts to separate agents, each specialized to one stage of the pipeline
  • Legal and compliance review pipelines where one agent extracts clauses, another flags risk against a policy checklist, and a third drafts a plain-language summary for a human reviewer
  • Personal AI assistants that delegate calendar management, email triage, and task tracking to specialized sub-agents coordinated by a single front-facing assistant the user actually talks to

Common Pitfalls

  • Cascading errors. One agent’s mistake propagates downstream and compounds — as the reliability math above shows, a chain of individually-solid agents can still produce an unreliable system overall
  • Over-engineering. Standing up an orchestrator plus three specialist agents for a task a single well-prompted agent would handle in one pass adds latency, cost, and failure surface for no real gain
  • Communication overhead. Every message passed between agents costs tokens and latency; peer-to-peer topologies in particular can spend more compute coordinating than actually solving the task
  • Context desynchronization. When state isn’t propagated carefully, one agent can act on stale or incomplete information another agent already superseded, producing contradictory outputs
  • Undefined termination conditions. Without an explicit stop signal, agents can loop indefinitely re-negotiating a decision or repeatedly “reviewing” and “revising” the same output
  • Cost multiplication. Every agent invocation is a separate model call; a five-agent pipeline that runs several iterations can cost an order of magnitude more than a single well-designed prompt
  • Prompt injection across the message bus. If one agent ingests untrusted content — a web page, a document — a malicious instruction embedded in it can propagate through shared memory and manipulate downstream agents that never touched the original untrusted source
  • Debugging difficulty. Non-determinism compounds across agents, so reproducing a failure means reproducing the exact sequence of LLM outputs across every agent involved
  • Duplicated or wasted work. Poorly scoped roles lead to agents redoing work another agent already completed, especially in loosely coordinated blackboard systems where visibility into “what’s already been done” is incomplete
  • Unclear ownership and role overlap. When two agents’ responsibilities aren’t cleanly separated, they can produce conflicting outputs with no clear rule for which one wins — a design failure, not a model failure
  • Treating more agents as inherently more capable. Adding agents without a clear reason each one exists tends to add failure surface faster than it adds capability — the coordination pattern matters more than the headcount
  • Skipping observability until something breaks. Teams that don’t instrument message passing and tool calls from day one find themselves unable to diagnose failures once the system is complex enough to actually need debugging
  • Misclassifying sequential work as parallel, or vice versa. Running dependent subtasks concurrently produces race conditions on shared state, while running truly independent subtasks in sequence wastes latency for no reliability benefit

Example

A team builds a multi-agent coding system to handle feature requests end-to-end. A user submits: “Add rate limiting to the public API.” The orchestrator receives the request and dispatches it to a planner agent, which breaks it into three subtasks: identify the API entry points, implement a token-bucket limiter, and add tests.

The orchestrator hands the first two subtasks to a code agent in sequence, writing a diff that adds middleware and modifies the route handlers. Before that diff is accepted, a review agent — running independently with only the diff and the original request as context, not the code agent’s reasoning trace — checks it against the requirement and flags a real problem: the limiter uses per-process memory instead of a shared store, so it won’t work correctly once the API is running behind a load balancer with multiple instances.

That flag goes back to the orchestrator, which routes a revision request to the code agent along with the review agent’s specific objection. The code agent rewrites the limiter to use the existing Redis client already present in the codebase. A test agent then runs the project’s test suite plus a new test it writes for the rate-limit boundary condition — the request exactly at the limit — confirms everything passes, and reports back. Only then does the orchestrator mark the task complete and return the final diff to the user, along with a short summary of what changed and why the first implementation attempt was revised.

The case illustrates the core value proposition concretely: the review agent caught a distributed-systems bug that the code agent, focused purely on “make the diff work,” had no reason to think about — its role and context were scoped narrowly to writing code, not to reasoning about deployment topology. A single monolithic agent handling planning, coding, and reviewing in one continuous context might have caught the same issue, or might not have, since the review step would have been just another paragraph of a much longer, more diluted prompt rather than a dedicated pass with its own focused attention.

The multi-agent structure made the verification step structurally guaranteed rather than incidental, and it left a clean audit trail: anyone inspecting the run later can see exactly which agent proposed what, which agent objected, and why the final diff looks the way it does — a byproduct of the architecture that a single continuous transcript would have made far harder to reconstruct.

Dig deeper