When Multi-Agent Is Worth It (Usually It Isn't)
The senior-engineer take interviews reward: multi-agent adds latency, cost, and compounding error rates, and most systems that ship as five agents should have shipped as one good agent. Learn the three legitimate justifications, the math of compounding failure, and how to run the baseline comparison that keeps you honest.
Here is the uncomfortable truth this module exists to teach: most multi-agent systems are worse than the single agent they replaced would have been. Every agent boundary you add introduces a handoff that can lose information, a scheduling step that adds latency, duplicated context that multiplies token cost, and an independent failure point. The architecture diagrams look like org charts, and that's seductive — "a researcher, a writer, an editor, just like a real team!" — but models aren't people. One model with good tools, a clean context, and a decent loop doesn't need meetings.
The error math is brutal and worth quoting in interviews. If a task flows through five sequential stages and each stage is 90% reliable, end-to-end reliability is 0.9 to the fifth power — about 59%. Your five nines-of-effort agents ship a coin flip. Worse, errors compound in content, not just probability: stage 2 doesn't merely fail sometimes, it feeds subtly-wrong output to stage 3, which builds confidently on the error. A single agent with the same 90% reliability on the whole task is... 90%. Chaining only wins when each stage is dramatically more reliable on its narrow slice than one agent is on the whole — which you must measure, not assume.
| Justified when… | Why it actually helps | NOT justified when… |
|---|---|---|
| Context isolation matters — each worker needs a clean window | A searcher with a 5k-token context containing exactly its subtask outperforms one agent dragging 100k tokens of accumulated transcripts, dead ends, and tool dumps — the mess actively degrades attention | "Separation of concerns" as an aesthetic — code modules give you that without runtime handoffs |
| True parallelism — subtasks are independent and I/O-bound | Three searchers finish in one wall-clock unit instead of three; only works when subtasks don't need each other's results | The subtasks are sequential anyway — you've added coordination and kept the latency |
| Distinct tools or permissions per role — the deploy agent has prod credentials, the researcher has read-only web | Smaller blast radius per agent; a compromised or confused researcher cannot touch prod (Module 7 will sharpen this) | All "agents" share the same tool set and permissions — that's one agent with extra steps |
Why does the clean-window worker win? Recall Module 4: model quality degrades as the context fills with low-relevance material — attention gets diluted, instructions in the middle get lost, and old errors get treated as established facts. A monolithic agent that has done nine searches carries every raw result and misstep into search ten. A fresh worker receives a brief: the subtask, the constraints, nothing else. Context isolation is the strongest technical argument for multi-agent because it attacks the actual failure mechanism, not the org chart. Notice the corollary: if your contexts aren't degrading — short tasks, small histories — this argument evaporates, and with it most of the case for splitting.
The baseline comparison — Lab 05's differentiator
The rule that keeps you honest: never ship a multi-agent system without benchmarking it against a single-agent baseline with the same tools. Same questions, same tools, same model; only the architecture differs. Measure three things — output quality (LLM-as-judge plus your own rubric), total cost (sum tokens across all agents and handoffs — people forget the orchestrator's tokens), and wall-clock latency. Signals that you should collapse to single-agent: quality is within noise of the baseline; most handoff payloads just restate prior state (the boundary adds no information); cost is a multiple of baseline; or your handoff log shows workers spending turns re-deriving context the monolith would simply have had.
# Colab cell 1 — pure Python; no API key needed. The two run_* functions
# are stubs standing in for your systems — Lab 05 wires in the real
# LangGraph pipeline and single-agent loop; the harness around them is
# exactly what you'd ship.
import json
import time
QUESTIONS = [ # the same set for both conditions
"Compare HNSW and IVF vector indexes.",
"When should you use hybrid search over dense retrieval?",
"What does a reranker buy you?",
]
def run_langgraph_system(q: str):
# stub: pretend the multi-agent pipeline ran; return (answer, usage)
return (f"[multi] answer to: {q}",
{"input_tokens": 9000, "output_tokens": 1200})
def run_single_agent_same_tools(q: str):
# stub: same task, one agent, same tools — cheaper context
return (f"[single] answer to: {q}",
{"input_tokens": 3000, "output_tokens": 900})
def run_condition(name: str, run_fn) -> list[dict]:
rows = []
for q in QUESTIONS:
t0 = time.time()
answer, usage = run_fn(q) # returns (text, token/cost totals)
rows.append({
"condition": name,
"question": q,
"answer": answer,
"latency_s": round(time.time() - t0, 1),
"input_tokens": usage["input_tokens"], # summed over ALL
"output_tokens": usage["output_tokens"], # agents + handoffs
})
return rows
multi = run_condition("multi_agent", run_langgraph_system)
single = run_condition("single_agent", run_single_agent_same_tools)
with open("comparison_raw.jsonl", "w") as f:
for row in multi + single:
f.write(json.dumps(row) + "\n")
for cond in (multi, single):
tok = sum(r["input_tokens"] + r["output_tokens"] for r in cond)
print(f"{cond[0]['condition']:12s} total tokens across all calls: {tok:,}")# Colab cell 2 — run cell 1 first (it defines multi, single). The judge
# call is stubbed; Lab 05 swaps in a real structured-output model call.
import random
JUDGE_PROMPT = """You are grading two research briefs answering:
{question}
Brief A:
{a}
Brief B:
{b}
Rubric: factual grounding and citations (40%), coverage of the
question (30%), clarity and structure (30%).
Score each brief 1-10 per rubric item, then declare a winner.
Return JSON: {{"a_scores": ..., "b_scores": ..., "winner": "A"|"B"|"tie"}}"""
def call_judge_model(prompt: str) -> dict:
# stub: Lab 05 makes this a real structured-output call —
# init_chat_model("anthropic:claude-sonnet-5") or "openai:gpt-5.5"
# returning JSON validated against the rubric schema. Canned here so
# the position-randomizing harness runs with no key.
return {"a_scores": [8, 8, 8], "b_scores": [7, 7, 7], "winner": "A"}
def judge(question: str, multi_ans: str, single_ans: str) -> dict:
# randomize position to cancel the judge's first-position bias
if random.random() < 0.5:
a, b, mapping = multi_ans, single_ans, {"A": "multi", "B": "single"}
else:
a, b, mapping = single_ans, multi_ans, {"A": "single", "B": "multi"}
verdict = call_judge_model(JUDGE_PROMPT.format(
question=question, a=a, b=b)) # structured output, JSON schema
verdict["winner"] = mapping.get(verdict["winner"], "tie")
return verdict
# judge the first question's two answers from cell 1:
print(judge(QUESTIONS[0], multi[0]["answer"], single[0]["answer"]))temperature/top_p/top_k entirely (Module 1), so "run it at temperature 0" isn't available; a narrow, structured rubric is what current models rely on for repeatable judging. Spot-check a sample of judgments by hand and report your own rubric scores alongside the judge's. If the single agent wins, say so in the README — an honest negative result is a stronger portfolio signal than a rigged win. The call_judge_model stub returns a canned verdict so the position-randomizing wrapper runs with no key; note the mapping step un-shuffles "A"/"B" back to "multi"/"single" so your bias fix doesn't scramble the results.Cost multiplication, worked
"Multi-agent costs more" is easy to say and easy to underestimate — the multiplier isn't the agent count, it's agent count × iterations per agent × handoff re-sends. Take a single agent averaging 8 loop iterations at roughly 3k tokens each: about 24k tokens for the task. Now the five-role research pipeline: planner runs 2 iterations, three parallel searchers run 4 each (12 total), the writer runs 3, the critic runs 2 for one revision cycle — 19 agent-iterations total, which sounds almost comparable to the single agent's 8. It isn't, because each of those iterations pays for context the single agent never re-sent: every searcher re-includes its brief and the tools it needs, the writer re-includes every finding, the critic re-includes the plan, the question, and the full draft. The realistic outcome on real workloads is 2–4× the single agent's token cost for the same task, not 19/8. This is exactly why the baseline harness above sums tokens across every call including the orchestrator's — a comparison that only counts "the interesting agents' " tokens will always make multi-agent look cheaper than it is.
Debugging multiplication
Cost isn't the only thing that scales with agent count — debugging surface area does too, and worse than linearly. A single agent has one loop to trace. A five-node pipeline has five nodes and four handoffs, and a wrong final answer could originate at any of those nine places — plus a tenth: the interaction between two of them, where each individually looks correct in isolation (Lesson 4's briefing-bug pattern). Failures also mask each other: a planner that silently drops a subtask doesn't look like a bug until three stages later, when the writer's draft is missing a section, and the instinct is to debug the writer — the actual defect is two hops upstream, in a component whose output looked fine because it was never checked against the original question, only consumed as-is by the next node. Multi-agent systems therefore have a non-negotiable observability tax on top of the token tax: every node and every handoff needs its own log entry (Lesson 4), because a system you can't fully instrument at N=5 is a system you can't debug at N=5, no matter how good any individual agent's prompt is.
Escalate on evidence, not vibes
The "Justified when…" table above tells you the three legitimate reasons; this is the protocol for finding out whether your system actually has one of them, instead of pattern-matching an architecture diagram to a vague feeling that "this task feels like it needs a team." Measure the single agent first, on the real workload, before designing a single node of a multi-agent replacement.
| Question to measure | How to measure it | What justifies escalating |
|---|---|---|
| Is context actually degrading? | Plot the single agent's error/quality rate against conversation length or tool-call count on real traces | Quality drops measurably past some context size — that's context isolation's evidence, not a hunch about 'long contexts are bad' |
| Are subtasks actually independent? | Profile the single agent's tool calls: do later calls depend on earlier results, or could they run with no shared state? | A real chunk of wall-clock time is spent on calls that don't depend on each other — that's true parallelism's evidence |
| Does blast radius actually matter here? | Ask what the worst-case action is if the model is confused or compromised, and who/what it can currently touch | The worst case is unacceptable and the tool set can be cleanly partitioned by role — that's the permissions justification's evidence |
| None of the above show up | — | Ship the single agent. Revisit only when new measurements say otherwise, not when a new pattern trends on social media |
Whiteboard drills
- ▸Default answer: one good agent. Multi-agent must earn its coordination cost with measurements.
- ▸Five sequential stages at 90% each ≈ 59% end-to-end — and errors compound in content, not just probability.
- ▸Three legitimate justifications: context isolation, true parallelism, distinct tools/permissions. "Separation of concerns" alone is hand-waving.
- ▸Clean 5k-token worker windows beat one 100k-token accumulated mess because context pollution degrades attention.
- ▸Always run a single-agent baseline: same tools, same model, same questions; compare quality, total cost (all agents' tokens), latency.
- ▸Collapse signals: quality within noise of baseline, handoffs that add no information, cost multiples, workers re-deriving context.