Multi-Agent Systems
Several LLM agents with separate contexts and roles working together - when that helps, how to structure it, and what it costs.
What is it?
A multi-agent system splits work across several agents, each with its own system prompt, its own tools, and - most importantly - its own context window. Typically a lead agent (the orchestrator or supervisor) breaks a task down and delegates pieces to subagents, which work independently and report back condensed results.
Why would several agents beat one? Three real reasons:
- Context isolation: a subagent can read 50 pages to answer one sub-question and return a 150-word summary. The lead agent's context stays small and focused instead of filling with raw material it will never need again.
- Parallelism: independent sub-questions (research five competitors, check four regions) can run at the same time, cutting wall-clock time for breadth-first tasks.
- Specialisation and least privilege: each agent gets only the prompt and tools for its job. A research subagent can have read-only search; only a separate, approval-gated agent can send emails.
And real costs - often larger than people expect:
- Tokens multiply: every agent has its own system prompt, tool definitions and loop, and subagents often explore more than strictly needed. Multi-agent runs commonly use several times more tokens than a single agent on the same task.
- Coordination is hard: the lead must write clear, self-contained task descriptions (objective, scope, output format, what not to do). Vague delegation leads to duplicated work, gaps, or subagents solving the wrong problem.
- Information loss: summaries drop details; subagents cannot see each other's findings unless you design shared state.
- Harder debugging and evaluation: failures spread across several transcripts; you need tracing that links parent and child runs.
- Conflicts: parallel agents that write (edit the same files, update the same record) can collide.
Common structures:
- Orchestrator with subagents (supervisor): the lead plans, delegates via a 'delegate' tool, and synthesises. The most widely used and easiest to reason about. Subagents are effectively tools that happen to be agents.
- Handoffs: a triage agent transfers the conversation to a specialist agent (billing, technical) that then owns it. Good for customer-facing flows with clear domains.
- Pipeline: agents in a fixed order (researcher, then writer, then editor) - really a workflow whose steps are agents.
- Critic / debate: one agent proposes, another challenges; a judge or rule decides. Useful for review, but costly.
Practical rules: start with a single agent and only split when you hit a concrete limit (context overflow, latency on parallelisable work, tool permissions). Give subagents tight, explicit tasks and a required output format. Return condensed results, not transcripts. Cap steps per agent and the number of subagents. Keep writes in one place (subagents read and propose; one agent or a human applies changes). Trace everything with a shared run id.
Multi-agent setups tend to shine on breadth-first tasks - research across many independent sources, broad surveys, large-scale review - and to struggle on tightly coupled tasks where every part depends on every other part, such as most coding changes, where a single agent with full context usually does better.
Explain like I'm 10
A newsroom. The editor (orchestrator) assigns three reporters (subagents) to cover different angles. Each reporter spends a day on interviews and notes - that is their private context - and files a 300-word piece. The editor never reads the raw notes, only the pieces, and writes the front page. It is faster than one reporter covering everything, but you pay three salaries, and if the editor's assignments are vague, two reporters cover the same story while nobody covers the important one.
Examples
Context isolation and its token price (runnable, toy numbers)
const words = (s) => s.split(" ").length;
const filler = (n) => Array(n).fill("detail").join(" ");
const OVERHEAD = 500; // system prompt + tool definitions per agent, in "tokens"
const corpus = {
pricing: ["Pricing: Pro plan is 20 per user per month. " + filler(300),
"Billing FAQ: annual billing gives two months free. " + filler(250)],
security: ["Security: data is encrypted at rest and in transit. " + filler(400),
"Compliance: audited yearly by an independent firm. " + filler(200)],
support: ["Support: 24/7 chat for the Pro plan. " + filler(300)],
};
// SUBAGENT: fresh context, reads raw docs, returns a short summary.
function subagent(topic) {
const context = OVERHEAD + corpus[topic].reduce((n, d) => n + words(d), 0);
const summary = corpus[topic].map((d) => d.split(". ")[0]).join(". ") + ".";
return { topic, summary, context };
}
// ORCHESTRATOR: delegates (in parallel in production), sees only summaries.
const results = Object.keys(corpus).map(subagent);
const leadContext = OVERHEAD + words("Compare our plan for a buyer") + results.reduce((n, r) => n + words(r.summary), 0);
results.forEach((r) => console.log("subagent " + r.topic.padEnd(8) + " context=" + r.context + " -> " + r.summary));
console.log("lead agent context =", leadContext);
// SINGLE AGENT: one context holds everything.
const single = OVERHEAD + Object.values(corpus).flat().reduce((n, d) => n + words(d), 0);
const multiTotal = leadContext + results.reduce((n, r) => n + r.context, 0);
const peak = Math.max(leadContext, ...results.map((r) => r.context));
console.log("single agent : peak context", single, "| total tokens", single);
console.log("multi-agent : peak context", peak, "| total tokens", multiTotal);
console.log("=> smaller, cleaner contexts, but", (multiTotal / single).toFixed(2) + "x the tokens");Each subagent's context holds only its slice, and the lead sees roughly 60 words of summaries instead of 1,500 words of raw text. But overhead is paid per agent, so total tokens rise. Real agents loop several times per task, so the multiplier in practice is usually larger than this toy shows.
Handoffs with least-privilege tool sets (runnable)
const agents = {
triage: { tools: [], handles: "classifies and hands off" },
billing: { tools: ["lookup_invoice", "issue_refund"], handles: "payments and refunds" },
technical: { tools: ["search_docs", "check_status"], handles: "bugs and outages" },
};
function triage(message) { // fake LLM classification
const m = message.toLowerCase();
if (m.includes("refund") || m.includes("invoice")) return "billing";
if (m.includes("error") || m.includes("down")) return "technical";
return "triage";
}
function callTool(agentName, tool) { // enforced in CODE, not just in the prompt
if (!agents[agentName].tools.includes(tool)) return "DENIED: " + agentName + " may not use " + tool;
return "ok: " + agentName + " ran " + tool;
}
for (const msg of ["I need a refund for invoice 881", "The dashboard is down with error 503"]) {
const owner = triage(msg);
console.log("'" + msg + "' -> handed off to " + owner + " (" + agents[owner].handles + ")");
}
console.log(callTool("billing", "issue_refund"));
console.log(callTool("technical", "issue_refund")); // a confused or hijacked agent is stoppedEach specialist gets only the tools for its domain, and the check lives in code. If the technical agent is tricked by text in a bug report into trying a refund, the call is refused regardless of what the model wants.
Orchestrator with parallel research subagents (Python)
import anthropic
from concurrent.futures import ThreadPoolExecutor
from rag_tools import SEARCH_TOOL, search_docs # the search tool from the Agentic RAG lesson
client = anthropic.Anthropic()
MODEL = "claude-opus-5-5"
def run_agent(system, task, tools, run_tool, max_steps=8):
"""The same manual agent loop as before, packaged as a function."""
messages = [{"role": "user", "content": task}]
for _ in range(max_steps):
resp = client.messages.create(model=MODEL, max_tokens=16000, system=system,
tools=tools, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
return "".join(b.text for b in resp.content if b.type == "text")
results = []
for b in resp.content:
if b.type == "tool_use":
try:
results.append({"type": "tool_result", "tool_use_id": b.id,
"content": run_tool(b.name, b.input)})
except Exception as e:
results.append({"type": "tool_result", "tool_use_id": b.id,
"content": f"Error: {e}", "is_error": True})
messages.append({"role": "user", "content": results})
return "[stopped at step limit]"
# ---- Subagent: read-only researcher with its own fresh context ----
RESEARCHER = ("You are a research subagent. Answer the assigned question using search_docs. "
"Return at most 150 words of findings, each with its [source id]. No preamble.")
def research(question: str) -> str:
return run_agent(RESEARCHER, question, [SEARCH_TOOL], lambda name, args: search_docs(**args))
# ---- Orchestrator: plans, delegates in parallel, synthesises ----
DELEGATE_TOOL = {
"name": "delegate_research",
"description": "Give ONE self-contained research question to a subagent that searches the "
"handbook and returns a short summary with source ids. The subagent cannot see "
"this conversation, so include all needed context in the question. For independent "
"questions, call this tool several times in the same turn; they run in parallel.",
"input_schema": {"type": "object",
"properties": {"question": {"type": "string"}},
"required": ["question"]},
}
LEAD = ("You are the lead researcher. Break the user's request into 2-4 independent questions, "
"delegate them, then write the final answer citing the subagents' source ids. "
"Do not delegate more than 4 questions in total.")
pool = ThreadPoolExecutor(max_workers=4)
def orchestrate(task: str, max_steps: int = 6) -> str:
messages = [{"role": "user", "content": task}]
for _ in range(max_steps):
resp = client.messages.create(model=MODEL, max_tokens=16000, system=LEAD,
tools=[DELEGATE_TOOL], messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
return "".join(b.text for b in resp.content if b.type == "text")
calls = [b for b in resp.content if b.type == "tool_use"]
answers = list(pool.map(lambda b: research(b.input["question"]), calls)) # parallel
messages.append({"role": "user", "content": [
{"type": "tool_result", "tool_use_id": b.id, "content": a} for b, a in zip(calls, answers)
]})
return "[orchestrator stopped at step limit]"
print(orchestrate("Compare the travel and expense rules for contractors vs employees."))Subagents are just agent loops exposed to the lead as a tool. Each one starts with an empty context, does its own searching, and returns only a summary. pool.map preserves order, so each answer is paired with the right tool_use id. Every subagent's tokens add to the bill - log usage per agent.
How it works
Mechanically, a multi-agent system is nested agent loops. The lead agent's loop treats 'run a subagent' as a tool call; that tool's implementation starts a new loop with a fresh messages list, a different system prompt and a restricted tool set, runs it to completion, and returns its final text as the tool_result. Parallel tool calls from the lead become concurrent subagent runs.
Information flows only through what you pass: the task description going down, the summary coming up, plus any shared store (a file, a database, a notes tool) you deliberately provide. This is both the benefit (isolation) and the risk (lost context). Designing the delegation message - objective, constraints, expected output format, effort limit - is the most important part of the system.
Operationally you need: per-agent step and token limits, a global budget, a trace id that ties child runs to the parent, and evaluation at both levels (did each subagent answer its question, did the final answer solve the task).
user task
|
+-------------+
| LEAD agent | small context:
| plan/merge | task + summaries
+------+------+
delegate (parallel tool calls)
/ | \
+---------+ +---------+ +---------+
| sub A | | sub B | | sub C | each: own
| search | | search | | search | context,
| loop | | loop | | loop | own tools,
+----+----+ +----+----+ +----+----+ step cap
\ summary | summary / summary
+-----------+-----------+
v
final answer + sourcesWhy does it exist?
Single agents hit limits on large, broad tasks: their context fills with raw material, sequential exploration is slow, and one agent holding every tool violates least privilege. Splitting work across agents with isolated contexts and narrow permissions addresses those limits - at the price of more tokens and more coordination.
When to use it
Use multiple agents for breadth-first work that splits into independent parts (multi-source research, surveys, reviewing many documents), when a single agent's context overflows with material it only needs briefly, when parallelism meaningfully reduces latency, or when distinct domains need different tools and permissions (handoffs between support specialists).
When not to use it
Do not use multi-agent designs for tightly coupled tasks where every step depends on shared, detailed context (most code changes, a single coherent document), for simple tasks a single agent or workflow already handles, or when token cost matters more than speed. Many 'multi-agent' demos are better as one agent with good tools, or as a workflow.
Common mistakes
Going multi-agent first, before a single agent has been shown to hit a real limit.
Vague delegation ('research pricing') that causes duplicated work or off-target subagents.
Returning full subagent transcripts to the lead, destroying the context-isolation benefit.
No per-agent step caps or global budget, so a fan-out multiplies runaway costs.
Several agents writing to the same files or records in parallel without coordination.
Relying on prompts alone to keep agents within their roles instead of enforcing tool permissions in code.
No tracing that links child runs to the parent, making failures impossible to locate.
Practice exercises
- Easy:
List three tasks from your work and decide for each whether one agent, a workflow, or a multi-agent design fits best. Justify using context size, parallelism and coupling.
- Easy:
In the context-isolation demo, double the size of every document and report how peak context and total tokens change for both designs.
- Medium:
Write a delegation template for subagents with fields for objective, scope, sources to use, output format and a step budget. Rewrite a vague task with it.
- Medium:
Add a third agent 'sales' with a 'create_quote' tool to the handoff demo and a test that the billing agent cannot call it.
- Hard:
Run the Python orchestrator and a single-agent version (search_docs directly on the lead) on the same 5 questions. Record tokens, wall-clock time and answer quality. Write up when multi-agent was worth it.
Interview questions
What are the main benefits of a multi-agent system?
Context isolation (subagents absorb raw material and return condensed results), parallelism on independent subtasks, and specialisation with least-privilege tool sets per agent.
What are the main costs and risks?
Much higher token usage, coordination overhead and the need for precise delegation, information loss through summaries, harder debugging and evaluation across multiple transcripts, and conflicts when parallel agents write to shared state.
How do you implement a subagent?
As a tool for the lead agent whose implementation runs a separate agent loop with a fresh message history, its own system prompt and restricted tools, and a step limit, returning only its final condensed answer as the tool_result.
Orchestrator-subagents vs handoffs?
With an orchestrator, a lead agent stays in control, delegates subtasks, and synthesises results. With handoffs, control of the conversation transfers to a specialist agent that then owns it, typically after a triage step.
When is a single agent better than multiple agents?
When the task is tightly coupled and needs shared detailed context (for example most coding changes), when it is small, or when cost matters more than latency. Splitting such tasks loses context and adds coordination errors.
How do you keep a multi-agent system safe and affordable?
Enforce tool permissions per agent in code, gate side effects behind approval, cap steps per agent and subagent count, set a global token budget, keep writes centralised, and trace every run with a shared id so costs and failures can be attributed.