Run a multi-agent system for a week. Then look at the bill.
You'll expect the cost to be split across your specialist agents — the one writing code, the one searching the web, the one analyzing data. But when you break it down, something unexpected jumps out: the coordinator — the agent whose only job is routing tasks and managing context — costs more than all the specialists combined.
This is the orchestration tax. And it's silently eating your budget.
What the Orchestration Tax Actually Looks Like
In a typical 4-agent setup (coordinator + 3 specialists), here's where the tokens go:
- Coordinator: 60-90% of total spend. Context summaries, task routing, result aggregation, state management.
- Specialist agents: 10-40% combined. The actual work — code generation, research, analysis.
The coordinator doesn't produce anything. It coordinates. Every time it routes a task, it needs the full context of what's happened so far. Every time a specialist returns a result, the coordinator re-ingests it, decides what to do next, and passes context to the next agent. That context grows with every step.
By the third round-trip, the coordinator is processing more tokens than any specialist ever will. By the tenth, it's processing more tokens than all specialists across the entire run.
Why This Happens
1. Context Duplication
Every agent call includes context. When a coordinator hands off to a specialist, it includes the task description, relevant history, and constraints. When the specialist returns, the coordinator re-reads all of that plus the new output. Then it builds a new context package for the next specialist.
The same information gets serialized, transmitted, and processed multiple times. A 2,000-token task description might get included in 8 different API calls throughout a single workflow.
2. State Reconstruction
Agents are stateless by default. Each API call starts from scratch. The coordinator has to reconstruct the full state of the workflow every time it wakes up:
- What tasks have been completed?
- What's still pending?
- What did previous specialists find?
- Are there any errors to handle?
- What's the overall goal again?
This state reconstruction is pure overhead. It produces no output. It exists solely because the system can't remember what happened 30 seconds ago.
3. Exploding Context Windows
As a multi-step workflow progresses, the coordinator's context grows linearly (at best) or quadratically (at worst). Each round adds results from the previous step plus the new instructions for the next step.
A 5-step workflow that starts at 3,000 tokens can easily reach 40,000+ by step 5. And you're paying for every token — input and output.
How to Measure Your Orchestration Tax
Before you can fix it, you need to see it. Here's what to track:
Per-Agent Token Accounting
Log input and output tokens for every API call, tagged by agent role. After a week, calculate:
orchestration_tax = coordinator_tokens / total_tokens * 100
If this number is above 60%, you have a problem. If it's above 80%, you're paying more for coordination than work.
Context Growth Rate
Track the coordinator's input token count across sequential calls within a single workflow. Plot it. If it's growing linearly, you're doing okay. If it's accelerating, context is compounding and you're heading for a wall.
Effective Work Ratio
effective_work_ratio = specialist_output_tokens / total_tokens
This tells you what percentage of your spend actually produced something. In healthy systems, this should be above 30%. In systems with a severe orchestration tax, it can drop below 10%.
5 Ways to Cut the Orchestration Tax
1. Compress Context Between Rounds
Don't pass raw results from specialists back to the coordinator. Summarize them first. A specialist that returns 3,000 tokens of analysis can be compressed to 300 tokens of key findings before the coordinator sees it.
This single change can cut orchestration cost by 50%+.
2. Use Structured State Instead of Natural Language
Instead of the coordinator maintaining a natural-language narrative of what's happened, use a structured state object:
{
"workflow": "data_analysis",
"step": 3,
"completed": ["fetch", "clean"],
"pending": ["analyze", "report"],
"key_findings": ["3 outliers in Q2", "missing data in col_F"],
"errors": []
}
This costs ~100 tokens instead of 2,000+ for a narrative summary. The coordinator gets the same information in 5% of the tokens.
3. Direct Agent-to-Agent Handoffs
Not every task needs to route through the coordinator. If Agent A's output is always Agent B's input, let them communicate directly. The coordinator only needs to know the handoff happened, not see every token that was exchanged.
4. Session-Based Context Instead of Stateless Calls
Use persistent sessions where agents maintain state across calls. This eliminates the state reconstruction overhead entirely. The coordinator doesn't need to re-explain the workflow because the session remembers it.
This is the single biggest architectural decision for cost control. Stateful > stateless for multi-step workflows, every time.
5. Tiered Model Selection
Your coordinator doesn't need your most expensive model. Routing decisions and state management don't require the same intelligence as code generation or analysis. Use a cheaper, faster model for coordination and save the expensive model for the specialists that need it.
# Example model allocation
coordinator: claude-haiku (fast, cheap, good enough for routing)
code_agent: claude-opus (needs deep reasoning)
research: claude-sonnet (good balance)
summarizer: claude-haiku (compression is simple)
This can cut coordination costs by 80%+ without any reduction in output quality.
The Numbers After Optimization
After applying these techniques to a 4-agent fleet running 24/7:
- Before: Coordinator = 87% of total spend. Effective work ratio = 8%.
- After: Coordinator = 31% of total spend. Effective work ratio = 44%.
- Total cost reduction: 64%.
Same output quality. Same agent capabilities. Just less money burned on the invisible middleware tax.
The Deeper Problem
The orchestration tax isn't just a cost problem. It's a design signal.
If your coordinator is consuming 90% of your budget, your system is spending 90% of its intelligence on thinking about thinking. That's not just wasteful — it's a sign that the architecture needs rethinking.
The best multi-agent systems minimize coordination. Each agent knows enough to do its job without constant supervision. The coordinator exists to handle exceptions, not to micromanage every step.
This is the same lesson from human management: if a manager spends 90% of their time telling people what to do, the problem isn't the team — it's the manager.
Don't optimize the coordinator. Reduce the need for coordination.
Want Us to Find Your Orchestration Tax?
We audit multi-agent systems for cost leaks, context drift, and architectural inefficiencies. We've done this across production fleets — from 3-agent setups to 12-agent orchestration layers.
Our Agent Operations Audit includes a full token accounting breakdown, cost allocation analysis, and specific recommendations to cut spend without cutting capability.