The Sol-to-Sol Trap: Why Multi-Agent V2 Defaults Kill Your ROI
The Sol-to-Sol Trap: Why Multi-Agent V2 Defaults Are Killing Your ROI
Most leaders believe moving to multi-agent systems will automatically save them money. They are wrong.
Without manual intervention, your Multi-Agent V2 setup will bleed budget on grunt work. When I built the first automation frameworks at Aniccai, I learned the hard way that vendor defaults are almost always the most expensive option.
Key Takeaways
- The Inheritance Trap: Sub-agents automatically inherit the expensive Sol model from their parent, leading to inflated token costs.
- The Hidden Schema: The
hide_spawn_agent_metadatasetting prevents the model from choosing cheaper models for routine tasks. - Context Bloat: Full-history forking forces every sub-agent to re-read the entire conversation, multiplying costs unnecessarily.
- The Pragmatic Fix: Exposing the schema and defining a model hierarchy (Sol, Terra, Luna) can cut operational costs by up to 80%.
The Promise of Delegation: Sol's Brain, Luna's Price Tag
The vision of Multi-Agent V2 is compelling: a high-tier orchestrator like GPT-5.6 Sol handles the strategy and delegates small tasks to faster, cheaper agents like GPT-5.6 Luna. This is a bespoke solution that balances quality and cost.
In reality, if you haven't touched your config, your system is just cloning itself. Instead of using cheap "workers," your orchestrator spawns expensive clones for every minor sub-task. I have seen single sessions spawn seven parallel Sol agents, consuming hundreds of millions of tokens in minutes.
The Hidden Schema Trap: Why Your Orchestrator Thinks It Can't Route
The problem starts with a parameter called hide_spawn_agent_metadata. By default, it is set to true. This effectively blinds the model to its own capabilities. It removes the fields for model selection, reasoning effort, and service tier from the tool definition the model sees.
The model looks at its available tools and doesn't see a "select model" button. Consequently, when it creates a new agent, it simply replicates its own settings. It is not a bug; it is a default designed for simplicity that happens to be disastrous for your ROI.
| Component | V2 Default | Business Impact |
|---|---|---|
| Model Selection | Hidden | Expensive Sol used for simple Luna tasks |
| History Forking | Full History | Paying double or triple for the same tokens |
| Reasoning Effort | Inherited | Wasted compute time on technical grunt work |
Token Bloat: The Hidden Cost of Full-History Forking
When a sub-agent is spawned in V2 without explicit overrides, it receives the parent's entire conversation history. This is Context Rot in action. A small agent meant to check if a file exists suddenly has to "read" a whole book of context before it can start.
The fix is using fork_turns = "none". This creates a clean agent that only receives the specific instructions for its task. This is a mindful step that saves money and prevents the agent from getting distracted by irrelevant data.
The Fix: Restoring Control via TOML Configuration
To stop the financial bleeding, you must update your .codex/config.toml. The first step is exposing the hidden metadata.
[features.multi_agent_v2]
hide_spawn_agent_metadata = false
Next, establish a clear hierarchy. Don't let the model guess. Set a default sub-agent model, such as GPT-5.6 Terra, which is more balanced in its costs.
The New Hierarchy: When to Use Sol, Terra, and Luna
Managing an agent fleet requires understanding roles. Sol is the CEO – expensive, smart, and shouldn't be writing test code. Terra is the project manager or senior dev. Luna is the fast scanner, the "eyes" on the ground.
Using Luna for read-only tasks can reduce task costs by 60% to 80% compared to using Sol throughout. This is the difference between a profitable project and a bottomless pit of cloud expenses.
Sources
- Sub-Agent Model Routing in Multi-Agent V2: Why Your Sol Orchestrator Spawns Seven Copies of Itself — and How to Fix It (Codex Knowledge Base)
- Restoring subagent roles, model and reasoning in multi_agent_v2 (OpenAI Developer Community)
- The Multi-Agent V2 Governance Playbook: From Encrypted Delegation to Fleet Cost Control (Codex Knowledge Base)
- OpenAI lets GPT-5.6 Sol delegate grunt work to cheaper Luna agents (RuntimeWire)
FAQ
Why does my model say it cannot specify sub-agent models?
This happens because the tool schema is hidden. The model genuinely believes it lacks the capability because it isn't in the tool definition it sees. You must set hide_spawn_agent_metadata = false in your configuration.
What is the difference between Sol and Luna in an agent system?
Sol is a high-reasoning, high-cost model best for planning and coordination. Luna is a fast, low-cost model designed for focused tasks like text scanning or simple execution without needing broad context.
Is history forking always bad?
Not always, but the V2 default of full-history forking causes redundant costs. For isolated tasks, always use fork_turns = "none" to save tokens and maintain precision.
Things to Remember
- Check your
config.tomltoday; the defaults are a financial trap. - Use Luna for read-only and search tasks to maximize ROI.
- Do not let agents inherit the full conversation history unless it is strictly necessary.
Do you know how many tokens your orchestrator is wasting right now on tasks a model ten times cheaper could handle?
Thinking about an agent for one of your workflows?
Most agent projects fail on scope, not on the model. A pilot picks one workflow and proves it end to end.
Related Articles
Explore all AI Agents
The Agent-Readiness Audit: Solving the Website Bottleneck
Is your website ready for AI agents? Discover why systems architecture is the new SEO and how to pass an agent-readiness audit with pragmatic technical fixes.

The End of the Blank Check: Why Agentic Governance Is the New Standard
Stop giving AI agents a blank check. Learn how session budgets, advisor models, and governance controls make automation sustainable for SMBs.

The Model Migration Tax: Why Chasing the Frontier Kills Velocity
Stop chasing every new AI model release. Learn how the 'Model Migration Tax' kills engineering velocity and why pinned versions are the key to operational maturity.