October 1, 2026
AI agents cost less: routing in Deep Agents cut costs by 74%
In LangChain's Aug 11, 2026 benchmark, an agent with model routing spent 74% less, but its success rate fell from 86% to 80%.

Harrison Chase
@hwchase17
Model routing: "nobody knows what it means, but it's provocative and it gets the people going" We're trying to figure out how to do it right 1. understand the tasks 2. understand the models 3. integrate the router **into the agent's runtime** 4. track results Lower costs without sacrificing quality
https://x.com/i/article/2105489324590645248

· 1.5K views
Across 145 multistep tasks, an agent with model routing spent $3 instead of $11.45 when using only Opus 4.8.
Cheaper model first. In escalation mode, the agent starts a task on Nemotron 3.5 Lightning. After 2 consecutive negative evaluations, a separate judge model switches it to Opus 4.8 for the rest of the session. Instead of using one expensive model throughout, the agent gets the expensive model after unsuccessful steps.
LangChain's Aug 11, 2026 benchmark:
| Metric | Opus 4.8 only | Nemotron with escalation to Opus | | --- | --- | --- | | Run cost | $11.45 | $3.00 | | Success rate | 86% | 80% |
Nemotron handled 93% of calls, while Opus handled 7%. The judge model accounted for 21.2% of routing costs, and each turn incurred roughly 700 ms of additional latency.
Model routing is built into the agent's runtime, where it performs tasks and calls tools. The proposed workflow consists of 4 steps:
1. Understand the tasks. 2. Understand the models' capabilities. 3. Integrate model routing into the agent's runtime. 4. Track results.
How to connect it. In Deep Agents, model routing is integrated through `SwitchyardRoutingMiddleware` in `create_deep_agent`. The instructions require cloning Switchyard and `langchain-nvidia`, installing them locally, and using Python ≥3.12 with `deepagents` ≥0.7.4.
Reproducing the benchmark requires a Switchyard server. Set its `base_url` for the agent, define the models in the config, and enable `llm_classifier` in `escalation` mode. The middleware integration example uses a different algorithm, so this specific setting matters when reproducing the benchmark.
The middleware stores the selected model and decision sequence in `response_metadata["switchyard"]`, in the `selected_model` and `decisions` fields.
Source
