Agent Routing (model, effort & cascade selection)
Decides which model, effort level, and cascade shape each subagent gets, routing on measured cost per completed task instead of per-token price.
AutomationAdvanced★ 150⑂ 6AI score 9/10Last updated: Oct 2, 2026
What it does
- Gives a task-shape routing table: two questions (short vs. long output, mechanically checkable vs. judgment) map onto
haiku@low,sonnet@medium,opus@high, and so on. - Routes on measured cost per completed task, not per-token price — with data showing a 5×-cheaper tier costing 30% more per solved task because it emitted 6.7× the tokens.
- Supplies cascade design rules: check the cheap rung is actually cheaper first, no verifier ⇒ no cascade, and the verifier's holder (the orchestrator) makes the escalation call because workers self-report success even when they failed.
- Documents the concision lever (−37% Sonnet / −27% Haiku output on long generation) and where it stops helping, plus how thinking suppression halves pass rates.
- A context handoff checklist: subagents inherit nothing, so artifact paths, verbatim commands (with interpreter path), anti-patterns, and an output spec must be serialized into every spawn prompt.
- Loop discipline (out-of-band evaluator, argmax selection, stop on first regression) and a procedure for watching a subagent fan-out live via per-thread streams.
Who it's for
- Claude Code / Claude Code on the Web users spawning subagents through Agent or Workflow tools
- Teams whose agent pipelines cost more than expected and who want to know why
- Orchestration designers who want an evidence-based tier policy rather than "just use the big model"
- (Not applicable to plain claude.ai chat usage)
Example uses
- 100-way JSON extraction fan-out: short, checkable output →
haiku@lowplus schema validation, with an explicit argument against up-tiering "to be safe" (3–5× cost, no measured gain). - Spec-to-module generation: rung 1
sonnet@low+ concision, and on test failure retry the same model atmediumcarrying the prior patch and raw test output — cheaper and more accurate than jumping to opus. - Four explore agents on a 2,300-file repo: hand each one an index slice, the exact command, and a no-
ls/glob rule so discovery cost drops to roughly zero.
· · · Install guide · · ·
Try it now, no install
Paste this into Claude to use the skill without installing anything.
Read the instructions in this file and follow them to help me: https://raw.githubusercontent.com/oaustegard/claude-skills/HEAD/agent-routing/SKILL.md What I want: (describe your task here)
If Claude can't open the link, open it yourself and paste the contents instead.
↓ If it works for you, download the ZIP below and install it. Then it runs on its own — no pasting each time.
Install in the Claude app (no terminal)
- Download the ZIP with the button below.
- In Claude, open Settings → Capabilities and turn on 'Code execution and file creation'. (one time)
- Go to Customize → Skills → + → 'Upload a skill' and upload the ZIP.
Install in Claude Code
Let Claude do it — paste this into Claude Code
Install the skill I found on Claude Skill Mart. Copy the agent-routing folder from the GitHub repo oaustegard/claude-skills into my ~/.claude/skills/agent-routing/. When it's done, tell me in one line what this skill can do.
Install with a command instead
git clone https://github.com/oaustegard/claude-skills.git && mkdir -p ~/.claude/skills && cp -r claude-skills/agent-routing ~/.claude/skills/⚠ This is a third-party skill. Check the source repository before installing.
- Open a terminal and create the skills directory:
mkdir -p ~/.claude/skills - Clone the repository:
git clone https://github.com/oaustegard/claude-skills.git - Copy just this skill:
cp -r claude-skills/agent-routing ~/.claude/skills/ - Confirm both
~/.claude/skills/agent-routing/SKILL.mdand itsreferences/folder came along — the reference files hold the calibration data. - Restart Claude Code and check that
agent-routingshows up in your skills list. - Trigger it by asking "which model and effort should this task get?" or whenever you ask Claude to fan out several subagents.
- Note that the price and token figures are snapshot measurements; if your models or pricing have changed, re-measure on your own workload instead of trusting the tables as-is.
View source on GitHub ↗License: MIT