model-router
A decision rubric that maps a task's shape to a model tier and an effort level, so the default is the cheapest model that can do the job and the expensive tier is a deliberate choice, gated at spawn time, not a habit.
Problem
Running every agent job on the top model is simple and wasteful; running everything on the cheapest one is cheap and unreliable. The useful behaviour is tiering: mechanical, well-specified work goes to a small model, genuine reasoning and fan-in synthesis go up a tier, and the top tier is reserved for the tasks that actually need it. The trap is that "just use the big model to be safe" is the path of least resistance, so tiering only holds if something makes the cheap tier the default that has to be argued out of.
Architecture
Design decisions
- A hard gate on opus, not a guideline. The rubric is advisory, but opus is
enforced:
cc-opus-gate.shruns before every spawn and exits with a refusal unlessCC_OPUS_REASONnames the architectural decision the run makes. The rejected alternative was documentation ("prefer sonnet") — which loses to convenience every time. A gate that blocks the spawn does not. - The gate keys on the token that actually drives cost. Opus costs a flat 5× per token on every line item, and because 90%+ of a run's volume is cache-read of the growing transcript, the model multiplier — not how tightly the task is written — sets the bill. So the gate asks the one question that changes the answer: is this run making a decision (opus) or executing one already made (sonnet)?
- Measure routing from telemetry, not intention. Every orchestrated run writes its model to a Postgres row, so the routing distribution is an observed fact rather than a claim about what the rubric "should" do.
Numbers
Observed on 2026-09-26. The model distribution is across the 426 runs that carry a model tag in telemetry; the cost figures come from the separate per-run cost logs:
Most work lands on the mid tier rather than defaulting to the top model — the tiering discipline showing up in the data. The gate is not decorative: of 83 opus decisions it logged, 18 were refused outright for lacking a stated reason and bounced to sonnet.
The cost lever is real and measurable. Across the 294 opus runs that carry a
cost breakdown, the actual spend was $8,776 against a sonnet-equivalent of
$1,755 — a ratio of exactly 5.00×, confirming the flat
per-token multiplier the gate is built around. That is the premium paid on runs judged worth opus;
it is stated as the cost of the tier, not as a saving, because the counterfactual (those runs on
sonnet) was never run. The telemetry DB's own cost_usd column is null for all 442 rows,
so this figure comes from the per-run cost logs, not the database.
Failure modes found in production
- Tier creep upward. Without an enforced gate, "just use the big model to be safe" becomes the default and the cheap tier goes unused; the gate exists to make the small model the default that has to be argued out of.
- Attribution is indirect. The distribution reflects routing decisions made by the drivers and the operator, informed by the rubric — it is not proof the rubric alone chose each model. Stated so the number is not over-read.
- Cost not in the database. Dollar cost lives only in per-run text logs, not the telemetry table, so any cost query has to scrape logs — brittle, and the reason the earlier version of this page claimed no dollar figure at all.
Limits
- The 5.00× and the $8,776 cover only the 294 opus runs that produced a cost log; 16 of the 442 telemetry runs carry no model tag at all (11 null, 3 synthetic, 2 unknown) and are excluded from the distribution.
- A per-week overpay figure could not be reconstructed reliably from the logs — the one prior estimate did not survive re-computation, so no weekly number is published here.
- The gate enforces opus specifically; sonnet-vs-haiku routing is rubric-guided but not gated, so the cheap-tier share is a floor a stricter gate could raise.
What's next
- A Tier-0 local model below Haiku for the most mechanical lanes is planned — it lives in the orchestrator only as a stub today.
- Write per-run cost into the telemetry column so tiering can be reported in dollars from one query instead of scraped logs.
Links
No public repository — this is a decision rubric plus a spawn-time gate, presented here for the routing discipline and its measured outcome rather than as a shipped codebase.