← all cases

model-router

A decision rubric that maps a task's shape to a model tier and an effort level, so the default is the cheapest model that can do the job and the expensive tier is a deliberate choice, gated at spawn time, not a habit.

cost routingmodel selection spawn-time gatetelemetry

Problem

Running every agent job on the top model is simple and wasteful; running everything on the cheapest one is cheap and unreliable. The useful behaviour is tiering: mechanical, well-specified work goes to a small model, genuine reasoning and fan-in synthesis go up a tier, and the top tier is reserved for the tasks that actually need it. The trap is that "just use the big model to be safe" is the path of least resistance, so tiering only holds if something makes the cheap tier the default that has to be argued out of.

Architecture

task shape+ effort estimate mechanical → Haiku reasoning → Sonnet hardest → Opus (gated) recorded per runmodel field in telemetry
The rubric outputs provider, tool, model, thinking mode and effort. Opus passes through a spawn-time gate that refuses it without a one-line architectural reason. The chosen model is stored on every orchestrated run, so routing in practice can be measured after the fact rather than assumed.

Design decisions

Numbers

Observed on 2026-09-26. The model distribution is across the 426 runs that carry a model tag in telemetry; the cost figures come from the separate per-run cost logs:

67%runs on Sonnet (287)
28%runs on Opus (120)
5%runs on Haiku (19)
18opus spawns refused by the gate (65 passed with a reason)

Most work lands on the mid tier rather than defaulting to the top model — the tiering discipline showing up in the data. The gate is not decorative: of 83 opus decisions it logged, 18 were refused outright for lacking a stated reason and bounced to sonnet.

The cost lever is real and measurable. Across the 294 opus runs that carry a cost breakdown, the actual spend was $8,776 against a sonnet-equivalent of $1,755 — a ratio of exactly 5.00×, confirming the flat per-token multiplier the gate is built around. That is the premium paid on runs judged worth opus; it is stated as the cost of the tier, not as a saving, because the counterfactual (those runs on sonnet) was never run. The telemetry DB's own cost_usd column is null for all 442 rows, so this figure comes from the per-run cost logs, not the database.

Failure modes found in production

Limits

What's next

Links

No public repository — this is a decision rubric plus a spawn-time gate, presented here for the routing discipline and its measured outcome rather than as a shipped codebase.