← all cases

abtest

A reasoning contract for treating any comparison across agent runs as a real A/B test — and the discipline to reject a comparison that cannot answer its own question before spending a single run on it.

experiment designevaluation agent runsceiling-first

Problem

"Which output style produces better agent runs" sounds measurable and usually is not. Runs differ in task, model and difficulty far more than in style, so a naive tally measures the workload, not the treatment. The interesting engineering here was refusing the easy answer: defining the arms, the control, the unit of analysis and the estimand up front, computing the ceiling of the effect before designing anything, and killing designs that only looked like evidence.

Architecture

run index nπ digit mod 3 arm: caveman arm: ponytail control: none outcome per runtokens + duration · RESULT filters
Arm assignment is by a digit of π at the run's running index (digit mod 3), so the control (~⅓) falls out on its own and assignment is done by the node wrapper, not by the model mid-chat. Failed runs act as a filter, not an outcome; the outcome is tokens-per-run and duration, compared only within a task type.

Design decisions

Numbers

The honest headline is what the data does not license, backed by real figures:

100style-tagged runs — 43 caveman / 48 ponytail / 9 control
0 / 442runs in the telemetry DB carry a style tag
91%mean cache-read share — the ceiling on any token effect
1183/1638/1974avg report chars: caveman/ponytail/control

No winning style is claimed — that is the result. Everything gathered under the earlier "style by task type" rule is treated as descriptive "period one", not an effect, precisely because assignment was confounded with difficulty. The one metric that isolates style rather than task size — the length of the final report — is monotonic in exactly the order the styles specify (1183 / 1638 / 1974 characters), which is the strongest evidence the style flag reaches the model at all. But report length is a rounding error against millions of transcript tokens, so it does not move the token-cost outcome. The clean control within the data confirms it: in two structurally identical design → impl → docs triples with swapped styles, the direction of the token difference flips between roles — an effect of the task, not the style — and the most comparable pair (docs, both caveman) differs by 4%.

Failure modes found in production

Limits

What's next

Links

No public repository — this is a reasoning contract applied to a private run dataset, presented here for the methodology rather than the code.