abtest
A reasoning contract for treating any comparison across agent runs as a real A/B test — and the discipline to reject a comparison that cannot answer its own question before spending a single run on it.
Problem
"Which output style produces better agent runs" sounds measurable and usually is not. Runs differ in task, model and difficulty far more than in style, so a naive tally measures the workload, not the treatment. The interesting engineering here was refusing the easy answer: defining the arms, the control, the unit of analysis and the estimand up front, computing the ceiling of the effect before designing anything, and killing designs that only looked like evidence.
Architecture
digit mod 3),
so the control (~⅓) falls out on its own and assignment is done by the node wrapper, not by the
model mid-chat. Failed runs act as a filter, not an outcome; the outcome is tokens-per-run and
duration, compared only within a task type.Design decisions
- Ceiling before design. The first gate computes the largest effect the treatment could physically have on the metric that matters. Output style only changes how the prompt and the final report are worded — a tiny slice of a run's tokens. Across 294 opus runs with a cost decomposition, 91% of tokens on average are cache-read of the growing transcript (min 24%, max 97%); style touches almost none of that. If the ceiling sits below the noise floor, the question is unanswerable at any N — so it is closed cheaply instead of funded. The rejected alternative was the usual one: design the experiment first, discover it was hopeless after burning the runs.
- Randomized assignment, not style-by-task. An earlier rule tied caveman to
mechanical tasks and ponytail to decision tasks — which confounds the treatment with difficulty
exactly. It was replaced (24 Aug) by a digit of π at a running index:
digit mod 3→ caveman / ponytail / control. The rejected alternative, simply inverting the mapping, keeps the confound; only assignment independent of the task removes it. - Two synthetic designs rejected outright. "Parallel lanes of one task" and "one task run three ways" were both refused (23 Aug) as counting one observation as three — they do not rule out coincident drift, so they inflate N without adding independent contrasts.
Numbers
The honest headline is what the data does not license, backed by real figures:
No winning style is claimed — that is the result. Everything gathered under the earlier
"style by task type" rule is treated as descriptive "period one", not an effect, precisely because
assignment was confounded with difficulty. The one metric that isolates style rather than task
size — the length of the final report — is monotonic in exactly the order the styles specify
(1183 / 1638 / 1974 characters), which is the strongest evidence the style flag reaches the model at
all. But report length is a rounding error against millions of transcript tokens, so it does not
move the token-cost outcome. The clean control within the data confirms it: in two structurally
identical design → impl → docs triples with swapped styles, the direction of the token
difference flips between roles — an effect of the task, not the style — and the most
comparable pair (docs, both caveman) differs by 4%.
Failure modes found in production
- Assignment confounded with difficulty. The original rule tied one style to mechanical tasks and another to decision tasks — so any observed difference could just be the task. Fixed by randomizing assignment independently of the task.
- Per-batch balancing under-sampled the control. Short batches systematically missed the control arm; the fix was a running index carried across batches, not reset per batch.
- Style not written to telemetry — sample leakage. The run telemetry DB has a
stylecolumn but it is empty for all 442 rows; roughly a third of runs were spawned directly, bypassing the wrapper that logs style, so their arm was never recorded. Tracking lives in a separate Craft table filled by hand — the largest structural weakness of the setup. - Treatment reaching the system is unverified. That the
outputStyleflag is accepted by the CLI and the file exists is confirmed; that it changes model behaviour at the transcript level is not — the report-length signal is the only indirect evidence.
Limits
- The 100 tracked runs are almost all opus and heavily touchviz-weighted; the within-task cells are small, and caveman still skews toward the shortest tasks, so residual confound remains even after randomization.
- Style compliance is not guaranteed: at least one caveman run on sonnet emitted a long narrative report against its own spec, so an arm's label does not certify the treatment took.
- No noise floor has been measured directly — the same task, same arm, repeated — so the required N to state an effect is still unknown. That measurement is the gate on any future claim.
What's next
- Write style into the telemetry DB at spawn time so the tracked sample stops leaking into a hand-maintained table.
- Measure a real noise floor (repeat one task under one arm), then decide whether the required N is reachable given the 91% ceiling — or record the question as closed.
Links
No public repository — this is a reasoning contract applied to a private run dataset, presented here for the methodology rather than the code.