← all cases

nightshift

A driver fleet for running unattended coding agents — sequential chains and parallel worktree swarms — with a machine verify gate on every unit of work, cost pre-flight before spawn, and full run telemetry in Postgres.

BashClaude Codegit worktrees PostgreSQLsystemdMIT

Problem

An unattended agent run has no human to catch it when it drifts. Left alone it will idle on a notification that never comes, report success on a job that failed its own acceptance test, or quietly burn budget with no ceiling. Scaling from one run to a fleet multiplies every one of those failure modes. The work needed a harness, not more prompting.

Architecture

plan filelane|task|verify|model lane 1 · worktreehaiku · verify-cmd lane 2 · worktreehaiku · verify-cmd lane N · worktreeMAXPAR ceiling fan-inSonnet Postgrestelemetry
Each lane runs in its own git worktree so lanes never share a working copy. A lane only counts as done when its verify-cmd exits 0. The optional fan-in reduces lane output on a dedicated branch. Every run — chain, swarm, lane, fan-in — is written to a Postgres schema.

Decisions and trade-offs

Numbers

Snapshot 2026-09-26. Every figure below has its source query in queries.sql alongside this run's artifacts; see Links. The telemetry is live, so counts move by a few between snapshots.

792run directories on the node
443runs in telemetry — 402 ok / 37 fail / 4 limit
3.67B / 22Minput / output tokens metered across 479 runs
5 / 18 minp50 / p90 run duration
$4,571metered spend, 421 runs · $3,493 of it on Opus
$2,795Opus-vs-Sonnet-equivalent overpay in that spend
28% / 76%Opus is 28% of runs but 76% of spend — the reason the opus gate exists
26commits · MIT · 13 test scripts

A representative parallel job — a design-compliance audit run on 2026-09-02 — fanned out to 11 lanes in two batches under a parallelism ceiling of two, and metered ~22.0M input and ~150.7K output tokens (11 lanes plus a Sonnet fan-in).

The pre-flight estimator is directionally useful but not precise: of 309 predictions, 231 have since been resolved against an actual run, with a median absolute error of ~81% (skewed by outliers to a 353% mean) — it beats a seed guess but nobody should budget a 5h/weekly cap against it to the dollar.

Machine verification is thin on real data so far — 26 verify-cmd runs recorded, all 26 passing. The 13 test scripts cover the harness guardrails themselves — max-parallelism, lane timeout, verify denylist, run-id substitution, escalation cap and the sync-only / background-wait rule.

What the guardrails target

Honest answer first: this harness has no clean before/after. Guardrails were added continuously, not at one cutover, and the 2026-09-22 consolidation into this repo postdates almost all of the telemetry — so a fail-rate drop can't be pinned on any one rule. What the data does support:

Failure modes found in production

Hung and phantom runs

Snapshot as of 2026-09-27: no hung runs. No orphaned claude -p or codex exec processes.

History is less clean, and the gap is real, not just unlucky timing. 11 runs since 2026-08-22 exit with code 4 — clean process exit, no RESULT: line in the log — and that mechanism works as designed: the driver marks them ambiguous, tags a fail_reason of BG-WAIT or NO-RESULT, notifies, and a chain halts on the branch for manual review rather than guessing.

What isn't caught: 41 run directories between 2026-08-05 and 2026-09-20 have no exit_code file at all and no RESULT: line anywhere in their log — the wrapper process itself died mid-run (SSH session drop, a kill, a reboot) before it ever reached its own exit trap. Checked by hand on four of them: one ends mid-turn on a blocked permission request, two never got past writing task.md (the spawn itself never started), one has a 0-byte log. None of these are flagged ok, fail, or ambiguous anywhere — they're invisible unless someone lists the directory tree by hand, which is what finding them for this page required. The mkdir isolation lock only stops a second run from starting in a busy directory; nothing currently revisits a directory a dead wrapper left behind. That's an open gap, not a solved one.

What I'd do differently

Running it

Every state path — run directories, worktrees, DB credentials, notification hook — is overridable by env var, with defaults pointing at the maintainer's node; nothing is hard-wired to this VPS. A single run is cc-run.sh <run-dir> against a directory holding a task.md; chains and swarms take a plan file (lane|task|verify|model per line). The full component list, gate exit-code table, and env-var contract are in the README.

What's next

Links