nightshift
A driver fleet for running unattended coding agents — sequential chains and parallel worktree swarms — with a machine verify gate on every unit of work, cost pre-flight before spawn, and full run telemetry in Postgres.
Problem
An unattended agent run has no human to catch it when it drifts. Left alone it will idle on a notification that never comes, report success on a job that failed its own acceptance test, or quietly burn budget with no ceiling. Scaling from one run to a fleet multiplies every one of those failure modes. The work needed a harness, not more prompting.
Architecture
verify-cmd exits 0. The optional fan-in reduces lane
output on a dedicated branch. Every run — chain, swarm, lane, fan-in — is written to a Postgres
schema.Decisions and trade-offs
- bash, not Python. The whole job is process control — spawn a
subprocess under
timeout, read its exit code, tail its log, hold a lock, fire acurlnotification. That is bash's native surface; a Python layer would add a dependency and an interpreter to the one thing that must keep working when a node is starved. The cost is paid in string discipline (set -euo pipefail, quoted paths) and in tests — 13 of them — rather than in a runtime. - A git worktree per lane, not a lock or branch-switch. Parallel
lanes each mutate a working copy. Sharing one copy behind a lock serializes the
thing that was supposed to be parallel; switching branches in place races on
.git/index. A worktree gives each lane its own checkout and its own branch off the same object store, so lanes are isolated without duplicating the repo. Cleanup is deferred (cc-gc.sh) and only ever deletes a branch already merged into its fan-in ormain. - A self-reported
RESULTline and a machineverify-cmd— not either alone. TheRESULT: ok|failcontract is how a run declares intent cheaply; it costs one line and lets the driver classify a run that did something no external check can see (a doc rewrite, a migration plan). But a model's self-report is not evidence — so a lane only counts as done when itsverify-cmdexits 0. The two catch different failures: a missingRESULTline is ambiguous (the run may have hung), a failingverify-cmdis a real fail. Conflating them is the trap the 2026-09-02 audit swarm fell into (below). - Postgres, not just the files on disk. Every run already leaves
files (
task.md,out.log,usage.json,cost.log). Telemetry is a best-effort UPSERT of one row derived from those files — it never fails the run that triggered it. The database earns its keep only for the cross-run questions the files can't answer cheaply: fail rate by week, cost by model, estimator error over 200+ resolved predictions. The disk stays the source of truth; the DB is the index over it.
Numbers
Snapshot 2026-09-26. Every figure below has its source query in
queries.sql alongside this run's artifacts; see Links. The telemetry is
live, so counts move by a few between snapshots.
A representative parallel job — a design-compliance audit run on 2026-09-02 — fanned out to 11 lanes in two batches under a parallelism ceiling of two, and metered ~22.0M input and ~150.7K output tokens (11 lanes plus a Sonnet fan-in).
The pre-flight estimator is directionally useful but not precise: of 309 predictions, 231 have since been resolved against an actual run, with a median absolute error of ~81% (skewed by outliers to a 353% mean) — it beats a seed guess but nobody should budget a 5h/weekly cap against it to the dollar.
Machine verification is thin on real data so far — 26 verify-cmd
runs recorded, all 26 passing. The 13 test scripts cover the harness guardrails themselves —
max-parallelism, lane timeout, verify denylist, run-id substitution, escalation cap and the
sync-only / background-wait rule.
What the guardrails target
Honest answer first: this harness has no clean before/after. Guardrails were added continuously, not at one cutover, and the 2026-09-22 consolidation into this repo postdates almost all of the telemetry — so a fail-rate drop can't be pinned on any one rule. What the data does support:
- Cost, not prompt-tuning, is the lever — and the gate targets it.
Opus is 120 of 431 metered run+lane rows (28%) but carries $3,493 of $4,571 spend
(76%). ~95% of a run's token volume is cache-read of a growing transcript, so model
choice dominates cost far more than how tightly
task.mdis written. That asymmetry is the entire reasoncc-opus-gate.shrefuses an Opus run unless a reason is set. Its effect is not measurable from telemetry — refused runs never spawn, so no row is ever written for them; there is no "before" set to compare. The gate is justified by the 28%/76% split it forces a decision against, not by a measured delta. - The $2,795 overpay figure has a counterfactual, stated. It is the sum of what those Opus runs would have cost at Sonnet tier (Sonnet is exactly 5× cheaper on every price component). It assumes the runs would have produced equivalent output at Sonnet — which is precisely the thing that isn't guaranteed. That uncertainty is why the gate demands a written reason rather than banning Opus outright.
- Ambiguous detection fires uniformly, not post-hoc. The 11 exit-code-4 runs are spread across the window — 6 in the week of 08-24, 3 in 08-31, 2 in 09-21 — i.e. the mechanism catches the pattern whenever it occurs, it wasn't bolted on after one bad week. The one visible fail spike (week of 08-24: ~22 fails of 132 run+lane rows, ≈17%, dropping to low single digits after) sits right after the first release and its 12-finding hardening pass; I won't over-claim which change drove the decline.
Failure modes found in production
- RESULT vs. audit semantics. A run can exit clean yet its
RESULT: ok|failline disagrees with what the audit task actually meant — the 2026-09-02 swarm above is recordedfailat the swarm level despite lanes completing. Parsing a self-reported line is not the same as verifying the outcome. - Missing RESULT block. A run that ends without emitting its final
RESULTline is unclassifiable; the driver has to treat "no result" as its own state rather than assume success. - Background-wait deadlock. A headless run that backgrounds a step and then
waits for a completion notification idles until timeout — there is no channel to deliver
that notification. Caught live twice on 2026-09-22 in the same chain
(
tf-gen-20260922-1954, thentf-topics-20260922-2041two steps later) and again on 2026-09-02 (pf-raglab-20260902-0914/-1150): the model's own log said, verbatim, that it would "wait for the background task notification" or had "scheduled a check-back" instead of polling synchronously. The harness gained an explicit sync-only rule and aBG-WAITmarker for it — the rule is appended to everytask.md, but it's advisory text, not a sandbox restriction, so the same pattern can still recur. - Two runs into one output tree. Two parallel runs writing the same web
directory (a staging vs. production collision) corrupt each other — resolved with a
mkdirlock that refuses to start a second run on a busy path. - Divergent driver copies. The driver scripts had drifted into three separate copies across a repo and two skills; consolidating them into this single-source repo was the reason nightshift was extracted as a standalone product.
Hung and phantom runs
Snapshot as of 2026-09-27: no hung runs. No orphaned claude -p or
codex exec processes.
History is less clean, and the gap is real, not just unlucky timing. 11 runs since
2026-08-22 exit with code 4 — clean process exit, no RESULT: line in the log — and
that mechanism works as designed: the driver marks them ambiguous, tags a
fail_reason of BG-WAIT or NO-RESULT, notifies, and a chain
halts on the branch for manual review rather than guessing.
What isn't caught: 41 run directories between 2026-08-05 and 2026-09-20 have no
exit_code file at all and no RESULT: line anywhere in their log — the
wrapper process itself died mid-run (SSH session drop, a kill, a reboot) before it ever reached
its own exit trap. Checked by hand on four of them: one ends mid-turn on a blocked permission
request, two never got past writing task.md (the spawn itself never started), one
has a 0-byte log. None of these are flagged ok, fail, or
ambiguous anywhere — they're invisible unless someone lists the directory tree by
hand, which is what finding them for this page required. The mkdir isolation lock
only stops a second run from starting in a busy directory; nothing currently revisits
a directory a dead wrapper left behind. That's an open gap, not a solved one.
What I'd do differently
- Make the sync-only rule a sandbox restriction, not advice. Today
the background-wait deadlock is prevented by appended text in every
task.md— advisory, so the same pattern still recurs (it did, twice on 2026-09-22). The right fix is to deny the background-notification tools at the harness level for unattended runs, so the rule can't be ignored rather than merely stated. - Sweep dead run-dirs instead of only blocking new ones. The
mkdirlock stops a second run entering a busy directory but nothing revisits a directory a dead wrapper left behind — hence the 41 phantom dirs below. A periodic reaper that reconciles run-dirs against telemetry would close that gap. - Populate the columns the schema already has.
cost_usdandstyleexist inswarm.runsbut the telemetry INSERT never writes them, so cost has to be reconstructed from disk by joiningcost.logper run. That reconstruction is why the cost figures cover a matched subset, not all rows — it should be written at run time. - Treat the estimator as a range, not a point. At ~81% median absolute error it's honest to publish, but a driver that budgets against it should consume the IQR band, not the median — and I'd gate spawns on the upper bound.
Running it
Every state path — run directories, worktrees, DB credentials, notification hook — is
overridable by env var, with defaults pointing at the maintainer's node; nothing is
hard-wired to this VPS. A single run is cc-run.sh <run-dir> against a
directory holding a task.md; chains and swarms take a plan file
(lane|task|verify|model per line). The full component list, gate exit-code
table, and env-var contract are in the README.
What's next
- Tier-0 local-model lanes (
local:<tag>) are planned — the code path exists as a stub today, not a working backend. - Idempotent run IDs so a re-run with an existing RESULT is a no-op.