← all cases

skill-governance-oss

A governance spec plus three linters for collections of agent skills. One skill needs no framework; a hundred skills sharing state, credentials and trigger space do.

Python 3.8+ stdlibClaude Skills static analysisCI gate

Problem

Agent skills accumulate. Each is written in isolation, but they run in one namespace: they claim overlapping trigger phrases, declare inheritance from skills that may not exist, and own shared state that nobody registered. None of that shows up when you read a single skill — it only emerges across the whole collection, and by then the collisions are silent bugs. A per-file linter cannot see this class of defect by construction: it has no notion of "the same state, described in two different files." The framework makes those cross-file relationships checkable.

Architecture

skills dir+ state registry audit_triggers.pyphrase collisions lint_dependencies.pyinheritance contracts validate_registry.pyone owner per state unit FAIL / WARNFAIL blocks CI
Three independent linters over a skills directory, standard library only, no install step. FAIL blocks; WARN informs. The registry linter is pointed at a separate ownership file, so state ownership is checked against an explicit source of truth rather than inferred.

Design decisions

Numbers

A fresh run on 2026-09-26 against the private collection the framework was extracted from, with the original 2026-09-02 figures shown for the delta. The collection grew from 116 to 127 skills over that window — the drift is the point:

42state-ownership failures (was 37 on 2026-09-02)
11unhonoured inheritance contracts (was 10)
16trigger-phrase / homoglyph warnings (was 12)
127skills scanned (was 116) · 11 shipped in the repo

The most telling number is the 42: the state-ownership registry has still never been filled in, now well over a year after the ownership convention was adopted. Every skill that declares owned state and appears in zero registry rows is a real, unregistered owner — a live Postgres schema or a bot identity behind a credential — not a hypothetical. It grew from 37 to 42 simply because eleven more skills arrived and none of them registered their state either. That is drift a static gate catches and a code review does not.

Run against the repo's own skills, the only failures come from the deliberately broken examples/collection/ — the shipped skills contribute zero. The example collection carries one of each defect on purpose, so anyone can verify the tooling before pointing it at their own collection.

Failure modes found in production

Limits

What's next

Links