skill-governance-oss
A governance spec plus three linters for collections of agent skills. One skill needs no framework; a hundred skills sharing state, credentials and trigger space do.
Problem
Agent skills accumulate. Each is written in isolation, but they run in one namespace: they claim overlapping trigger phrases, declare inheritance from skills that may not exist, and own shared state that nobody registered. None of that shows up when you read a single skill — it only emerges across the whole collection, and by then the collisions are silent bugs. A per-file linter cannot see this class of defect by construction: it has no notion of "the same state, described in two different files." The framework makes those cross-file relationships checkable.
Architecture
FAIL blocks; WARN informs. The registry linter is pointed at a
separate ownership file, so state ownership is checked against an explicit source of truth rather
than inferred.Design decisions
- Standard library only, no install step. The linters import nothing outside Python 3.8+ stdlib and run as three plain scripts. The rejected alternative was a packaged tool with a YAML config and a frontmatter-schema dependency — cleaner on paper, but it makes the linters something you install and pin rather than something a CI job or a pre-commit hook runs by copying one file. A governance gate that is annoying to adopt does not get adopted.
- State ownership checked against a separate registry, not inferred. The
registry linter compares each skill's declared
## Encapsulationblock against an explicit ownership file. The rejected alternative was inferring ownership from the skills themselves — but "who owns this Postgres schema" is exactly the fact that a skill written in isolation gets wrong, so inferring it from the same files reproduces the bug instead of catching it. An external registry is the one place a second, conflicting owner becomes visible. - De-hardcoded meta-skill and registry names (2.2.0). The framework originally baked in its own meta-skill list and registry filename. That was moved out into the private registry and env overrides so the tool works on any collection, not just the one it was extracted from.
Numbers
A fresh run on 2026-09-26 against the private collection the framework was extracted from, with the original 2026-09-02 figures shown for the delta. The collection grew from 116 to 127 skills over that window — the drift is the point:
The most telling number is the 42: the state-ownership registry has still never been filled in, now well over a year after the ownership convention was adopted. Every skill that declares owned state and appears in zero registry rows is a real, unregistered owner — a live Postgres schema or a bot identity behind a credential — not a hypothetical. It grew from 37 to 42 simply because eleven more skills arrived and none of them registered their state either. That is drift a static gate catches and a code review does not.
Run against the repo's own skills, the only failures come from the deliberately
broken examples/collection/ — the shipped skills contribute zero. The example
collection carries one of each defect on purpose, so anyone can verify the tooling before pointing
it at their own collection.
Failure modes found in production
- Duplicate trigger owners. Two skills claim the same phrase, so which one
loads is undefined. The example collection ships this on purpose:
'shared-cache' claimed by 2 skills. - Dangling dependency. A skill declares
Inherits from:an upstream that does not exist — the contract can never be honoured. - Homoglyph in a trigger-adjacent word. A Latin letter mixed into a Cyrillic word renders identically in nearly every font. In body prose it is harmless; one character over in a quoted trigger phrase it produces a phrase that looks present but matches nothing — a collision the linter can no longer see, because the string comparison it depends on now fails on a byte level nobody can read. Only a codepoint check finds it.
- Silent state ownership. Shared state with no registered owner: nothing fails at runtime until two skills write it, which is exactly when it is hardest to debug.
Limits
- The registry linter is only as good as the registry. Here it runs against the blank template, so 42 is "everything that declares state" — an upper bound on drift, not a diff against a maintained ownership file (which does not yet exist for this collection).
- The linters detect declared relationships — an
Inherits from:line, an## Encapsulationblock. A skill that owns state but never declares it stays invisible; the framework raises the floor, it does not prove the ceiling. - The trigger check is lexical (shared phrases, mixed scripts). It does not model semantic overlap — two skills that trigger on paraphrases of the same intent pass clean.
What's next
- Ship the linters as a live skill that loads and self-enforces per the repo's own
ENFORCEMENT.md, rather than being run by hand. - Fill in a real ownership registry for the source collection, turning the 42 upper bound into a maintained diff.
Links
GitHub — 0mandrock1/skill-governance-oss CASE.md — full writeup