AI Lab · Agent Skills
Claude Code skills that make an agent follow a method.
A prompt solves a problem once. A skill writes the method down, so it holds every time after.
Is this repo ready for an agent to work in?
I ran every command in the guide instead of reading it.
- Guide found. CLAUDE.md lists build, test and run.
- Build and run executed. Both exit 0, copy-pasted as written.
- No pass: tests documented. The command fails as written, so the check blocks.
- Traps and business rules. Documented, with self-check commands.
Result: one blocking check is open. Fix the test command, then re-run for a score.
Reply to Claude…
Claude can make mistakes. Please double-check responses.
- Role
- Wrote and maintain them
- Format
- Claude Code skills
- Gates
- Refusals, not advice
- Status
- In daily use
Skills I've built
Each one asks a question first, and refuses a shortcut.
Agent-readiness audit
24 checks · 10 block · scored /100
Could a stranger open this repo, run it, change it and prove the change worked?
Refuses: A pass on a command nobody ran.
Stuck-order investigator
5 phases · cohort before mechanism
Is the field you are blaming actually different from the orders that completed fine?
Refuses: An explanation before a comparison.
Story builder
14 prompts · ~30 criteria
Can engineering build from this ticket without guessing what you meant?
Refuses: Creating while an unconfirmed assumption still holds it up.
How a skill works
skills/stuck-order-investigator/SKILL.md
- ---
- name: stuck-order-investigator
- description: Investigate a stuck order. Establish whether the cause is knowable from the data — especially when the temptation is to explain it from fields alone.
- ---
- ## Hard rules
- Cohort before mechanism.
- No write proposals before the verdict.
- State confidence explicitly.
- ## Phase 1 — Is this anomalous?
- Build the comparison first. Classify by signature, not symptom.
The trigger
The only part the model reads before it fires.
The refusals
Each names a thing the agent may not do.
The order
Each phase must land before the next may start.
Why I built them
An agent with good instructions still improvises. What's missing isn't capability. It's method.
Why did this fail?
A plausible answer, built from fields
Can an agent work in this repo?
Whoever read it last, deciding
Is this ticket ready to build?
As good as its prompt
Did that change make it worse?
Nobody could tell
Three theories, all wrong
One order. Three explanations built from its fields. Each arrived with a fix ready to run.
The exchange made no payment record, so the replacement is unfunded.
Killed by
8,580orders like itThe original order's payment status is stale. A sync will fix it.
Killed by
0changed by the syncA fully-returned original blocks its replacement.
Killed by
~2,900replacements that finished
Every one would have changed a healthy order in production. One question killed all three: how many comparable orders are in this state? The skill asks it first now.
Every change is tested before it ships
The Story builder runs fourteen test prompts on every edit, and a second agent grades the output.
- EditA change to the skill
- Run14 test prompts
- GradeA second agent checks ~30 criteria
- ShipOnly when every prompt passes
The fourteen prompts
- 7Ticket types
- 2Conversions
- 1Must not fire
- 4Edge cases
Needed to ship
14 / 14
What it catches that a code review wouldn't
Over-triggering
One prompt has no ticket in it. It passes only if the skill stays out.
Format setting the verdict
Same content in two formats must get the same rating.
An assumption as a fact
An unconfirmed claim blocks the ticket until it is answered.
The gate going quiet
Every draft still waits for approval.
What changed
- Checks, every repository
- 0Checks, every repository
- Prompts, every change
- 0Prompts, every change
- Phases before a verdict
- 0Phases before a verdict
- Skills in daily use
- 0Skills in daily use
- Before
- After
- A prompt that worked once
- A method that survives a handover
- A plausible explanation
- A comparison first
- Quality tracked the prompt
- The same rigour every time
- Silent drift
- Fourteen prompts, every change
My role
Each one is something I was already doing by hand and re-explaining to the next person.
- Wrote the method as prose first, then turned it into a file.
- Tuned each trigger until it fires on the real request and nothing near it.
- Turned advice into refusals, because only those hold.
- Added the suite after a change degraded the output and nothing caught it.
Built with
- Claude Code
- Model Context Protocol
- Markdown
- Bash
- Git
- JSONL records