Back to AI Labs

AI Lab · Agent Skills

Claude Code skills that make an agent follow a method.

A prompt solves a problem once. A skill writes the method down, so it holds every time after.

Portrait of Antonis KatsonisAntonis KatsonisAuthor and maintainer
Repo readiness checkShare

Is this repo ready for an agent to work in?

agent-ready-auditSkill · 24 checks

I ran every command in the guide instead of reading it.

  • Guide found. CLAUDE.md lists build, test and run.
  • Build and run executed. Both exit 0, copy-pasted as written.
  • No pass: tests documented. The command fails as written, so the check blocks.
  • Traps and business rules. Documented, with self-check commands.

Result: one blocking check is open. Fix the test command, then re-run for a score.

Reply to Claude…

Opus 4.5

Claude can make mistakes. Please double-check responses.

Illustrative runs, redacted.
Role
Wrote and maintain them
Format
Claude Code skills
Gates
Refusals, not advice
Status
In daily use

Skills I've built

Each one asks a question first, and refuses a shortcut.

  • Agent-readiness audit

    24 checks · 10 block · scored /100

    Could a stranger open this repo, run it, change it and prove the change worked?

    Refuses: A pass on a command nobody ran.

  • Stuck-order investigator

    5 phases · cohort before mechanism

    Is the field you are blaming actually different from the orders that completed fine?

    Refuses: An explanation before a comparison.

  • Story builder

    14 prompts · ~30 criteria

    Can engineering build from this ticket without guessing what you meant?

    Refuses: Creating while an unconfirmed assumption still holds it up.

How a skill works

SKILL.md

skills/stuck-order-investigator/SKILL.md

  1. ---
  2. name: stuck-order-investigator
  3. description: Investigate a stuck order. Establish whether the cause is knowable from the data — especially when the temptation is to explain it from fields alone.
  4. ---
  5. ## Hard rules
  6. Cohort before mechanism.
  7. No write proposals before the verdict.
  8. State confidence explicitly.
  9. ## Phase 1 — Is this anomalous?
  10. Build the comparison first. Classify by signature, not symptom.
  1. The trigger

    The only part the model reads before it fires.

  2. The refusals

    Each names a thing the agent may not do.

  3. The order

    Each phase must land before the next may start.

Why I built them

An agent with good instructions still improvises. What's missing isn't capability. It's method.

  • Why did this fail?

    A plausible answer, built from fields

  • Can an agent work in this repo?

    Whoever read it last, deciding

  • Is this ticket ready to build?

    As good as its prompt

  • Did that change make it worse?

    Nobody could tell

Three theories, all wrong

One order. Three explanations built from its fields. Each arrived with a fix ready to run.

  • The exchange made no payment record, so the replacement is unfunded.

    Killed by

    8,580orders like it
  • The original order's payment status is stale. A sync will fix it.

    Killed by

    0changed by the sync
  • A fully-returned original blocks its replacement.

    Killed by

    ~2,900replacements that finished

Every one would have changed a healthy order in production. One question killed all three: how many comparable orders are in this state? The skill asks it first now.

Every change is tested before it ships

The Story builder runs fourteen test prompts on every edit, and a second agent grades the output.

  1. EditA change to the skill
  2. Run14 test prompts
  3. GradeA second agent checks ~30 criteria
  4. ShipOnly when every prompt passes

The fourteen prompts

  • 7Ticket types
  • 2Conversions
  • 1Must not fire
  • 4Edge cases

Needed to ship

14 / 14

What it catches that a code review wouldn't

  • Over-triggering

    One prompt has no ticket in it. It passes only if the skill stays out.

  • Format setting the verdict

    Same content in two formats must get the same rating.

  • An assumption as a fact

    An unconfirmed claim blocks the ticket until it is answered.

  • The gate going quiet

    Every draft still waits for approval.

What changed

Checks, every repository
0Checks, every repository
Prompts, every change
0Prompts, every change
Phases before a verdict
0Phases before a verdict
Skills in daily use
0Skills in daily use
Before
After
A prompt that worked once
A method that survives a handover
A plausible explanation
A comparison first
Quality tracked the prompt
The same rigour every time
Silent drift
Fourteen prompts, every change

My role

Each one is something I was already doing by hand and re-explaining to the next person.

  • Wrote the method as prose first, then turned it into a file.
  • Tuned each trigger until it fires on the real request and nothing near it.
  • Turned advice into refusals, because only those hold.
  • Added the suite after a change degraded the output and nothing caught it.

Built with

  • Claude Code
  • Model Context Protocol
  • Markdown
  • Bash
  • Git
  • JSONL records