Your skill now gives specific, evidence-backed findings. It also dumps everything it found at once, with equal confidence. This lesson fixes both.
Progressive disclosure
Don’t load everything up front. Load context only when the task needs it.
A skill can be a folder instead of one file. SKILL.md loads when the skill runs. Other files sit next to it, and SKILL.md says when to read each one:
repo-roast/
SKILL.md loaded when the skill runs
references/
scoring-rubric.md read when scoring findings
recommendation-templates.md read when writing fixes
This only works with separate files. If everything lives in SKILL.md, the agent sees all of it every time. A separate file is a real gate: the agent opens it only when the condition is met.
Your roast is small enough to stay in one file today. Keep this in your pocket for when a skill grows past a page or two.
What costs context
Progressive disclosure works one level up, too. Your agent doesn’t load every skill you have. In Claude Code, it loads in stages (docs):
| What | When it loads |
|---|---|
| Name and description of every skill | At the start of every session |
The full SKILL.md |
When the skill runs, then it stays for the session |
Files next to it (references/, scripts) |
Only when the agent reads them |
All those descriptions share a budget of about 1% of the context window, and each one is capped at 1,536 characters. Install too many skills and the least-used ones get dropped from the list. The agent can’t pick a skill it can’t see.
That’s why your description has to be short and specific. It’s also what this frontmatter line does:
disable-model-invocation: true
A skill with this line stays out of the list completely. It costs nothing until you type /name, and the agent will never run it on its own. Use it for skills you only want on purpose: big outputs, side effects, or anything you’d rather call by name. (The reverse is user-invocable: false: hidden from your / menu, but the agent can still use it.)
Try it: add disable-model-invocation: true to your roast, start a new session, and ask “roast this.” It won’t fire. Type /repo-roast or /draft-roast and it runs. Remove the line before you move on.
Phases
Without phases, the agent does everything in one pass. With phases, it gathers evidence first, judges it second, and writes recommendations last. Each step builds on the one before.
Self-assessment
The model is good at reasoning and bad at knowing when it’s guessing. So make it score its own findings before it shows them to you. Anything with weak evidence gets dropped or flagged.
The result: the skill tells you when it’s guessing, instead of presenting everything with the same confidence.
Your turn: add phases and confidence
Repo Roast track
Add these two sections after ## Constraints:
## Workflow
Work through these phases in order. Do not skip phases.
1. Run all context scripts. Gather raw data. Summarize counts and hotspots.
2. Categorize findings by type (complexity, coverage, dependencies, documentation, churn). Score severity. Run self-assessment. Drop or flag weak findings.
3. Build prioritized recommendations. Run constraints checklist. Present final assessment.
## Self-Assessment
Rate each finding before presenting:
- Evidence quality (1-10): Is this backed by script output or inference?
- Severity accuracy (1-10): Is this actually a problem, or could it be intentional?
- Actionability (1-10): Can someone act on this recommendation today?
If any finding scores below 6 on evidence quality, drop it or flag as "needs investigation."
Show the scores for each finding so the reader can see your reasoning.Add a line to ## Structure: “Confidence summary (how many findings were high-confidence vs flagged)”.
Run Roast this repo on the same repo as before.
Draft Roast track
Add these two sections after ## Constraints:
## Workflow
Work through these phases in order. Do not skip phases.
1. Read the script output. Summarize length, the opening, and the worst offenders.
2. Categorize findings by type (buried lede, missing ask, vague claim, jargon, length, tone). Score severity. Run self-assessment. Drop or flag weak findings.
3. Write the fixes. Run constraints checklist. Present final assessment.
## Self-Assessment
Rate each finding before presenting:
- Evidence quality (1-10): Is this backed by a quote or script output, or by inference?
- Severity accuracy (1-10): Will this hurt the reader, or is it just a style preference?
- Actionability (1-10): Could the author paste your fix in right now?
If any finding scores below 6 on evidence quality, drop it or flag as "needs investigation."
Show the scores for each finding so the reader can see your reasoning.Add a line to ## Structure: “Confidence summary (how many findings were high-confidence vs flagged)”.
On the web, step 1 becomes “Count words and find the longest sentences and paragraphs.”
Run Roast this draft on the same draft as before.
Compare
Put this run next to your last one:
- Did the phases change how the answer reads?
- Did self-assessment drop at least one finding?
- Is the output better, or just longer?
Be honest. The before and after is the lesson.
Tune it
The first version is never the good one. Pick one change, re-run, and compare:
- Move the threshold. Try 7 instead of 6. Did you agree with what got dropped?
- Add a dimension. For example, “Specificity (1-10): does this finding name a file?” or “…quote a full sentence?”
- Disagree with a score. If a finding scored high and you think it’s wrong, write a constraint that would have caught it.
That loop, edit, run, read, adjust, is how every skill gets good.
Behind?
Repo Roast track
./setup.sh repo --checkpoint 2Or read checkpoint 2.
Draft Roast track
./setup.sh draft --checkpoint 2On the web, copy checkpoint 2 into your editor.
Checkpoint 2 also adds more scripts. You’ll pick your own in the next lesson.