Level up your scripts
Pick one script that sounds fun. Add it to ## Context, re-run, and see how the new signal changes the roast.
Repo Roast track
Two of these use awk, so add Bash(awk:*) to your allowed-tools first.
Bus factor. Files only one person has ever touched:
Bus factor: !`git log --format='@%an' --name-only | awk '/^@/{a=substr($0,2);next} NF{k=$0 SUBSEP a; if(!(k in seen)){seen[k]=1; n[$0]++; who[$0]=a}} END{for(f in n) if(n[f]==1) print f " (" who[f] ")"}' | head -15`Commit crimes. The laziest commit messages:
Commit crimes: !`git log --oneline --since='6 months ago' | grep -iE '^[a-f0-9]+ (fix|wip|temp|asdf|update|stuff|changes)$' | head -15`Zombie branches. Remote branches nobody touches, oldest first:
Zombie branches: !`git for-each-ref --sort=committerdate --format='%(committerdate:relative) %(refname:short)' refs/remotes/ | head -15`3am commits. Code written after midnight:
Late-night commits: !`git log --format='%ad %s' --date=format:'%H:%M' --since='6 months ago' | awk '$1 < "06:00"' | head -10`README vs reality. Commands the README tells you to run:
README commands: !`grep -oE '(npm|yarn|pnpm|make) (run )?[a-z:-]+' README.md 2>/dev/null | sort -u | head -15`Ask the skill to check those commands against package.json or the Makefile.
Draft Roast track
Checkpoint 2 already has these. If you built your own, add one:
Hedge words. The words that make a claim sound unsure:
Hedge words: !`grep -oiwE 'just|really|very|actually|basically|quite|truly|potentially' draft.md 2>/dev/null | tr 'A-Z' 'a-z' | sort | uniq -c | sort -rn`Passive voice. Sentences that hide who did the thing:
Passive voice: !`grep -oiwE '(is|are|was|were|be|been|being) [a-z]+ed' draft.md 2>/dev/null | head -20`Longest paragraphs. Where readers start skimming:
Longest paragraphs: !`awk 'BEGIN{RS=""} {print "paragraph " NR ": " NF " words"}' draft.md 2>/dev/null | sort -t: -k2 -rn | head -3`Links. Is there anything for the reader to click?
Links: !`grep -nE 'https?://|\]\(' draft.md 2>/dev/null || echo "none found"`Then write your own. What do you always check before you hit send? Your team’s banned words? Whether the first line says what changed? Turn it into a script.
On the web, turn your favorite into an instruction: “Before critiquing, list every passive-voice phrase.”
Every script above is one pipeline that always exits successfully. Keep yours that way (see the gotchas in Build the foundation).
Same skill, everywhere
The file you wrote is plain markdown. Where it runs changes what it can do:
- Claude Code and Codex. Scripts run. You get the full evidence-backed roast. Codex looks for skills in
~/.agents/skills/instead of~/.claude/skills/. - Claude on the web or desktop. No shell, so the
!lines don’t run. The constraints and structure still shape the answer. Paste the evidence in and the skill reasons over it. - Agents you build. The Claude Agent SDK and other harnesses load skills too. The skill becomes the brains of a program, a CI job, or a product.
Skills at production scale
The WorkOS CLI installer works this way. It runs on the Claude Agent SDK, and every decision is a skill: which framework you’re using, what to install, how to check that it worked. Small skills call other skills.
It’s the same markdown format you’ve been writing all session. The domain changes. The patterns don’t.
Measurement matters
A skill can look helpful and still make things worse. Correct evidence plus missing context can produce confident, wrong recommendations.
If you rely on a skill, measure it. Even a light version beats vibes: keep three real inputs, save the output before a change, run them again after, and compare.
In the wild
Here are a few of my own skills. Each one is a markdown file shaped like the one you just built.
- grill. Interviews you about a plan until nothing is left assumed. It asks every question it can in rounds, gives a recommended answer for each, and looks up facts itself instead of asking you. The whole skill is 23 lines.
- explain-this. Explains a paper, article, or codebase based on a learner profile you build once. Then it quizzes you and schedules review cards. It’s the clearest progressive disclosure example I have: it reads
references/adapters/paper.md,article.md, orcode.mddepending on what you hand it, and opens the quiz rules only when it’s time to quiz. - brainrot. Turns a doc into a split-screen, TikTok-style lesson: narrated beats on top, an endless parkour game underneath. The model writes the lesson. Bundled scripts build the page and record the voice. The work is split cleanly: judgment in markdown, everything repeatable in code.
- wizard. Writes an interactive bash script that walks a person through steps only a human can do, like copying an API key out of a dashboard. A bundled template handles the prompts, progress, and
.envwrites. The skill only decides the steps.
One thing you’ll notice in their frontmatter: most of them set disable-model-invocation: true. That means the agent never picks them on its own. They only run when called by name. That’s a Claude Code setting, and it’s a trade-off. A skill that routes automatically is easy to reach, but it can also fire when you don’t want it. For a skill with a big output or side effects, calling it by name is often the better default.
Check
You added at least one new script or instruction and can point to a finding it produced that wasn’t there before.