And who are the judges? Thanks to AI judges, Claude stopped deleting my databases
We give an agent more and more freedom, and one day it does the wrong thing. A list of bans does not save you: a ban knows words, not meaning. So in front of a dangerous action we put another Claude, a judge. It did not see how the first one reasoned: it gets only the intent, what is going to be done, why, and by which exact command line. There are several judges, they answer separately, and one vote against is enough. The permission is narrow: fifteen minutes, one thing, one use. The recording of the masterclass and the working kit of files that goes with it.
The walkthrough is for subscribers
The write-up above is free and stays free. The recorded walkthrough is what a subscription adds to it.
See plansArtifact
AI judges kit, 9 files
- how-to-repeat.md The instruction: six steps, about fifteen minutes. Start here. 11 KB
- judge.py The judge itself: it takes the intent, collects several votes and issues the pass. 5 KB
- judge-prompt.txt What the judge is told: the grounds it decides on. 4 KB
- judge-hook.sh The guard in front of a command: nothing dangerous passes without a fresh pass. 6 KB
- kinds.txt The list of guarded actions. Remove a line and that kind is no longer guarded. 2 KB
- check.py Thirty two test cases. It changes nothing in your system. 9 KB
- settings-home.json The piece of settings for a kit that sits in the home folder. 1 KB
- settings-project.json The piece of settings for a kit that sits inside a project. 1 KB
- rule-for-claude.md Lines for your own rules file, so the agent takes the refusal of the guard seriously. 1 KB
The files themselves are for subscribers. The list above is the whole of it: that is everything in the set.
See plansAsk about this case
The question goes to the author, not onto this page. Nothing sent here is published, and the answer comes back to the address you leave.