Method Note · The Fire Drill

The rules, locked before the score.
Then the measurement killed the test.

Killed 2026-08-27 · zero seeds ever ran

The plan below was to plant mistakes in the AI's work and score how often I catch them. Standing it up meant first measuring the stream we would be planting into, and that measurement ended the experiment: roughly half of my review turns already carry a correction, and the log had been keeping about one in twenty-seven of them. You don't seed a field that's already full.

The seeding math died on its own terms too. The planted odds put synthetic mistakes into real work often enough to be their own risk, for a number the natural stream hands over free.

So the drill is dead and the question lives: "does the human actually catch it" is now scored on the real mistakes instead of planted ones. It is done when the number exists and its interval excludes zero. Same finish line, no seeds.

Everything below stays published exactly as written, dated, because that is what a pre-registration is FOR. When the design dies, you watch it die in public. The alternative is a page that quietly forgets it made a promise.

The pre-registration as locked, 2026-08 · now a record, not a plan

The test: real mistakes are planted into real work, in secret, at random, and the human reviewer's catch rate is scored. Caught or missed, the result publishes.

Pre-registered, before any trial exists:

  • At least 40 seeded trials before any rate is quoted; the stop rule is fixed now.
  • 20% odds per session, at most one seed per session. The reviewer never knows which session carries one.
  • Selection is made by mechanism, not by a person, inside operator-consented ranges. The administrator is a sealed manifest, not someone who knows the answers.
  • The manifest is hash-committed at seed time and sealed until scoring. Once seeding starts, only manifest hashes publish, never contents.
  • Seeds touch internal surfaces only, with a tripwire that stops any seed from reaching outbound work.

Known limit, disclosed up front: planted mistakes are easier to catch than natural ones, and knowing a drill exists sharpens the reviewer. The measured rate is therefore an upper bound on real oversight, and it will be published as one.

Final state: zero seeds run, ever. The scoreboard stayed blank from lock to kill, which is the receipt working: a number that never existed is a number that could not have been picked.

Back to the So What

Read this before you believe any of it. The platform is a running prototype, not an accredited or certified system. The register, the war room, and the brain are receipts, not revenue. Every claim on this site links to something you can check. If one doesn't, tell us, and we'll fix the page.