Verification tooling for AI agents.
AI coding agents report work as finished when it isn’t — and they do it fluently enough that you can’t tell by reading. A plausible narration and verified work look identical at the moment you read them. The cost arrives later.
This is not fabrication. The agent’s account is its honest best guess. The problem is that nothing sits between its fluency and your trust. I build the layer that goes there, and release it openly.
Rules and drop-in configs that stop coding agents from reporting work as finished when it isn’t — and from quietly doing something other than what you asked. Includes a model-specific failure taxonomy and verification gates shaped to match each one, for Claude Code and Codex.
Three ideas the tooling is built on:
These come from running several frontier models daily on real work and keeping records of what went wrong — not from benchmarks, which measure success rates rather than failure shapes.