Skip to Main Content

Proofread

Built at The Harness Engineering & Model Wrangling Hackathon · Sep 26, 2026 · New York, NY

Proofread — Demo video

Goodhart's law says that when a measure becomes a target, it stops being a good measure. Language models obey this pretty well; as ImpossibleBench documents, models edit tests and tamper with evaluation setups to create the illusion of improvement. Proofread stops this at the root. Every change a recursive harness makes to itself has to show two things. One, that it is allowed, which is verified formally against Lean 4 policies for every action the agent takes. And two, that it actually helps, which is measured on real tasks against a statistical bar fixed in advance. This bifold technique makes self-improvement trustworthy. The harness can rewrite its own prompts and workflow, but never the rules it is measured against.

Team