Skip to Main Content

n = 3, and Proud

Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote

n = 3, and Proud — Demo video

Everyone building research agents runs into the same quiet question: once you have one, how should you set it up to do its best work? Usually the answer is a hunch. This project turned the hunch into an experiment. A researcher used one AI research agent to study how to make that same agent produce better science — running a blank coding agent as a driver in a closed, self-improving loop, and pitting each loop design against a raw, unscaffolded baseline on word-identical questions, scored blind by an automated reviewer panel and a human domain expert. Across three rounds, the matured design scored highest on every metric and for every scorer, and beat its predecessor on all three test questions. The effect is modest and the sample is small — three questions, and one open-ended question the baseline won — and the report says so plainly. The lasting contribution is not the winning configuration; it is a reusable, auditable way to measure whether an agent's setup actually helps, rather than asserting that it does, paving the way for individual harness and workflow research to produce better science.

Team