# Dima

- **Event:** [Built with Opus 4.7: a Claude Code hackathon](https://cerebralvalley.ai/e/built-with-4-7-hackathon)
- **When:** Apr 21 at 12:00 PM – Apr 27 at 2:00 AM (EDT)
- **Where:** Online
- **Team:** [Dmitry Golovchits](https://cerebralvalley.ai/u/golovchits)
- **GitHub:** https://github.com/golovchits/RunItBack
- **Demo video:** https://drive.google.com/file/d/1c0yJFpPJjAiKCjmJxmjYwY-jLp1eLEur/view?usp=sharing
- **Gallery:** https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery/98

Even top ML PhDs cap out at a 41.4% replication score after 48 hours of dedicated effort per paper (PaperBench, OpenAI, 2025). Full replication is too slow and too hard, even for experts — which is exactly why static diagnosis has to be the first line of defense. The bottleneck isn't compute, it's diagnosis. The bug — an augmentation leak, a train-test overlap, a metric mismatch, a missing `model.eval()` — was almost always a static read away.

RunItBack does the static read. Point it at a paper, a repo, and some data; four Opus 4.7 agents (Paper Analyst, Code & Data Auditor, Validator, Reviewer) running on Claude Managed Agents return a diagnostic report with a verdict, claim-by-claim verification, severity-ranked findings, and unified-diff fixes — in minutes, before any GPU starts.

Opus 4.7 ingests PDFs natively (tables, figures, equations — no OCR, no `pdftotext`), carries the full reproducibility-failure taxonomy across 60-turn tool-using sessions, and enforces a ≥ 2-agent cross-check rule that surfaces genuine disagreements instead of hallucinating consensus. Managed Agents provides the sandbox, tools, sessions, and streaming on day one — RunItBack adds only the brain.

---

Markdown version of https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery/98. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
