Skip to Main Content

Dima

Built at Built with Opus 4.7: a Claude Code hackathon · Apr 21, 2026 · Remote

Dima — Demo video

Even top ML PhDs cap out at a 41.4% replication score after 48 hours of dedicated effort per paper (PaperBench, OpenAI, 2025). Full replication is too slow and too hard, even for experts — which is exactly why static diagnosis has to be the first line of defense. The bottleneck isn't compute, it's diagnosis. The bug — an augmentation leak, a train-test overlap, a metric mismatch, a missing `model.eval()` — was almost always a static read away. RunItBack does the static read. Point it at a paper, a repo, and some data; four Opus 4.7 agents (Paper Analyst, Code & Data Auditor, Validator, Reviewer) running on Claude Managed Agents return a diagnostic report with a verdict, claim-by-claim verification, severity-ranked findings, and unified-diff fixes — in minutes, before any GPU starts. Opus 4.7 ingests PDFs natively (tables, figures, equations — no OCR, no `pdftotext`), carries the full reproducibility-failure taxonomy across 60-turn tool-using sessions, and enforces a ≥ 2-agent cross-check rule that surfaces genuine disagreements instead of hallucinating consensus. Managed Agents provides the sandbox, tools, sessions, and streaming on day one — RunItBack adds only the brain.

Team