Skip to Main Content

eBBD

Built at Built with Opus 4.7: a Claude Code hackathon · Apr 21, 2026 · Remote

eBBD — Demo video

Paper Trail is a Claude-powered agent and web dashboard that investigates why an ML paper's published result fails to reproduce from its public repo. It identifies the root cause, applies the minimal fix, re-runs the evaluations, and opens a real GitHub PR with an evidence-backed scientific dossier. It targets the boring, expensive layer of the reproducibility crisis: data leakage, train/test contamination, label-aware preprocessing, and metric-implementation drift; the failures that quietly inflate headline numbers and rarely get caught in review. The product runs in two modes. Deep Investigation is the autonomous flow: paste a paper URL and a repo URL then watch the agent generate ranked hypotheses, run discriminating checks via tool use, converge on a verdict with cited evidence, patch the code, and post a before/after metric delta. Quick Check is a chat-style sidebar where a researcher asks targeted, bounded questions ("is this imputation fit on train only?") and gets a verdict (confirmed OR refuted OR unclear) with at least one code citation in under 30 seconds. Together they reframe the agent from "autonomous scientist" to "verification intern"; one a researcher can actually trust to do the unglamorous audit work.

Team