# eBBD

- **Event:** [Built with Opus 4.7: a Claude Code hackathon](https://cerebralvalley.ai/e/built-with-4-7-hackathon)
- **When:** Apr 21 at 12:00 PM – Apr 27 at 2:00 AM (EDT)
- **Where:** Online
- **Team:** [Enam Biswas](https://cerebralvalley.ai/u/ebis)
- **GitHub:** https://github.com/e-biswas/paper-trail
- **Demo video:** https://youtu.be/j_11iV3FKdw
- **Gallery:** https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery/24

Paper Trail is a Claude-powered agent and web dashboard that investigates why an ML paper's published result fails to reproduce from its public repo. It identifies the root cause, applies the minimal fix, re-runs the evaluations, and opens a real GitHub PR with an evidence-backed scientific dossier. It targets the boring, expensive layer of the reproducibility crisis: data leakage, train/test contamination, label-aware preprocessing, and metric-implementation drift; the failures that quietly inflate headline numbers and rarely get caught in review.

The product runs in two modes. Deep Investigation is the autonomous flow: paste a paper URL and a repo URL then watch the agent generate ranked hypotheses, run discriminating checks via tool use, converge on a verdict with cited evidence, patch the code, and post a before/after metric delta. Quick Check is a chat-style sidebar where a researcher asks targeted, bounded questions ("is this imputation fit on train only?") and gets a verdict (confirmed OR refuted OR unclear) with at least one code citation in under 30 seconds. 

Together they reframe the agent from "autonomous scientist" to "verification intern"; one a researcher can actually trust to do the unglamorous audit work.

---

Markdown version of https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery/24. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
