Skip to Main Content

4nY4

Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote

4nY4 — Demo video

Thirteen real BCG trials. Pooled risk ratio 0.49. AskBench refuses to call it one number. I² is 92.1%. No model in the loop. One command reproduces it. That is the product: a bench scientist asks their data a plain-English question and gets a verdict, not a confident paragraph. SOLID or FLAGGED, with the Skeptic's reason on top. Six cells with a strong effect? Still flagged. A combined maternal risk that would imply 673 per 1000 pregnancies? Refused. The stats come from a fixed Python toolkit; Claude reads messy questions and narrates the argument. It does not touch a p-value. We measured it. Two hundred seeds, planted traps, zero API credits: structural traps caught every time; statistical traps land in the low 90s; 1.58% false positives after Benjamini-Hochberg, reported in the README, not buried. Same Skeptic, new traps it was never tuned on: still catches them. Judges can rerun python3 eval.py and python3 real_data.py themselves. Shipped as a live demo and an MCP server, so Claude can call the Skeptic inside a session. Built this week with Claude Code.

Team