# 4nY4

- **Event:** [Built with Claude: Life Sciences](https://cerebralvalley.ai/e/built-with-claude-life-sciences)
- **When:** Jul 7 at 12:00 PM – Jul 14 at 12:00 AM (EDT)
- **Where:** Online
- **Team:** [Anya Chueayen](https://cerebralvalley.ai/u/4nY4)
- **GitHub:** https://github.com/anyapages/askbench, https://askbench-weld.vercel.app
- **Demo video:** https://youtu.be/heNUcS1DftI?si=hyB76aD4fNKJz7-U
- **Gallery:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/148

Thirteen real BCG trials. Pooled risk ratio 0.49. AskBench refuses to call it one number. I² is 92.1%. No model in the loop. One command reproduces it.

That is the product: a bench scientist asks their data a plain-English question and gets a verdict, not a confident paragraph. SOLID or FLAGGED, with the Skeptic's reason on top. Six cells with a strong effect? Still flagged. A combined maternal risk that would imply 673 per 1000 pregnancies? Refused. The stats come from a fixed Python toolkit; Claude reads messy questions and narrates the argument. It does not touch a p-value.

We measured it. Two hundred seeds, planted traps, zero API credits: structural traps caught every time; statistical traps land in the low 90s; 1.58% false positives after Benjamini-Hochberg, reported in the README, not buried. Same Skeptic, new traps it was never tuned on: still catches them. Judges can rerun python3 eval.py and python3 real_data.py themselves.

Shipped as a live demo and an MCP server, so Claude can call the Skeptic inside a session. Built this week with Claude Code.

---

Markdown version of https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/148. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
