# Anees Ahmed

- **Event:** [Built with Claude: Life Sciences](https://cerebralvalley.ai/e/built-with-claude-life-sciences)
- **When:** Jul 7 at 12:00 PM – Jul 14 at 12:00 AM (EDT)
- **Where:** Online
- **Team:** [Anees Ahmed Mahaboob Ali](https://cerebralvalley.ai/u/aneesahmed)
- **GitHub:** https://github.com/ahmedanees-m/mmc
- **Demo video:** https://youtu.be/28YprpGq8Hs
- **Gallery:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/116

MMC - the Mechanistic Model Compiler
AI models that predict what happens when you switch off a gene have a problem: the good ones are black boxes, and they still barely beat a linear baseline (Ahlmann-Eltze, Nature Methods 2025).

So we asked a different question. Not only that, can Claude predict better? But can Claude build a model that a biologist can actually open up, poke, and know when to trust?
What we built: MMC is a Claude-driven loop that reads the newest human T-cell atlas (Zhu 2025, 22 million cells). Claude proposes a wiring diagram of how genes control one another, compiles it into a runnable simulation, tests it against real experiments, reads where it fails, reasons about why it fails, and rewrites its own wiring. The output is not a number. It's a circuit you can interrogate.

The moment to watch. In our Th2 circuit, knockdown of GATA3 → IL-5 drops by 5.0. Knock down GATA3 and STAT6 together → IL5 drops by 5.0 again. Not more. Because STAT6 acts through GATA3, hitting both changes nothing. The additive baseline says −5.4. It's wrong, and it can't tell you why. That's mechanistic reasoning a black box structurally cannot do.

What we found, and it's the point. Claude proposed a new hypothesis: STK11 represses chemokines. It looked good. It was grounded, the data genuinely support that edge, and it passes the same interpretability checks that textbook biology passes.
But it didn't work. Across 76 hypotheses, in every regime we tested, single-gene knockdowns and combinatorial double-knockouts grounded mechanisms were never once incorporated into a model that beat a simple linear baseline (0 of 76; 95% CI [0, 4.8%]). MMC's own held-out test refused to certify STK11. We killed our own best result.

Why does it matter? The scary failure mode of AI in science isn't hallucination. It's the opposite: hypotheses that are grounded, interpretable, and pass every check we usually rely on and still don't predict. Interpretability didn't catch it. Plausibility didn't catch it. Only the held-out performance against a strong baseline did.

We're shipping the method, an interrogable T-cell circuit, and a map of exactly when mechanistic modelling helps and when it doesn't, a referee for a field where prediction is either stuck or bought with proprietary data.
Who this is for, 
Biologists get a T-cell circuit they can open, poke, and argue with, and a map of exactly which regulatory questions mechanistic models can answer today, so they stop spending months where the data can't support them.
AI builders get the harder lesson: the model's hypotheses can be grounded, interpretable, and pass every sanity check you have and still be worthless. Only held-out performance against a strong baseline catches it. Ship that gate, or ship a plausible lie.
And next, the loop is general. Point it at any system with a simulator and intervention data, and now we know exactly which regimes are worth pointing it at, because we measured where mechanism earns its keep and where it doesn't.

---

Markdown version of https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/116. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
