# The CD4 Perturbation Predictability Audit Team

- **Event:** [Built with Claude: Life Sciences](https://cerebralvalley.ai/e/built-with-claude-life-sciences)
- **When:** Jul 7 at 12:00 PM – Jul 14 at 12:00 AM (EDT)
- **Where:** Online
- **Team:** [Lukas Weidener](https://cerebralvalley.ai/u/LSW)
- **GitHub:** https://github.com/VibeCodingScientist/CD4-Perturbation-Predictability-Audit
- **Demo video:** https://youtu.be/AO0A4UhhMZE
- **Gallery:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/62

Everyone builds bigger models to predict what a gene knockdown does. I asked the prior question: how much is even predictable? Using a genome-scale CRISPRi Perturb-seq screen in primary human CD4⁺ T cells (~22M cells, 4 donors), I ran seven pre-registered investigations, each gated on CPU to fail cheaply before spending GPU.

Six failed — and the failures are the finding. Under honest, reliability-ceiling-calibrated measurement, the per-perturbation signal sits at the noise floor: only ~17% of knockdowns have an effect you can trust (SNR > 3), and reaching reliability would need ~12× more sequencing depth. The one method that worked — a causal do-operator (C2 +0.118/+0.162) — beats its twin only within-distribution; on external causal data it's correlational, not causal.

I wrapped all seven probes into a reliability-ceiling-calibrated predictability audit: a model-agnostic pre-flight scorecard, anchored by a positive control that proves it detects signal when present.

Why it matters: perturbation models exist to prioritize expensive bench experiments. This audit tells a lab what its data can and can't support — before a single experiment runs.

---

Markdown version of https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/62. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
