The CD4 Perturbation Predictability Audit Team
Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote

Everyone builds bigger models to predict what a gene knockdown does. I asked the prior question: how much is even predictable? Using a genome-scale CRISPRi Perturb-seq screen in primary human CD4⁺ T cells (~22M cells, 4 donors), I ran seven pre-registered investigations, each gated on CPU to fail cheaply before spending GPU. Six failed — and the failures are the finding. Under honest, reliability-ceiling-calibrated measurement, the per-perturbation signal sits at the noise floor: only ~17% of knockdowns have an effect you can trust (SNR > 3), and reaching reliability would need ~12× more sequencing depth. The one method that worked — a causal do-operator (C2 +0.118/+0.162) — beats its twin only within-distribution; on external causal data it's correlational, not causal. I wrapped all seven probes into a reliability-ceiling-calibrated predictability audit: a model-agnostic pre-flight scorecard, anchored by a positive control that proves it detects signal when present. Why it matters: perturbation models exist to prioritize expensive bench experiments. This audit tells a lab what its data can and can't support — before a single experiment runs.