# Shadi Shafighi

- **Event:** [Built with Claude: Life Sciences](https://cerebralvalley.ai/e/built-with-claude-life-sciences)
- **When:** Jul 7 at 12:00 PM – Jul 14 at 12:00 AM (EDT)
- **Where:** Online
- **Team:** [Shadi Shafighi](https://cerebralvalley.ai/u/dshshadi)
- **GitHub:** https://github.com/shafighi/mpra-estimand-audit.git
- **Demo video:** https://youtu.be/BUdYVQLRvGc?si=SklH8MIdYE11U27G
- **Gallery:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/195

### What I built and investigated

Regulatory-variant prediction models are commonly evaluated against massively parallel reporter assays, or MPRAs, as though the assay were direct ground truth. I built an estimand-aware benchmark audit that first asks a more fundamental question: do the model and experiment estimate the same biological quantity?

I applied this framework to 3,976 psychiatric-risk regulatory variants studied in the developing human cortex. I compared three sequence-to-function models—AlphaGenome, Borzoi, and Enformer—with a genomic language model, Evo 2; a developing-cortex lentiMPRA; and a fine-mapped fetal-cortex eQTL atlas used as an external endogenous anchor.

With Claude, I implemented custom Bayesian measurement-error and latent-variable models, coordinated model scoring, checked the analysis for reproducibility and overclaiming, and produced a calibrated per-variant concordance map. The latent agreement estimator was validated through simulation-based calibration.

### What I found

The three activity-predicting models showed strong agreement with one another, with rank correlations of approximately 0.55–0.81, but essentially no agreement with the lentiMPRA, with correlation around −0.02. Accounting for the reporter’s reported measurement uncertainty did not restore agreement, indicating that this uncertainty alone is insufficient to explain the gap.

Evo 2 did not join either group, showing that “sequence model” does not represent a single estimand.

The three activity models also showed modest agreement with the fine-mapped eQTL effects, with correlations of approximately 0.24–0.28, while the reporter showed essentially none. Among the 25 variants with the strongest fine-mapping support, model and eQTL directions agreed for 20.

### Why it matters

The conclusion is not that computational models beat experiments. It is that a precise experimental measurement may still be a misaligned benchmark when it targets a different biological quantity.

Here, the lentiMPRA measures an isolated DNA element in an ectopically integrated construct, while the models predict endogenous functional signals from native-locus sequence. Their disagreement is therefore consistent with a consequential estimand mismatch, although native-locus editing is still required to distinguish this explanation from model error and shared bias.

The result is not a model-versus-assay winner. It is a calibrated map showing where the measurements agree, where they diverge, and which native-locus experiments would be most informative next.

---

Markdown version of https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/195. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
