# Anuj Dev Singh

- **Event:** [Built with Claude: Life Sciences](https://cerebralvalley.ai/e/built-with-claude-life-sciences)
- **When:** Jul 7 at 12:00 PM – Jul 14 at 12:00 AM (EDT)
- **Where:** Online
- **Team:** [Anuj dev Singh](https://cerebralvalley.ai/u/anujdevsingh)
- **GitHub:** https://github.com/anujdevsingh/regulatory-direction-independence
- **Demo video:** https://youtu.be/pgvrCK41Xu8?si=5bRzg5qvzTGyiBhV
- **Gallery:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/45

I tested a specific, high-stakes question for disease genetics: can today's best sequence-to-function deep-learning models (AlphaGenome, Borzoi, and the DeepSEA/Basset lineage) predict the DIRECTION — up or down — of a noncoding variant's regulatory effect in a defined cell type? This is the quantity a variant-interpretation pipeline actually needs, yet it is rarely benchmarked directly.

I assembled a chromosome-split benchmark of noncoding variants with laboratory-measured effect direction across nine cell contexts (6,377 variant×context measurements), plus a native human-microglia chromatin-accessibility QTL map from 95 donors as an on-target test. On the fairest possible test — accessibility direction in the native cell type — AlphaGenome (0.537), Borzoi (0.551), motif-PWM floor baselines, and models trained in-distribution on the caQTL data itself were ALL statistically indistinguishable from chance (majority-class 0.565).

I then found the mechanism. In the same 95 donors, a variant's effect on chromatin accessibility and its effect on gene expression are statistically INDEPENDENT (sign concordance 0.506, 95% CI 0.455–0.556, p=0.87). A model that reads one functional layer cannot recover the direction of a layer it does not read. The Alzheimer's risk variant rs6733839 (upstream of BIN1) anchors it: the risk allele opens chromatin AND raises BIN1 expression, yet represses episomal enhancer activity — three "directions" at once. A Boltz-2 co-fold shows a tighter MEF2A–DNA interface at the risk allele (+12.5% contacts).

Why it matters: this is a well-powered NEGATIVE result with a mechanism. It tells the field that directional predictions from single-layer sequence models should not be trusted for variant interpretation — and explains precisely why. I release the harmonized benchmark, the paired-layer independence statistic, and all per-variant scores as a reusable community resource.

---

Markdown version of https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/45. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
