The Nineteen
Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote

# Project description ## What we built / investigated We mined a genome-scale CRISPRi Perturb-seq atlas of about 22 million primary human CD4⁺ T cells to discover **novel drug targets**. (Per the project's dataset documentation, the atlas is the CD4⁺ T-cell Perturb-seq resource from the Marson and Pritchard labs, distributed on the CZI Virtual Cells Platform; this attribution is as recorded in the project files and has not been independently re-verified here.) The core idea is that a CRISPRi knockdown is a genetic model of drug-induced loss of function, so we read the perturbation-to-transcriptome map backwards: starting from a therapeutically desirable T-cell program and asking which upstream regulators control it. Those regulators become drug-target hypotheses. On top of this we built an end-to-end computational-to-translational pipeline: foundation QC and reproducibility filtering, then six orthogonal discovery lenses, then an integrated ranked scorecard, then single-cell validation across donors, and finally structure, genetics, and chemistry-based translational dossiers. ## What we found The pipeline narrowed roughly 34,000 perturbations down to a short list of **19 single-cell-validated targets**, and this is the main result (full detail in the Appendix table). It splits into two therapeutic groups. **A TCR-proximal signaling group** (LAT, PLCG1, CD247, CD3E, ZAP70, VAV1, SIK3) contains the highest-ranked hits and mostly points toward suppressing inflammation. It recovers known immune-drug biology on its own, which validates the logic, and it surfaces **SIK3**, a novel kinase that became our structurally validated lead (redocking under 2 Å, actives-versus-decoys enrichment AUC 0.905). **A drug-naive chromatin and transcriptional group** (a SAGA/Mediator axis plus remodelers: SMARCE1, STAT6, SGF29, MED24, MED12, TADA2B, NSD1, CHD4, SMARCB1, TRIP12, ARNT, SEL1L) is the central novelty. All 12 of these validate at single-cell resolution and form a mechanism distinct from acute TCR signaling (89% cluster recovery, 2.4× smaller state-space displacement). Of the full set of 19, **18 are clinically unprecedented** (only CD3E is already drugged). The single top opportunity is **STAT6** (rank 5): a boost-immunity target backed by strong asthma and allergy genetics (55 immune GWAS associations), rich chemical matter (552 ChEMBL bioactivities, max pChEMBL 9.15), and small-molecule tractability, yet with no approved drug. The remaining standouts are the drug-naive SAGA/Mediator readers **MED24**, **SGF29**, and **TADA2B**. Two targets (CHD4, SMARCB1) are flagged for essentiality and toxicity risk on weaker knockdown. The complete ranked list, with validation and translational evidence for each target, is given in the Appendix. ## Why it matters By recovering known targets from scratch, which validates the logic, and then nominating a genetically-supported, single-cell-validated, structurally-tractable set of *novel* regulators, the project turns a 22-million-cell atlas into a short list of de-risked, mechanistically-distinct therapeutic entry points for autoimmune and allergic disease and for immuno-oncology. STAT6 and a previously undrugged chromatin/transcriptional axis head that list. More broadly, the work demonstrates a reusable in-silico funnel that goes from genome-scale perturbation data to prioritized, translation-ready drug targets. --- ## Appendix: the 19 validated targets Sorted by nomination rank, with validation and translational evidence for each target. | Rank | Gene | Axis | Direction | Novelty | Dirs | KD% | Concord. r | Donors | Best PDB (n) | SM bucket | ChEMBL act. | max pChEMBL | Immune GWAS | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | 1 | **LAT** | TCR-proximal | Suppress | novel-druggable | 4 | 85% | 0.67 | 4/4 | — | — | 2 | — | — | | 2 | **SMARCE1** | Chromatin/TF | Suppress | novel-druggable | 3 | 88% | 0.74 | 3/3 | 9 | Structure with Ligand | 8 | 6.95 | 37 | | 3 | **PLCG1** | TCR-proximal | Suppress | novel-druggable | 4 | 88% | 0.46 | 4/4 | 6 | High-Quality Ligand | 438 | 6.75 | — | | 4 | **CD247** | TCR-proximal | Suppress | novel-druggable | 3 | 72% | 0.51 | 4/4 | 38 | Structure with Ligand | — | — | 46 | | 5 | **STAT6** | Chromatin/TF | Boost | novel-druggable | 3 | 82% | 0.57 | 4/4 | 7 | High-Quality Ligand | 552 | 9.15 | 55 | | 6 | **CD3E** | TCR-proximal | Suppress | known-drug-target | 3 | 96% | 0.85 | 4/4 | 44 | Med-Quality Pocket | — | — | 3 | | 7 | **ZAP70** | TCR-proximal | Suppress | novel-druggable | 3 | 91% | 0.78 | 3/3 | 15 | High-Quality Pocket | 2390 | 8.1 | 1 | | 9 | **VAV1** | TCR-proximal | Suppress | novel-druggable | 2 | 92% | 0.66 | 3/3 | 10 | Structure with Ligand | 1 | 8.98 | — | | 19 | **SIK3** | TCR-proximal | Suppress | novel-druggable | 3 | 87% | 0.52 | 3/3 | 5 | High-Quality Ligand | 809 | 9.34 | — | | 22 | **TRIP12** | Chromatin/TF | Suppress | novel-druggable | 3 | 84% | 0.70 | 4/4 | 5 | Structure with Ligand | — | — | — | | 25 | **NSD1** | Chromatin/TF | Suppress | novel-druggable | 3 | 92% | 0.65 | 3/3 | 4 | High-Quality Ligand | 147 | 6.96 | 1 | | 31 | **SGF29** | Chromatin/TF | Suppress | novel-druggable | 3 | 94% | 0.69 | 4/4 | 8 | High-Quality Pocket | 1 | — | 5 | | 34 | **MED24** | Chromatin/TF | Mixed | novel-druggable | 3 | 96% | 0.80 | 4/4 | 10 | Structure with Ligand | — | — | 12 | | 44 | **ARNT** | Chromatin/TF | Mixed | novel-druggable | 3 | 89% | 0.69 | 4/4 | 46 | High-Quality Ligand | 25 | — | 1 | | 45 | **TADA2B** | Chromatin/TF | Suppress | difficult | 3 | 92% | 0.74 | 4/4 | — | — | — | — | — | | 66 | **CHD4** | Chromatin/TF | Boost | novel-druggable | 3 | 52% | 0.38 | 4/4 | 12 | Structure with Ligand | 270 | 8.52 | — | | 80 | **SMARCB1** | Chromatin/TF | Suppress | novel-druggable | 3 | 48% | 0.59 | 4/4 | 18 | Structure with Ligand | 8 | 8.03 | — | | 99 | **MED12** | Chromatin/TF | Suppress | difficult | 3 | 85% | 0.76 | 4/4 | 3 | — | 6 | 6.89 | — | | 124 | **SEL1L** | Chromatin/TF | Suppress | novel-druggable | 3 | 94% | 0.42 | 3/3 | 5 | — | — | — | — | Columns: **Axis** = TCR-proximal signaling vs the novel chromatin/transcriptional axis. **Direction** = therapeutic direction implied by loss of function (Suppress inflammation / Boost immunity / Mixed). **Dirs** = number of independent discovery directions that nominated the gene. **KD%** = single-cell on-target knockdown (mean across powered donors). **Concord. r** = Pearson concordance between the single-cell knockdown signature and the pseudobulk DE signal. **Donors** = powered donors out of those tested. **Best PDB (n)** = count of experimental structures. **SM bucket** = best Open Targets small-molecule tractability bucket. **ChEMBL act. / max pChEMBL** = chemical-matter depth and best measured potency. **Immune GWAS** = immune-disease GWAS associations. Dashes mark zero or no available data.