# Ray

- **Event:** [Built with Claude: Life Sciences](https://cerebralvalley.ai/e/built-with-claude-life-sciences)
- **When:** Jul 7 at 12:00 PM – Jul 14 at 12:00 AM (EDT)
- **Where:** Online
- **Team:** [Rui Qiao](https://cerebralvalley.ai/u/rqiaorrr)
- **GitHub:** https://github.com/volpato30/lifescience-hackathon/blob/main/single_cell_best_practices.pdf
- **Demo video:** https://www.loom.com/share/ca3b199628a9446f8175bfc1073cbf0b
- **Gallery:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/125

Recent progress in mass spectrometers and softare drive protein identification from a single HeLa cell to more than 6000. However, how many false positives there are remains an open problem.  In this study we evaluated three DIA-NN 2.6.0 features that are on or encouraged by default and that all inflate identification counts — the refined q-value procedure, multi-species database searching for non-human samples, and match-between-runs (MBR) — asking in each case whether the extra identifications are trustworthy. The core method was a two-species entrapment search: the database held the sample's true species plus a foreign species known to be absent, so any precursor mapping only to the foreign proteome is a ground-truth false positive, giving an empirical FDR independent of the engine's own estimate. Spectral quality was measured directly from the raw Thermo files by mapping each precursor to its apex MS2 scan and counting matched b/y fragment ions at ±15 ppm (mass model validated to 0.06 ppm against DIA-NN's reported precursor m/z).
What we found

    Refined q-value trades peptide confidence for protein depth. Precursors rose +4.7% (36,989 → 38,736) but protein groups fell −5.4% (5,069 → 4,794) — the signature of protein-informed rescoring. The added precursors had ~30% fewer matched fragments (median 7 vs 10), though entrapment FDR stayed controlled (0.51% → 0.37%). Shared precursors kept identical RTs, so refined-q only recalibrates q-values, not peak picking.
    DIA-NN identification is systematically human-biased for non-human samples. In a mouse cell, 6.6% of precursors mapped only to human (~12% size-adjusted) vs ~0.4% for human HeLa. Across 14 mouse liver single cells the human false-positive rate was 5.05–6.18% (median 5.62%, SD just 0.32%) — reproducing to ±0.3% across independent cells proves it is systematic: weak SCP spectra preferentially match the larger, better-annotated human proteome.
    MBR's homogeneity assumption does not hold well for single cells. Before MBR, mean pairwise cell overlap was only Jaccard 0.54, and more precursors were cell-unique (28.4%) than universal (23.5%) — the "identified in one ⇒ present in others" premise is true less than half the time. MBR added +3–46% precursors (most in the sparsest cells) and lowered entrapment FDR, but second-pass-only precursors carried the same ~30% fragment deficit (median 7–8 vs 10–12). Post-MBR apparent homogeneity (24%→48% "present in all cells") is manufactured by the imputation itself. Crucially, entrapment FDR certifies the aggregate ID list, not whether a precursor is truly present in the specific cell it was transferred into.

Why it matters

All three features buy depth by admitting lower-confidence identifications, and all three are most hazardous precisely where SCP is most valuable — measuring real cell-to-cell biological differences. A low FDR, even a low empirical entrapment FDR, certifies the aggregate list, not any single per-cell measurement. Practical guidance: turn refined q-value off when confident peptide-level IDs are the goal (it also yields fewer proteins); for non-human data, don't trust the nominal 1% FDR — run an entrapment check, prefer species-specific databases, scrutinize human-annotated hits; for MBR, report both passes and treat transferred values as imputed for any per-cell biological claim, reserving MBR for building complete quantitative matrices for clustering/integration. The cross-cutting principle: report depth and confidence together, and validate rescored or transferred identifications before building single-cell conclusions on them.

---

Markdown version of https://cerebralvalley.ai/e/built-with-claude-life-sciences/hackathon/gallery/125. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
