Sanjana T.
Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote
Labs don't sample variants randomly. They measure hotspots, clinically observed positions, the ones the models argue about, and then compute a predictor's correlation on that set and treat it as that predictor's skill. We show that number is badly biased. On NUDT15 (Suiter et al. 2020), a non-random variant set understates a predictor's true correlation by 0.29 Spearman. Inverse-propensity weighting recovers it to within 0.009, a thirty-fold correction, verified to machine precision. We also settled the question we started with. We hypothesized that measuring where predictors disagree would identify the best predictor faster than random sampling. We pre-registered the test, committed it before writing the code, and it failed. We sharpened it, pre-registered again, and it failed again, because modern unsupervised predictors, benchmarked across ProteinGym, are statistically distinguishable but practically identical on any given assay. This corroborates Fawzy & Marsh (2024), who found per-gene predictor rankings largely uninformative. We add that the meaningful unit is protein × assay. So: if you've measured non-randomly, here's your correction. If you're choosing what to measure next, randomize.