SVBench AI
Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote

SVBench AI is a Claude verdict layer for structural-variant callers. One reviewer — Claude Opus 4.8 — reads each call's alignment image and genomic context and returns an evidence-cited verdict, with or without a truth set, corroborated against independent population catalogs (gnomAD-SV, dbVar, HGSVC). Against HG002 it showed 93% of Sniffles' and 71% of SVIM's "false positives" are benchmark artifacts, not caller errors (raw 0.886→0.992; 0.694→0.886) — and SVIM's corrected precision equals Sniffles' raw, exposing a gap the benchmark hid. On genomes with no truth, the same reviewer still delivers grounded verdicts. Trustworthy SV evaluation, everywhere.