Skip to Main Content

Team Coincidence

Built at Built with Claude: Life Sciences · Jul 7, 2026 · Remote

Team Coincidence — Demo video

Proteins are spelled in a twenty-letter alphabet, and every one of those letters is an ordinary Roman letter. So English words — HEALTH, ELVIS, VEGAN — fall out of human protein sequences by pure chance. The question isn't whether they exist. It's whether they land anywhere that matters. **The finding.** This is not an absence of a result. Words *avoid* the parts of a protein that matter, in a specific and measurable direction, and the cause is letter frequency. Right on a protein's active site — the tiny pocket where it does its chemistry — you find only 82% of the words chance would put there. Across 10,000 composition-preserving shuffles of the reviewed human proteome, not one placed as few words on active and binding residues as reality does: a depletion of −7.02σ, holding across all 27 languages tested. The cause is compositional and unmystical: functional sites are built from cysteine, histidine and aspartate; common words from A, L, S and E. Same alphabet, opposite frequencies — so the stretches that spell words and the stretches that do chemistry are drawn from nearly disjoint pools. Words also show no thematic match to their host genes: HEART sits in a ubiquitin hydrolase, LIVER in a vesicle protein. The pipeline goes further than the headline. Words are traced through Ensembl orthologs across a billion years — six survive intact all the way back to yeast, sitting in proteins the cell cannot live without. They're mapped to their genes' true GRCh38 coordinates (the gene-rich X spells 95 words; the gene-poor Y, just 8). And they're tracked through ClinVar variants and splice isoforms, where a documented disease mutation writes a word the reference proteome never spells. **You can walk it yourself via a standalone developed game.** The explorer is a single offline HTML file: type your name, or any word, and travel the protein it hides in — residue by residue through the real sequence, with the functional sites marked. Type SLUG and you land on the catalytic selenocysteine of glutathione peroxidase, the very atom the enzyme uses to protect your cells from damage. Type CARAVAN and you're in a proteasome subunit, unchanged for a billion years, back to yeast. Type JACOB and you learn it can never exist in any protein — J and O aren't amino acids. Five of our letters aren't, and between them they rule out 53% of English. **Every objection is tested, not asserted.** Counts are compared against composition-preserving permutations run within each protein, so uneven annotation cannot manufacture the signal; the depletion survives restriction to structurally ordered residues (AlphaFold pLDDT ≥ 70); and featured words are ranked by a hashed, tamper-checked scoring rule, so nothing can be quietly tuned after the fact. The full threats-to-validity analysis is in the manuscript. **Why it matters.** First, it's a null result reported plainly, with every threat to validity tested rather than buried — confirming that a coincidence really is a coincidence, rigorously, is a finding worth publishing. Second, it's useful: a lexicon of strings that look like biological signal and provably carry none is a calibrated negative control for short-linear-motif discovery, a field chronically vulnerable to false positives. Third, it's an educational tool: sequence, function and chance become something a non-scientist can hold in ten seconds. Genetics rarely offers a handle this tangible — a word, sitting on a binding groove, in a protein that is really inside you. Everything is open — pipeline, result tables, manuscript, and the explorer. Reproducible from public data with fixed seeds; every figure regenerated from tables.

Team