Skip to Main Content

Shibboleth

Built at The Harness Engineering & Model Wrangling Hackathon · Sep 26, 2026 · New York, NY

Shibboleth — Demo video

Shibboleth is a harness trust layer for open-weight AI models. Anyone can download a model, quietly strip out its safety ("abliterate" it), and re-upload it looking completely normal, so testing its outputs won't reveal it. Shibboleth reads the model's internal activations instead: it fingerprints how strongly the model refuses harmful prompts at each layer, then compares that to a trusted baseline. The system runs entirely on MongoDB. A change stream fires a scan the moment a checkpoint lands, the verdict is written as a document, and $vectorSearch matches each fingerprint against the imposters already caught. That growing library is the flywheel: every catch makes the next scan sharper, so it's a recursive harness that improves its own guardrail as it runs. It's a live eval and gate you can put in front of a model hub or a CI pipeline, so you can trust an open-weight model before you run it. On the Qwen2.5-1.5B family it caught an abliterated model at an AUC of 1.0 while clearing a legitimate finetune.

Team