Skip to Main Content

Rulebook

Built at The Harness Engineering & Model Wrangling Hackathon · Sep 26, 2026 · New York, NY

Rulebook — Demo video

Rulebook — the harness that writes its own rulebook, from the one you already wrote by hand. Every team using a coding agent has been building a recursive harness by hand: a CLAUDE.md or AGENTS.md full of dated rules, each one earned by something that went wrong. Those rules came through blood, sweat and tears, so throwing them away to start an "automated" harness would be silly. Rulebook takes your manual labor, turns it into a system, and builds on it automatically from then on. How it works, all inside MongoDB Atlas: paste a rules file and a stream processor splits it into sections and calls a model from an $https stage to extract one lesson per rule (what was learned, what happened, how often it recurred, whether a mechanism was built) — the extractor is a pipeline definition, not a service. Nothing of ours runs between the file and the lessons. Automated Embedding embeds every lesson on insert; $vectorSearch finds what a team keeps re-learning; the count is the hard metric. A rule is proposed with the lessons that earned it as receipts, and nothing enters the rulebook until a human pushes the button. The rulebook renders back out as the agent's CLAUDE.md. When a rule keeps failing while followed, the harness escalates: it proposes a mechanism instead of another sentence. Seeded with 105 lessons distilled from two months of one team's real rules file. Two measured findings from today: embeddings group lessons by subject, not by idea — the same subject learned again is found by similarity (one group of eight), while the same idea across subjects is found by a rule's own wording; and every extracted lesson carries its own model cost, read from the response, so the price of a rulebook is a query. I wanted to tackle Statement Two, but I needed to tackle Statement One so I could build the memory layer for Statement Two. A long-horizon agent's context is the rules its record earned, not its transcript, and recurrence is the metric that says when memory failed. I solved the first on day one, and would've solved the second if there was a day two.

Team