Bonsai
Built at AI Engineer World's Fair Hackathon 2026 · Jun 27, 2026 · San Francisco, CA

Every domain expert who owns "what good looks like" — a compliance lead, a security reviewer — is drowning, hand-checking AI answers one at a time, and the ones that bite are the confident answers whose citation doesn't actually hold. Bonsai turns those experts into bonsai gardeners. They shape the ideal form once — a small, frozen standard — and the harness does the patient tending: it watches a cited-answer agent (Gemini 3.5 via Vertex AI), catches every claim whose citation doesn't trace to its source, clusters that failure by meaning (MongoDB Atlas $vectorSearch over Voyage embeddings), grows a new branch — one general check that catches that family of mistakes — and prunes the checks that overfit or go stale. Any real decision turns on dozens of factors; the rubric captures that accumulated judgment branch by branch, instead of one brittle hand-written test. And the growth stays honest: a build-time test fails CI if the loop ever reads the gardener's frozen standard (a guardrail, not an unbreakable wall). On that held-out set the loop can't read, agreement rose 9→14 of 15 (Wilson 95% CI [70.2%, 98.8%], sign-test p=0.031). The aim: one gardener, many trees.