Skip to Main Content

Hindsight

Built at The Harness Engineering & Model Wrangling Hackathon · Sep 26, 2026 · New York, NY

Hindsight — Demo video

Results snapshot (read-only): https://hindsight-xi-two.vercel.app Hindsight turns real coding-agent history into a long-horizon memory benchmark on MongoDB Atlas, then lets the agent's harness tune its own memory configuration. We load real Claude Code sessions from the SWE-chat dataset (Hugging Face) into Atlas. An LLM labels "moments" in the chat history: constraints, decisions, and facts the developer gave the agent, which a memory-enabled coding harness should remember. Another LLM turns these moments into eval cases, and we keep only the cases an agent fails without memory but passes with it. We use the open-source pi.dev coding harness with a memory service on top of Atlas. It searches both the mined moments and the raw chat history, using vector search and BM25 together in hybrid retrieval. Each eval case has a cutoff date, and the memory service filters out everything after that date inside Atlas, so the agent can't see the future. The eval cases are split into a dev set and a test set. The harness tunes itself using only the dev set, and final scores are reported on the test set. On the test set, pi without memory scored 4%, our hand-designed memory harness scored 41%, and the self-tuned harness scored 71%, while doubling its reliability (passing all 3 attempts) from 27% to 56%.

Team