# Kekule

- **Event:** [Built with Opus 4.6: a Claude Code hackathon](https://cerebralvalley.ai/e/claude-code-hackathon)
- **When:** Feb 10 at 12:00 PM – Feb 17 at 10:00 AM (EST)
- **Where:** Location TBA
- **Team:** [Ansh Tulsyan](https://cerebralvalley.ai/u/ansht)
- **GitHub:** https://github.com/archi-max/kekule/tree/data
- **Demo video:** https://youtu.be/N3V0VAV1k4E
- **Gallery:** https://cerebralvalley.ai/e/claude-code-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/claude-code-hackathon/hackathon/gallery/190

I agent benchmarks today are static: run agents, score results, manually tweak prompts, repeat. Researchers spend more time diagnosing failures and hand-tuning strategies than running experiments. The feedback loop between "what went wrong" and "what to try next" is entirely manual -- and it doesn't scale.

Kekule is a Claude Code skill that closes this loop. You define a task set and point kekule-bench at it. A swarm of Claude agents solves tasks, then the system runs its own retrospective: a failure analyst diagnoses why each fix broke, a coordinator distills generic lessons and invents new verification strategies, and the next epoch runs with a rewritten playbook. 

The verification engine evolves autonomously. In our first experiment, the coordinator invented 5 oracle strategies that didn't exist at startup. We targeted SWE-bench tasks the current SOTA couldn't solve and resolved 2.

---

Markdown version of https://cerebralvalley.ai/e/claude-code-hackathon/hackathon/gallery/190. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
