Kekule
Built at Built with Opus 4.6: a Claude Code hackathon · Feb 10, 2026

I agent benchmarks today are static: run agents, score results, manually tweak prompts, repeat. Researchers spend more time diagnosing failures and hand-tuning strategies than running experiments. The feedback loop between "what went wrong" and "what to try next" is entirely manual -- and it doesn't scale. Kekule is a Claude Code skill that closes this loop. You define a task set and point kekule-bench at it. A swarm of Claude agents solves tasks, then the system runs its own retrospective: a failure analyst diagnoses why each fix broke, a coordinator distills generic lessons and invents new verification strategies, and the next epoch runs with a rewritten playbook. The verification engine evolves autonomously. In our first experiment, the coordinator invented 5 oracle strategies that didn't exist at startup. We targeted SWE-bench tasks the current SOTA couldn't solve and resolved 2.