Skip to Main Content
Finalist

SolverFit Labs

Built at WeaveHacks 3: Self-Improving Agents Hackathon with Weights & Biases · Jan 31, 2026 · San Francisco, CA

Demo video · www.canva.com/…

Frontier LLMs are generalists trying to squeeze every benchmark. We showcase a deployed application of "Learning to Discover at Test Time" (Yuksekgonul et al. 2026) which uses RL to overfit an LLM to maximise a metric for a particular task. We demonstrate that GPT-oss-120B using our TTT harness can outperform Claude Code and even the best humans when it comes to cuda kernel wrtiing.

Team