# Self-Taught Trainer

- **Event:** [The Harness Engineering & Model Wrangling Hackathon](https://cerebralvalley.ai/e/mongodb-nyc-hackathon)
- **When:** Sat, Sep 26 at 9:00 AM – 10:00 PM (EDT)
- **Where:** New York, NY
- **Team:** [Devin Dawson](https://cerebralvalley.ai/u/unclemusclez), [Ryan Sherman](https://cerebralvalley.ai/u/beginnerbot), [Eric Ko](https://cerebralvalley.ai/u/ERIC_EX), [Daniel Joo](https://cerebralvalley.ai/u/danielsjoo)
- **GitHub:** https://github.com/ERICEX2025/self-taught-trainer
- **Demo video:** https://www.loom.com/share/d8f1ff2b4a594a6485ac96e79f8cb40f
- **Gallery:** https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/31

Self-Taught Trainer is a self-improving harness for an LLM that plays Pokémon Showdown (Gen 1 OU). The player, GPT-5.4 mini, is never retrained. Instead, a coach AI studies its lost games and rewrites the harness around it: its rules, its instructions, and small Python tools it writes itself, which run in a sandbox. Every change is an A/B test: the new version and the current best play fresh games side by side, the top idea is re-tested head to head to rule out luck, and it is kept only if it wins by 5+ points.

Starting from one generic sentence with no Pokémon knowledge in the prompt, it taught itself Gen 1 quirks from its own losses (freeze is permanent, only one status at a time, a Hyper Beam that knocks out skips its recharge turn). Win rate against our training bots went from 52% to 71% (Run 1), and up to 80% in Run 2. Every harness version, battle, coach decision and lesson is a document in MongoDB Atlas. Lessons are auto-embedded with Voyage and recalled by vector search before each coach decision.

We also tested it honestly on the PokéAgent Challenge ladder (NeurIPS 2025) against opponents it never trained on. The gains didn't carry over yet (13% to 14%), which taught us the key lesson: a self-improving agent is only as good as its sparring partners. The dashboard shows the real Showdown replay next to the player's reasoning and the coach-written tool hints behind each move.

## More from The Harness Engineering & Model Wrangling Hackathon

- [Sudo Squad](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/28)
- [emirates](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/29)
- [Second Shift](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/30)
- [Crucible](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/32)
- [LILI](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/33)
- [Bisect](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/34)

---

Markdown version of https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/31. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
