# AlphaWolf

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Team:** [Andres Nino](https://cerebralvalley.ai/u/andres), [Lex Hackett](https://cerebralvalley.ai/u/lexhacke)
- **GitHub:** https://github.com/lexhacke/alphawolf
- **Website:** https://colab.research.google.com/drive/1qJRHAZDny_Da0Sf_d_4DQh5cuhxOjniK?authuser=2#scrollTo=GXzZKpim8SoI
- **Demo video:** https://youtu.be/aL03D9GTj50
- **Hugging Face:** https://huggingface.co/spaces/detectivejoewest/alphawolf
- **Gallery:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/77

AlphaWolf is a self-play pipeline that teaches LLMs social deception through the game of Werewolf. A wolf agent plays, reflects on losses, rewinds to the critical decision point, and replays with an improved strategy. This approach resulted in Qwen3-8B-Thinking going from a 0.3 win rate to a 0.7 win rate over the span of only 45 games and 9 total gradient steps. The winning and losing transcripts become Direct Preference Optimization training pairs, while real-time consensus polling (every player votes on "who's suspicious?" after each speech) provides a novel auxiliary signal trained via ZIP-RC (ICLR 2026) to bootstrap a value function to the model. The result: a recursive self-improvement loop where the wolf gets better at deception, forcing villagers to adapt, driving emergent theory-of-mind.

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/77. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
