# Bears

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Team:** [Masaki Tanaka Allwardt](https://cerebralvalley.ai/u/Masaki1), [Ronok Tanvir](https://cerebralvalley.ai/u/ronoktanvir), [Aryan Gupta](https://cerebralvalley.ai/u/Guptaman)
- **GitHub:** https://github.com/ronoktanvir/overseer
- **Website:** https://colab.research.google.com/drive/1sw6rxqgt15AEkJULX352wnD8tNMLoQ5h?usp=sharing
- **Demo video:** https://youtu.be/yfn8JjM45n8
- **Hugging Face:** https://huggingface.co/spaces/ronoktanvir/overseer-openenv
- **Gallery:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/67

Our project is an environment designed for an agent to oversee 7 agents playing a popular strategy game Diplomacy. Players in this game can have discussions, collude, and betray each other using complex, long-horizon strategies: reasoning that is hard to interpret from the observation space of an outside agent.

The overseer learns via RL, rewarded whenever it correctly infers a player's hidden strategy, determined by an LLM acting as a judge. Our agent powers alternative interpretability of LLMs: Instead of focusing on reading model weights and activations, we learn intent from the agent's actions.


As multi-agent systems become standard infrastructure, we need an oversight mechanism that can monitor many agents simultaneously and detect nuanced and hidden misalignment. Our overseer is exactly this, a dedicated monitoring model trained specifically to infer strategy in a complex multi-agent environment.

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/67. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
