Skip to Main Content

Bears

Built at OpenEnv Hackathon SF · Mar 7, 2026 · San Francisco, CA

Bears — Demo video

Our project is an environment designed for an agent to oversee 7 agents playing a popular strategy game Diplomacy. Players in this game can have discussions, collude, and betray each other using complex, long-horizon strategies: reasoning that is hard to interpret from the observation space of an outside agent. The overseer learns via RL, rewarded whenever it correctly infers a player's hidden strategy, determined by an LLM acting as a judge. Our agent powers alternative interpretability of LLMs: Instead of focusing on reading model weights and activations, we learn intent from the agent's actions. As multi-agent systems become standard infrastructure, we need an oversight mechanism that can monitor many agents simultaneously and detect nuanced and hidden misalignment. Our overseer is exactly this, a dedicated monitoring model trained specifically to infer strategy in a complex multi-agent environment.

Team