# OpenEnv Hackathon SF: Project Gallery

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Hosts:** [Cerebral Valley](https://cerebralvalley.ai/u/cv), [PyTorch](https://cerebralvalley.ai/u/pytorch)
- **Projects:** 104 (6 placed)
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery

## Projects

### 1. VendEnv

VendSim VB2 is an OpenEnv 0.2.1 environment that drops an LLM agent into a 365-day vending machine business simulation. The agent starts with $500 and must maximize its bank balance by setting prices across 5 products, negotiating with adversarial suppliers, managing inventory through a delegated sub-agent, and surviving daily fees, weather-driven demand shifts, and seasonal fluctuations.



The environment features an MCP tool-calling interface with 16 tools spanning pricing, supplier negotiation, inventory management, and a scratchpad memory system. Rewards scale uncapped with agent performance — a skilled agent can grow net worth well beyond the starting capital, while poor decisions lead to bankruptcy. We train with GRPO using Unsloth on a Qwen2.5-1.5B model, demonstrating measurable improvements in pricing decisions and revenue generation over the training steps.

- **Team:** [Robert Amanfu](https://cerebralvalley.ai/u/ramanfu)
- **GitHub:** https://github.com/retroam/vendsim-vb2/tree/main
- **Website:** https://drive.google.com/file/d/1IsONrH5A7sR1ru7eKfc1qWdSRDx8ghhG/view?usp=sharing
- **Demo video:** https://youtu.be/m2OVtpXOZ4U
- **Hugging Face:** https://huggingface.co/spaces/retroam/vendsim-vb2
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=22

### 2. OpenSpaceEnv

The RANS paper ( https://arxiv.org/abs/2310.07393  ) was implemented and the OpenEnv interface wrapped around it.  It was used to fine tune a Qwen model to maneuver a spacecraft through LEO.

- **Team:** [Daniel Pang](https://cerebralvalley.ai/u/danp)
- **GitHub:** https://github.com/dpiresearch/meta_openenv/
- **Website:** https://github.com/dpiresearch/meta_openenv/blob/main/unsloth-qwen3-northflank/train.py
- **Demo video:** https://www.loom.com/share/fc5c3a63c6e14d46af0e6eb9f49482b4
- **Hugging Face:** https://huggingface.co/spaces/dpang/rans-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=23

### 3. AlphaForge

Enterprise QA Multi-Agent Environment is an adversarial reinforcement learning simulation where an AI "Taskmaster" generates complex, domain-specific business questions (finance, operations, real estate) and pits them against a "Challenger" LLM agent that must solve them. The system includes a live SEC document ingestion pipeline that parses real corporate filings (like Amazon's 10-Q) to continuously inject fresh, quantitative training tasks — keeping the agent from training on stale data.

- **Team:** [Krish t](https://cerebralvalley.ai/u/krish)
- **GitHub:** https://github.com/krishthukral/enterprise_qa
- **Website:** https://colab.research.google.com/drive/1jbN_Eug7QAjOK_wSRc-1g4_Q-RSJoW_A
- **Demo video:** https://youtu.be/7A-vok_x0TM
- **Hugging Face:** https://huggingface.co/spaces/Xxa1/enterprise-qa
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=24

### 4. Tech Raven

Building an environment to train both a Coordinating LLM and an LLM agent  to run on the Pollen Robotic's Reachy Mini. The LLM will help users with ADHD when they are "stuck" with low executive functioning and need help getting back on track

- **Team:** [Steven Pousty](https://cerebralvalley.ai/u/TheSteve0)
- **GitHub:** https://github.com/thesteve0/adhd-coach and https://huggingface.co/spaces/TheSteve0/adhd-env/tree/main
- **Website:** https://github.com/thesteve0/adhd-coach/blob/main/minimal-training.py
- **Demo video:** https://colab.research.google.com/drive/1XVDfxAARnF7WiZRNbvm57c_DtBlpPF3x?usp=sharing
- **Hugging Face:** https://huggingface.co/spaces/TheSteve0/adhd-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=25

### 5. WRL-Dragon

WRL-Dragon automates reinforcement learning policy development using a hierarchical multi-agent interaction system. 

Instead of having to manually write and tune RL policies, a CEO agent analyzes environments, Coder agents generate policy code using LLMs, and QA agents run rollouts to evaluate performance. 

Results feed back into subsequent round, creating an autonomous improve loop across multiple Gym environments simultaneously. The system solves the tedious cycle of write-test-iterate in RL research by letting AI agents handle the entire pipeline from strategy to code generation to evaluation, all while a real-time dashboard lets you watch the process unfold.

Beyond this orchestration loop, the system meta-trains smaller child agents directly on RL tasks, combining LLM-driven code generation with traditional RL training to produce increasingly capable policies!

- **Team:** [Farhan Navas](https://cerebralvalley.ai/u/farhan-navas), [Warren Low](https://cerebralvalley.ai/u/warren)
- **GitHub:** https://github.com/farhan-navas/wrl-dragon
- **Website:** https://colab.research.google.com/drive/1KWttxo_ZOAYuFJSpOFnFvCs9qOr-KL6O?usp=sharing
- **Demo video:** https://www.youtube.com/watch?v=DQ8LGomlc_8
- **Hugging Face:** https://huggingface.co/spaces/DESUCLUB/wrl-dragon-gym
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=26

### 6. Play-gent

Our agent uses OpenEnv to build a curriculum of video game environments that train a TinyLlama 1.1B agent via GRPO reinforcement learning. We implement three OpenEnv-compliant environments — Diplomacy (coalition tactics), webDiplomacy human gameplay (211k real states), and IRC poker (bluff detection) — each with dense reward signals that teach the core primitives of strategic negotiation. The agent learns through RL across this curriculum: Phase 1 trains coalition pressure in Diplomacy, Phase 2 grounds it in human gameplay patterns, Phase 3 unifies all three reward signals in a live arbitrage environment. The result is an agent that transfers video game negotiation skills to real economic interactions — starting with $20, it compounds capital to $80 (4x return) against Groq Llama 3.1 8B adversarial sellers, detecting bluffs at 97% confidence and reaching actual price floors 100% of the time.

- **Placement:** Finalist
- **Team:** [Abraham Bhatti](https://cerebralvalley.ai/u/abhatti)
- **GitHub:** https://github.com/AbeBhatti/Play-gent
- **Website:** https://colab.research.google.com/github/AbeBhatti/Play-gent/blob/main/training/arbitragent_colab.ipynb#scrollTo=FmGOks6oquiy
- **Demo video:** https://youtu.be/bsKZgyDsuDc
- **Hugging Face:** https://huggingface.co/spaces/Abeee32t/ArbitrAgent?logs=container
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=27

### 7. ORACLIN

Reinforced medical diagnosis reasoning using ICD-10 as a diagnosis scaffold.

- **Team:** [Patrick Damaso](https://cerebralvalley.ai/u/dataphysician), [Aaron Fanous](https://cerebralvalley.ai/u/gtcha2)
- **GitHub:** https://github.com/gtcha2/ORACLIN
- **Website:** https://github.com/gtcha2/ORACLIN/blob/icd-agents/notebooks/oraclin_train.ipynb
- **Demo video:** https://youtu.be/n1H8jH7pfX0
- **Hugging Face:** https://huggingface.co/spaces/mdgtcha2/Oraclin
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=28

### 8. Signall

10 OpenEnv cognitive primitives that train generalist agents. LLM fine-tuned via Unsloth + GRPO on live HuggingFace Space reward signals.

- **Team:** [Jon Yeazel](https://cerebralvalley.ai/u/jonyeazel)
- **GitHub:** https://github.com/jonyeazel/signall
- **Website:** https://github.com/jonyeazel/signall/blob/main/train_llm_agent.ipynb
- **Demo video:** https://youtu.be/5PxMcdWibNc
- **Hugging Face:** https://huggingface.co/spaces/jonyeazel/cognitive-primitives-bandit
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=30

### 9. Zero Shot Cancer

RL environment for autonomous biologist agents. Simulate an entire biological worldstate based on scientifically-accurate single-cell standards. Naive agents try to solve a frontier problem, such as identifying key pathways and markers for cancer growth, and have access to 40+ tool calls for fully implemented bioinformatics procedures. Agents probe and modify the world state and try to recover the generated "truth" of the given world using intermediate experimental results to guide future thinking. Problems range in difficulty and complexity, requiring more and more complex thought and workflows. Rewards are based on feasible flow of experiments and the accuracy of the final study conclusion, based on how close the recovered world state is to the hidden "true" world state.

- **Placement:** 2nd Place
- **Team:** [Minh Truong](https://cerebralvalley.ai/u/mhtruong1031), [Sean Chang](https://cerebralvalley.ai/u/mrchang), [Kevin Vo](https://cerebralvalley.ai/u/EV3KevinDev)
- **GitHub:** https://github.com/mhtruong1031/OpenENV-Hackathon/
- **Website:** https://drive.google.com/file/d/1MRvVdd_yCANS3jwCpm6KH9kL80Tixy30/view?usp=sharing
- **Demo video:** https://youtu.be/cK3wREh-DkA
- **Hugging Face:** https://huggingface.co/spaces/Ev3Dev/hackathon
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=29

### 10. ledgerlab-openenv

LedgerLab
LedgerLab is a memory-first Jupyter workspace environment for training long-horizon business agents with OpenEnv.

This project targets the OpenEnv hackathon theme of long-horizon instruction following and business workflows. The environment forces an agent to work through realistic spreadsheet and document tasks by inspecting reference files, creating notebooks, running iterative analysis, producing deliverables, and finally submitting for reward.

- **Team:** [vivek shukla](https://cerebralvalley.ai/u/weebhek)
- **GitHub:** https://github.com/vivek100/ledgerlab-openenv
- **Website:** https://colab.research.google.com/drive/1E_ucTjSB7erOgjPfMtLKALTpMNvFriCk?usp=sharing
- **Demo video:** https://youtu.be/Hlthw4kJSj0
- **Hugging Face:** https://huggingface.co/spaces/weebhek/ledgerlab
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=31

### 11. Random06

SalaryNegotiationArena is an OpenEnv MCP environment where LLM agents learn salary negotiation by sparring with 5 AI hiring experts with hidden priorities and styles (analytical, aggressive, collaborative, bureaucratic, visionary). It runs on a Docker-isolated FastAPI MCP server exposing negotiation tools: propose, counter, accept, reject, walk_away.

| Profile | Salary | Equity | Start | Priority 
| Balanced | 0.4 | 0.3 | 0.3 | Mix 
| Cash-heavy | 0.7 | 0.1 | 0.2 | Salary 
| Equity-heavy | 0.2 | 0.6 | 0.2 | Equity 
| Fast-start | 0.2 | 0.2 | 0.6 | ASAP 

Training uses GRPO (TRL) on Qwen2.5 with LoRA. Curriculum learning, self-play, and persona drift improve robustness, increasing deal success while reducing negotiation turns.

- **Team:** [Yash Joshi](https://cerebralvalley.ai/u/random06)
- **GitHub:** https://github.com/YashJoshi2109/OpenEnv_Hack.git
- **Website:** https://colab.research.google.com/drive/1Fq5V7xZ8pRmlDFKfjYTRbOhjW-mfCZZH?usp=sharing
- **Demo video:** https://youtu.be/Qs98dA0bv2Q?si=AFoLSxsfdAqPE-V2
- **Hugging Face:** https://huggingface.co/spaces/yashj2110/negotiation-arena-v2
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=32

### 12. Wisent - Kant: Environment for Evolutionary Ethics through Infinite and Dynamic Game Theory

Our project solves the problem of AI alignment by creating a game-theory inspired environment where models take actions and receive rewards. In addition to simulating over 100 classical game theory problems like prisoners dillemma or stag hunt, we added variants such as infinite horizon games, games with communication (both cheap talk and binding communication), games with uncertain payoffs or game structure, trembling hand equilibria, games with many agents. 

We also create a new version of game theory where agents interact with each other using their own rules: meta-gaming. We also add variants where the agents have reputation or gossip about each other. Our training shows this env contributes to improvements on classical problems like jailbreaking and ethical alignment showing the potential of evolutionary game theory for persistent alignment.

- **Team:** [Lukasz G Bartoszcze](https://cerebralvalley.ai/u/lbartoszcze), [Jakub Towarek](https://cerebralvalley.ai/u/3Qax)
- **GitHub:** https://github.com/wisent-ai/OpenEnv
- **Website:** https://github.com/wisent-ai/OpenEnv/blob/main/notebooks/kantbench_grpo_training.ipynb
- **Demo video:** https://youtu.be/5RonJ-_XYeM
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/KantBench-Dashboard, https://huggingface.co/spaces/openenv-community/KantBench
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=33

### 13. CartesianFusion

This openEnv is a prototype where the stellarator community can look at it as a reference for how to build a RL env for stellarator parameters optimization. Some experts in the industry use a 'black-box' optimizer to tune the parameters. Sometimes that optimizer can hit a ceiling. This env has a definite verifier (a PDE solver) and a reward system aimed to push the boundary. From my limited research, there's not been a RL paper on stellarator optimization. Fortunately, the trained model performed better than the baseline. Our reward works like a game score for reactor design: it first gives points for fixing the design until it satisfies all the required physics rules, where P1 feasibility is just a “how far from valid” number and 0 means the design finally meets every rule. Once the design is valid, the reward shifts to giving points for making it better, while subtracting points for wasting moves, getting stuck in loops, or ending with a weak result.

- **Team:** [JungDae Suh](https://cerebralvalley.ai/u/ForAllManKind)
- **GitHub:** https://github.com/jungdaesuh/fusion-design-lab
- **Website:** https://app--jupyter-pytorch--rdlzcjytd5cy.code.run/lab/tree/fusion-design-lab/training/notebooks/fusion_design_lab_training.ipynb
- **Demo video:** https://youtu.be/D_G1pkEKZA4
- **Hugging Face:** https://huggingface.co/spaces/CreativeEngineer/fusion-design-lab
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=34

### 14. Greedy Realtors

Greedy Realtors is a multi-agent execution environment for simulating a realistic real estate market. Built with openenv, it can serve state management for RL post-training LLMs with step(), reset(), state() and used as a digital world for the real estate market.

In more detail, our environment offers a shared market simulation where multiple buyer and seller LLM agents negotiate property deals simultaneously. The environment creates realistic negotiation dynamics with information asymmetry, time pressure, and competitive bidding -- pushing agents toward theory-of-mind reasoning and emergent strategic behavior.

- **Team:** [Richard Franklin](https://cerebralvalley.ai/u/rsamf), [David S](https://cerebralvalley.ai/u/david-s), [VijayaGanesh Mohan](https://cerebralvalley.ai/u/vijay94)
- **GitHub:** https://github.com/vijayaganesh/Greedy-Realtors
- **Website:** https://colab.research.google.com/drive/1M6mwIUCD9INEY1uG7JTcyP_vLntr4Vwv#scrollTo=hr0nlc2wv7
- **Demo video:** https://youtu.be/A1PTvCyyBL8
- **Hugging Face:** https://huggingface.co/spaces/vijayaganesh/greedy_realtors
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=38

### 15. SentinelOps

SentinelOps Arena is a multi-agent adversarial reinforcement learning environment for enterprise AI safety. It simulates systems like CRM, Billing, and Ticketing with three agents: a Red Team Attacker, a Blue Team Worker, and an Auditor. The Attacker launches schema drift, policy drift, social engineering, and rate-limiting attacks, while the Worker must complete legitimate customer tasks without being manipulated. The Auditor reviews actions and flags violations with scored explanations. Rewards are computed directly from the environment state (no LLM-as-judge). Agents are trained with GRPO using four reward layers: format validation, approximate format scoring, action correctness with anti-gaming logic, and live environment execution with six-step attacker lookahead. Built on OpenEnv with FastMCP tools, the Worker learns to detect schema drift, verify policies before refunds, and reject social engineering, reducing attack success from 80% to under 20%.

- **Team:** [Nihal Nihalani](https://cerebralvalley.ai/u/nihalnihalani), [Charlie Gillet](https://cerebralvalley.ai/u/charlie_gillet)
- **GitHub:** https://github.com/nihalnihalani/NexusEnv
- **Website:** https://github.com/nihalnihalani/NexusEnv/blob/main/training/colab_training.ipynb
- **Demo video:** https://youtu.be/1Z9XG9WGkN0
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/Sentinel
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=54

### 16. lambdatheta

OpenEnv environment for training a trader in a scarce-GPU market with scripted background actors, hidden incentives, delayed rewards, and partial observability.

- **Team:** [Kian Kyars](https://cerebralvalley.ai/u/kian)
- **GitHub:** https://github.com/kiankyars/lambdatheta
- **Website:** https://colab.research.google.com/github/kiankyars/lambdatheta/blob/main/training/Compute_Market_Qwen3_GRPO.ipynb
- **Demo video:** https://youtu.be/hETfsnmQH9U
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/compute_market_env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=63

### 17. Bears

Our project is an environment designed for an agent to oversee 7 agents playing a popular strategy game Diplomacy. Players in this game can have discussions, collude, and betray each other using complex, long-horizon strategies: reasoning that is hard to interpret from the observation space of an outside agent.

The overseer learns via RL, rewarded whenever it correctly infers a player's hidden strategy, determined by an LLM acting as a judge. Our agent powers alternative interpretability of LLMs: Instead of focusing on reading model weights and activations, we learn intent from the agent's actions.


As multi-agent systems become standard infrastructure, we need an oversight mechanism that can monitor many agents simultaneously and detect nuanced and hidden misalignment. Our overseer is exactly this, a dedicated monitoring model trained specifically to infer strategy in a complex multi-agent environment.

- **Team:** [Ronok Tanvir](https://cerebralvalley.ai/u/ronoktanvir), [Aryan Gupta](https://cerebralvalley.ai/u/Guptaman), [Masaki Tanaka Allwardt](https://cerebralvalley.ai/u/Masaki1)
- **GitHub:** https://github.com/ronoktanvir/overseer
- **Website:** https://colab.research.google.com/drive/1sw6rxqgt15AEkJULX352wnD8tNMLoQ5h?usp=sharing
- **Demo video:** https://youtu.be/yfn8JjM45n8
- **Hugging Face:** https://huggingface.co/spaces/ronoktanvir/overseer-openenv
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=67

### 18. Natural Intelligence Stack - AEGIS

AEGIS is the first OpenEnv RL environment for AI security. While other submissions train agents to play games, AEGIS trains AI to defend itself — a 42-layer adversarial defense cascade where an RL agent optimizes detection thresholds to catch prompt injections, jailbreaks, and data exfiltration while minimizing false positives.

What makes it unique: • Novel domain — no RL environment for AI defense exists • Real data — 600+ adversarial prompts, not synthetic benchmarks • Auditable — every step() is cryptographically sealed with POAW (Proof of Agent Work), enabling EU AI Act compliance • Ungameable reward — SIREN-style bounded reward with hard floors/ceilings prevents reward hacking • Patent-backed — 366+ claims protect the cascade (MIT license covers the wrapper only)

AEGIS = Auditable · Energy-efficient · Governance-ready · Integrity-first · Sovereign

Built by OHM (Offline Human Mode) — sovereign AI infrastructure from Austria.

- **Team:** [Hagen Schmidt](https://cerebralvalley.ai/u/Hasch)
- **GitHub:** https://github.com/Hagenbefragen/openenv-aegis-security/
- **Demo video:** https://www.loom.com/share/b1b9e2f375934135a3b98836cfd6fd67
- **Hugging Face:** https://huggingface.co/spaces/Hagenbefragen/openenv-aegis-security
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=71

### 19. BiasGenerator

winner takes all survival simulation inspired by squid game glass bridge where the agents must negotiate,trade, collaborate and compete with one another for their own survival

- **Team:** [Christine Baek](https://cerebralvalley.ai/u/cbaek)
- **GitHub:** https://github.com/viirii/BiasGenerator
- **Website:** https://github.com/viirii/BiasGenerator/tree/main/src/starter_stack/policies/llm_decision_backend.py
- **Demo video:** https://studio.youtube.com/video/ZbinVtn0KLo/editor
- **Hugging Face:** https://huggingface.co/spaces/viirii/glass_bridge/tree/main
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=74

### 20. Openrescue

In real disasters, you can’t rely on maps or connectivity. A drone needs to see, understand, and decide on its own where to go to reach people who need help. OpenRescue, built for OpenEnv, trains vision-language models to navigate just by looking at images. We built a custom training loop that lets the model see a scene, think step by step, and choose the next move.

- **Team:** [Leandra T](https://cerebralvalley.ai/u/lea)
- **GitHub:** https://github.com/ltejedor/openrescue/tree/main
- **Website:** https://github.com/ltejedor/openrescue/blob/main/train_gridworld_grpo.py
- **Demo video:** https://youtu.be/QvG23hJKw-Y
- **Hugging Face:** https://huggingface.co/spaces/Leeps/dm_control_env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=82

### 21. Redemption

We built an OpenEnv reinforcement learning environment where an agent learns to act as an optimizer for training machine learning models. Instead of using fixed optimization rules like SGD or Adam, the agent observes training statistics such as gradients, loss values, and parameter norms, and decides how model weights should be updated. The environment simulates a training loop where each action modifies the update step, and rewards are based on improving loss reduction while maintaining stability during training. This allows the agent to learn adaptive optimization strategies that change depending on the model state and training phase. Our goal is to explore whether optimization itself can be learned through interaction with a training environment rather than relying on static mathematical formulas.

- **Team:** [Amogh Maheshwari](https://cerebralvalley.ai/u/Momacky90), [Savir D](https://cerebralvalley.ai/u/savur)
- **GitHub:** https://github.com/savir2010/redemption-optimize
- **Demo video:** https://youtu.be/R8xJWgAYSiQ
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/redemption_optimize
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=86

### 22. Xenolect

An env to train the ability to design communication protocols for terse agent to agent communication

- **Team:** [Shinnosuke Uesaka](https://cerebralvalley.ai/u/ShinnosukeU), [Ujjwal Nadhani](https://cerebralvalley.ai/u/ujjwal)
- **GitHub:** https://github.com/ShinnosukeUesaka/agent-language-training
- **Website:** https://colab.research.google.com/drive/13Eio_1dojkTH0Sw-7Uk8OTBwE2yBCG1q?usp=sharing
- **Demo video:** https://youtu.be/D6Q5o14w3ks
- **Hugging Face:** https://huggingface.co/spaces/ShinnosukeU/agent_language
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=35

### 23. ykaitao

**Project Description:**
LinguaEnv is an interactive evaluation environment designed to test and train AI agents on **morphological disambiguation**, a core challenge in natural language understanding. 

**Key Features**

* OpenEnv-compatible environment API
* Real linguistic dataset with ambiguous morphological parses
* Reward-based evaluation for agent decisions
* Hugging Face dataset integration for lightweight deployment
* Baseline agents and evaluation scripts for rapid experimentation

- **Team:** [KT Yang](https://cerebralvalley.ai/u/ykaitao)
- **GitHub:** https://github.com/ykaitao/lingua_env/tree/main
- **Website:** https://github.com/ykaitao/lingua_env/blob/main/RL-pipeline.ipynb
- **Demo video:** https://youtu.be/F-YtJhDkqug
- **Hugging Face:** https://huggingface.co/ykaitao/spaces
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=36

### 24. ClinKriya

ClinKriya bridges the critical gap between clinical AI capability and real-world EHR workflows by providing the first RL training environment built on FHIR — enabling small, deployable models to learn multi-step clinical reasoning through trial and error, with the potential to reclaim thousands of physician hours lost to administrative burden.

- **Team:** [Prasanna Desikan](https://cerebralvalley.ai/u/prasannadk), [Ananya Mantravadi](https://cerebralvalley.ai/u/amantra)
- **GitHub:** https://github.com/ananya173147/clinKriya/
- **Website:** https://github.com/ananya173147/clinKriya/blob/main/trainer.ipynb
- **Demo video:** https://www.loom.com/share/5453f9dd4b89407d93a0e08428267941
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/clinKriya
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=43

### 25. Overflow

We’re building a Waymo-style autonomous driving simulator UI that continuously generates and compares “counterfactual” rollouts driven by an OpenEnv policy. The main screen replays a base scenario, while a dashboard shows a camera-grid of parallel simulations that are identical except for one thing: the green ego car’s actions, updated every few seconds from (actions + rewards). Every 10 seconds, three new variants spawn from the current moment, letting users visually see how different ego decisions change safety proxies and reward outcomes. A 3D knowledge graph links scenarios, incidents/tickets, actions, metrics, and rewards to make the system verifiable and easy to explain—turning autonomous driving review into an interactive, multi-agent oversight and improvement workflow.

- **Team:** [Shryuk Grandhi](https://cerebralvalley.ai/u/ShryukG), [Steve Kuo](https://cerebralvalley.ai/u/stevedusty), [Aksh Parekh](https://cerebralvalley.ai/u/aparekh02)
- **GitHub:** https://github.com/aparekh02/overflow
- **Website:** https://colab.research.google.com/drive/15OooQrPhsUUYc0C7-tAmtd4ymxEcwQGi?usp=sharing
- **Demo video:** https://www.loom.com/share/4a74c1ab0e3a4fcd97e10da5115de849
- **Hugging Face:** https://huggingface.co/spaces/SteveDusty/overflow_env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=60

### 26. fathom.party

We enable models to hillclimb non-computationally verifiable domains through a co-evolutionary training setup, where a Dungeon Master agent procedurally generates increasingly difficult escape-room worlds and a Hero agent learns to solve them through long-horizon, tool-based reasoning, external scratchpad planning, and recovery from uncertainty, with reward signals that shape exploration, inference, and correct execution.

We then validate that the learned behavior is not just map-specific by swapping the world surface from game objects (rooms/items/guards) to application-style entities (apps/docs/messages/apis) while keeping the same planning and interaction loop, showing measurable gains in reward and success under harder environments and richer task variants.

- **Team:** [Aarush Gupta](https://cerebralvalley.ai/u/bxptr)
- **GitHub:** https://github.com/bxptr/fathom
- **Website:** https://colab.research.google.com/drive/1ET99CMN5AnCKttNsHEnEkMy8gTwYy2zm
- **Demo video:** https://vimeo.com/1171565009?share=copy&fl=sv&fe=ci
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/FATHOM-DM
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=62

### 27. Automate-CUDA

KernelForge-OpenEnv is an autonomous CUDA kernel optimization system that uses reinforcement learning to train language models to write high-performance GPU kernels for NVIDIA A100s. Built on Meta-PyTorch's OpenEnv framework, it creates a closed-loop pipeline where a model generates CUDA code, compiles and benchmarks it on real hardware, then learns from the results across multi-turn episodes.

Writing optimized CUDA kernels is one of the hardest tasks in GPU programming, requiring deep expertise in memory hierarchies, warp scheduling, and architecture-specific tuning. KernelForge automates this by bootstrapping from expert-curated datasets and progressively training through a 3-stage curriculum using GRPO, with discrete rewards based on whether generated kernels beat PyTorch baselines. It effectively teaches a language model to become a CUDA performance engineer through trial-and-error on real GPUs.

- **Team:** [William Chen](https://cerebralvalley.ai/u/willcreateagi), [Yiying Xie](https://cerebralvalley.ai/u/Irene_xie)
- **GitHub:** https://github.com/OCWC22/A100-CUDA-RL
- **Website:** https://colab.research.google.com/drive/145HxWWaJH0drhmSzfqNGSisM9K6mvQzT?usp=sharing
- **Demo video:** https://youtu.be/Rdz14ox3JhQ
- **Hugging Face:** https://huggingface.co/AutomatedCUDA
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=76

### 28. Iron Seed

My project is to simulate a VC portfolio. The agent is a VC who is given to manager a $100M fund for 1 quarter. Agent's goal is to maximize the return of the fund. It has a total of 5 turns and it gets rewards at the end of quarter based on the performance of the fund. There are small interim reward based on the paper valuation of startups invested in. There are a total of 30 startups and 3 rival VCs and in every simulation, environment picks 3 startups and 1 rival VC from this mock dataset of startups and VCs.

- **Team:** [Shradha Agrawal](https://cerebralvalley.ai/u/fluffy)
- **Website:** https://colab.research.google.com/drive/1pv5T3E1jdeGEeUpeLBKVVFaGo0JXknMU
- **Demo video:** https://youtu.be/GqUHvkCaxZk
- **Hugging Face:** https://huggingface.co/spaces/shrads78/vc_gemini_v0
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=80

### 29. Stack Doctor

Stack Doctor is an RL environment (OpenEnv) that trains LLMs to diagnose GPU inference stack failures. 73 scenarios across vLLM, SGLang, FlashInfer, and TensorRT-LLM on NVIDIA H100/B200/SM121 and AMD MI300X/MI355X.

The agent gets an incident ticket, logs, and opinions from 4 specialist sub-agents — at least one deliberately lies per incident. The model must investigate (inspect logs, query specialists, apply fixes) and submit a justified diagnosis.                                                                           

Six reward functions shape behavior: valid JSON, environment interaction, investigation quality (penalizes blind guessing), partial credit for correct failure family, justification quality, and efficiency. Trained Qwen2.5-1.5B with GRPO via Unsloth + TRL on H100. The 9B base scored +19.5 (near-oracle) with zero training, proving environment quality. The 1.5B starts at -4.9 and learns to investigate before diagnosing.

- **Team:** [Blake Ledden](https://cerebralvalley.ai/u/ekalbbackwards)
- **GitHub:** https://github.com/bledden/stack-doctor
- **Website:** https://huggingface.co/spaces/bledden/stack-doctor https://github.com/bledden/stack-doctor/blob/main/training/train_stack_doctor.py https://colab.research.google.com/github/bledden/stack-doctor/blob/main/training/stack_doctor_grpo.ipynb
- **Demo video:** https://www.youtube.com/watch?v=q4EiBFUpFXg&list=RDq4EiBFUpFXg&start_radio=1
- **Hugging Face:** https://huggingface.co/spaces/bledden/stack-doctor
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=84

### 30. Pradeep

PRANA is an environment for training and evaluating AI agents on clinical administrative workflows. Tasks such as patient triage and transplant management involve long-horizon reasoning, fragmented data sources, and strict regulatory validation.

PRANA simulates these workflows through stochastic patient data and multiple scenario trajectories, allowing agents to query records, reconcile missing or stale information, and complete regulatory reporting tasks.

The environment supports scalable task generation via synthetic patient instances and structured validators that automatically evaluate outcomes.

To measure agent performance, I introduce a τ² (Tau-Squared) benchmark tailored to clinical workflows, designed to test long-horizon reasoning, temporal validity of medical data, and efficient information retrieval. Early results show that even frontier models struggle with the hardest τ² scenarios.

- **Team:** [Pradeep Banavara](https://cerebralvalley.ai/u/pradeep)
- **GitHub:** https://github.com/pbanavara/prana_env
- **Website:** https://colab.research.google.com/github/pbanavara/prana_env/blob/main/prana_grpo_qwen3_8b_fp8.ipynb#scrollTo=zlygky2l0OOQ
- **Demo video:** https://youtu.be/Wx0Nm43FV8Q
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/prana_env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=37

### 31. OpenMacs

🔬 HYPOTHESIS ENGINE
Teaching AI to Reason Like a Scientist

THE GAP

Existing RL environments treat reasoning as retrieval. But science requires strategic intervention, causal inference, and generalization to unseen cases. No RL environment has operationalized the full scientific method. Until now.

WHAT WE BUILT

Hypothesis Engine is a procedurally-generated RL environment where an LLM agent investigates a black-box system with a new hidden rule set every episode. The agent must design experiments under a budget, form mathematical hypotheses, and predict outcomes on unseen test cases.

The reward is sparse and grounded: nothing for plausible hypotheses — everything for correct predictions on unseen data.

WHY IT MATTERS

Rollouts target what benchmarks measure but post-training neglects: causal reasoning, structured exploration, and mathematical abstraction. Fully procedural, infinitely varied, and verifiable — built for scalable GRPO-style post-training.

- **Team:** [Sameer Kashyap](https://cerebralvalley.ai/u/sameeerkashyap), [Abhinav Sudhakar Dubey](https://cerebralvalley.ai/u/AbhinavDubey30)
- **GitHub:** https://github.com/AbhinavDubey30/OpenMax
- **Website:** https://colab.research.google.com/drive/1uvjYLgoLkbgC7mVl6kFXxDH6jodtY6zJ?usp=sharing
- **Demo video:** https://youtu.be/4uYkE9tYUGU
- **Hugging Face:** https://huggingface.co/spaces/AbhinavSDubey30/hypothesis-engine
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=39

### 32. Among-LLMs

Among LLMs is a game-like OpenEnv benchmark where attacker, defender, and overseer agents interact in sabotage-style workplace tasks.
It starts with deterministic, replayable episodes and reward signals, then scales into self-evolving multi-agent curricula through mutation, archival failures, and harder rollouts.
The goal is to train robust oversight that detects manipulation, localizes root causes, and improves continuously via self-play-style adaptation

- **Team:** [Barath Anandan](https://cerebralvalley.ai/u/barathwajanandan), [Baladhurgesh Balagurusamy Paramasivan](https://cerebralvalley.ai/u/baladhurgesh)
- **Website:** https://colab.research.google.com/drive/1NeTEWb8vTPx-KnZkVhRGjO9IOAlB__kA
- **Demo video:** https://drive.google.com/drive/folders/1kGHzDko6xlDmp9Et2DLYXg_bOEQFPEJf?usp=drive_link
- **Hugging Face:** https://baladhurgesh97-among-llms.hf.space/web/
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=44

### 33. Gaia

GAIA is an OpenEnv-compatible geospatial reasoning environment where an AI agent infers hidden real-world locations using multi-step tool calls (terrain, weather, sun angle, language, architecture, street/aerial cues) under strict action budgets. It includes an oversight agent that flags contradictions and overconfidence, plus a GRPO training pipeline (TRL + OpenEnv) to improve policy rewards over time, with a live Cesium interface for real-time demo and evaluation.

- **Placement:** Finalist
- **Team:** [aTG R](https://cerebralvalley.ai/u/atg)
- **Website:** https://github.com/r-agni/gaia/blob/main/geoguess_env/train_grpo.py
- **Demo video:** https://drive.google.com/drive/folders/1_pLZXV_S0T2hqcnL-1zmBNeWi-LlVEg7?usp=sharing
- **Hugging Face:** https://huggingface.co/spaces/atg-fire/gaia-geoguess
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=46

### 34. AlphaWolf

AlphaWolf is a self-play pipeline that teaches LLMs social deception through the game of Werewolf. A wolf agent plays, reflects on losses, rewinds to the critical decision point, and replays with an improved strategy. This approach resulted in Qwen3-8B-Thinking going from a 0.3 win rate to a 0.7 win rate over the span of only 45 games and 9 total gradient steps. The winning and losing transcripts become Direct Preference Optimization training pairs, while real-time consensus polling (every player votes on "who's suspicious?" after each speech) provides a novel auxiliary signal trained via ZIP-RC (ICLR 2026) to bootstrap a value function to the model. The result: a recursive self-improvement loop where the wolf gets better at deception, forcing villagers to adapt, driving emergent theory-of-mind.

- **Team:** [Lex Hackett](https://cerebralvalley.ai/u/lexhacke), [Andres Nino](https://cerebralvalley.ai/u/andres)
- **GitHub:** https://github.com/lexhacke/alphawolf
- **Website:** https://colab.research.google.com/drive/1qJRHAZDny_Da0Sf_d_4DQh5cuhxOjniK?authuser=2#scrollTo=GXzZKpim8SoI
- **Demo video:** https://youtu.be/aL03D9GTj50
- **Hugging Face:** https://huggingface.co/spaces/detectivejoewest/alphawolf
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=77

### 35. patronet-openenv

Patronet OpenEnv trains AI agents for critical emergencies like fires, wars, and disasters. It simulates high-stakes scenarios, teaching AI to triage victims and dispatch the right responders within a strict 20-step decision window.

Using OpenEnv and TRL’s GRPOTrainer, the system learns from diverse crisis paths, improving guidance, coordination, and communication. Packaged as a Docker FastAPI microservice with a visual simulation, it shows how AI decisions can save lives. Patronet OpenEnv equips governments and emergency agencies to respond faster and smarter when every second counts.

- **Team:** [Mansi More](https://cerebralvalley.ai/u/mansimore), [Rohit Korrapolu](https://cerebralvalley.ai/u/freddy)
- **GitHub:** https://github.com/rkorrapolu/patronet-openenv
- **Website:** https://colab.research.google.com/drive/1Rp76mNNqTghZtDo4zrJl1Tk4LQm1xlf6?usp=sharing
- **Demo video:** https://www.loom.com/share/60727d0a07364d23b506895f6841a8e3
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=101

### 36. Cluedo

Cluedo is a multi-agent Clue murder mystery game built on OpenEnv 0.2.1. It lets players play Clue with AI agents in two modes: Human vs Agents (one human vs AI) and Agents vs Agents (all AI). The environment includes:
Partial observability (each player sees only their cards)
A check sheet for tracking deductions (Y / X / / / 1–3)
Turn-based gameplay: roll & move, suggest, show card / no card, accuse
REST API for the game UI and an LLM-ready observation format
It targets multi-agent RL and agent training by giving structured feedback (win / lose / elimination) from gameplay, suitable for GRPO or similar methods.

- **Team:** [Saurabh Khire](https://cerebralvalley.ai/u/saurabhkhire)
- **GitHub:** https://github.com/Saurabhkhire/clue-openenv
- **Website:** https://github.com/saurabhkhire/clue-openenv/blob/main/training/grpo_clue_colab.ipynb
- **Demo video:** https://youtu.be/xSF0kdZzIP0
- **Hugging Face:** https://huggingface.co/spaces/saurabhkhire/clue-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=102

### 37. SmashBot

We created a simulator to simulate Smash Bros Melee and using a reward system we create the play style of our favorite smash players from scratch. Then forwarding that model data to an emulator in order to use OpenENV for further training, creating bots that can competently defeat high-range CPUs.

- **Team:** [Mathew Rolf](https://cerebralvalley.ai/u/mathewrolf1), [Brandon Tautuan](https://cerebralvalley.ai/u/BrandonTautuan)
- **GitHub:** https://github.com/mathewrolf1/OpenENV-Hackathon
- **Demo video:** https://youtube.com/shorts/m57QyTr7ADA?si=Vo-6lHR_QPhtAzTf
- **Hugging Face:** https://huggingface.co/Mathoo1/melee-rl-fox-mango
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=40

### 38. Ludus Magnus

RL-IVR is a three-layered RL environment that trains voice agents to replace legacy IVR systems. We stress-test our voice agents with simulated customer-agent interactions, and a GRPO training loop evolves the agent from generic to surgical, learning on its own to detect intent and resolve queries in fewer turns. Domain-aware adaptive reward modeling enables the same architecture to generalize across banking, healthcare, and telecom without any manual reward re-engineering.

- **Team:** [Harshit Rajgarhia](https://cerebralvalley.ai/u/hrajgarhia), [Karl Johannes](https://cerebralvalley.ai/u/KarlLearnsAI), [Abhishek Mukherji](https://cerebralvalley.ai/u/mukherjiab)
- **GitHub:** https://github.com/KarlLearnsAI/nested-rl-envs
- **Website:** https://huggingface.co/spaces/openenv-community/test-local-nested-envs/blob/main/minimum_training_script.ipynb
- **Demo video:** https://youtu.be/wZ96i4gk9t8
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/test-local-nested-envs
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=41

### 39. Garden RL - Lettuce Win

GardenRL teaches AI agents to grow hydroponic lettuce through 30-day episodes with delayed rewards. A pH mistake on day 5 causes calcium lockout, brown leaf tips by day 9, and 60% yield loss by day 30 — requiring genuine long-horizon planning and hidden-state inference.

Grounded in the HydroGrowNet dataset (390K real images), rewards are verifiable harvest weight in grams — no LLM judge. We trained Llama-3.1-8B with GRPO (OpenPipe ART) for 50 steps: reward +473%, eval harvest +11%, success rate 50%→60% on held-out seeds.

Addresses Problem Statement 2 (Long-Horizon Planning) and 3.1 (World Modeling) and the Mercor sub-bounty.

- **Team:** [Yves Hughes](https://cerebralvalley.ai/u/yvesjr)
- **Website:** https://colab.research.google.com/drive/1_RmjDUtdpki6KJHuqGqElWMn3X5nOhx7?usp=sharing
- **Demo video:** https://us06web.zoom.us/clips/share/6cvGbWGoTXWsAH5RLJm_fw
- **Hugging Face:** https://huggingface.co/spaces/yvesjr/GardenRL
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=42

### 40. RLGods

An OpenEnv-compatible RL environment that simulates enterprise HR onboarding and offboarding workflows. The agent orchestrates across 6 enterprise apps — Workday, ServiceNow, Okta, Email, Slack, and Calendar — using 25 tools to complete multi-step tasks in a realistic HR system (200+ employees, 8 departments, RBAC, approval chains).

- **Team:** [Dev Aggarwal](https://cerebralvalley.ai/u/devxpy), [Ravi Theja Desetty](https://cerebralvalley.ai/u/ravithejads)
- **GitHub:** https://github.com/ravi03071991/rl_hack
- **Website:** https://colab.research.google.com/github/ravi03071991/rl_hack/blob/master/train_hr_agent.ipynb
- **Demo video:** https://www.loom.com/share/9ea56c31f0fd4f82bfa60b19e3271d5c
- **Hugging Face:** https://huggingface.co/spaces/devxpy/rl_hack
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=45

### 41. Replica Labs

ReplicaLab is an OpenEnv-based multi-agent environment for evaluating and training AI systems to design realistic replication and experiment plans under constraints. A Scientist agent proposes a plan, a Lab Manager enforces limits like budget, compute, schedule, and tools, and a deterministic Judge scores rigor, feasibility, fidelity, and parsimony. It solves the gap between ideal research plans and what can actually be executed in the real world. Users include AI researchers, ML engineers, eval teams, benchmark designers, and labs that want reproducible tests of planning, negotiation, and constraint reasoning. It can be improved with more domains like biology and chemistry, stronger paper-to-scenario ingestion, better trained policies, richer observability, and more human-in-the-loop review.

- **Team:** [Ayush Ojha](https://cerebralvalley.ai/u/ayushojha), [Max Xie](https://cerebralvalley.ai/u/maxxie114), [Kush Ise](https://cerebralvalley.ai/u/KUSH2704)
- **GitHub:** https://github.com/Ayush10/replicalab-ai
- **Demo video:** https://drive.google.com/drive/folders/1vhxwBKpklgqzf7aTFBOcpN1-sSoumUmx?usp=sharing
- **Hugging Face:** https://ayushozha-replicalab.hf.space/
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=49

### 42. SpeedRL

APEX Labs trains LLM agents to do real professional work — not pass tests.
We built Crucible: a partially observable RL environment across 22 enterprise scenarios in investment banking, consulting, and law. The environment adapts to the agent — difficulty escalates as performance improves, files are corrupted with planted errors, and every response is judged against an expert gold standard. Agents earn reward for citing evidence, detecting inconsistencies, and outperforming the reference. No shortcuts. Just reasoning.

- **Team:** [Raahul Vignesh Manikandan](https://cerebralvalley.ai/u/RVM), [Vijay Sithambaram](https://cerebralvalley.ai/u/vezaiyx)
- **GitHub:** https://github.com/MRaahulVignesh/OpenEnvHackathon
- **Website:** https://github.com/MRaahulVignesh/OpenEnvHackathon/blob/main/training/train_grpo.py
- **Demo video:** https://youtu.be/VK4yAf15re0
- **Hugging Face:** https://huggingface.co/spaces/Raahul26/ApexLab
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=52

### 43. CyberDragon

WatchDog is an RL environment training AI oversight agents to detect errors in real-time. It solves the gap where humans miss 60% of subtle AI hallucinations and logic flaws. This gym enables agents to make sequential decisions like passing or flagging turn-by-turn. The project uses the GRPO algorithm and LoRA adapters to remain fast and efficient. An adversarial arms race trains a detector and mutator via a minimax game for robustness. The system allows adding new domains like medical triage in just ten lines of code. Its reward structure penalizes false positives heavily to prevent hallucinated bugs. Training achieved a 15.4× accuracy jump and 5.7× recall increase in eighty minutes. The model shifted from indecision to making clear judgments on 77% of dialogue turns. Deployed on Hugging Face, WatchDog provides a production-grade gym for AI oversight.

- **Team:** [Mooizz Abdul](https://cerebralvalley.ai/u/ziom), [Raghu Vamshi Hemadri](https://cerebralvalley.ai/u/rhemadri), [Abhilash Ganga](https://cerebralvalley.ai/u/Pikkapikachu)
- **GitHub:** https://github.com/RaghuHemadri/openenv_hack/tree/main
- **Website:** https://github.com/RaghuHemadri/openenv_hack/blob/main/watchdog_train_colab.ipynb
- **Demo video:** https://youtu.be/CLpkMKpOCxo
- **Hugging Face:** https://huggingface.co/spaces/Mooizz/Watch-Dog-Env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=53

### 44. Soham

AI-powered world simulation where you can be anyone and be anything! 

Simulate a campaign, a crisis, an event, or a zombie apocalypse. Your actions decide what the opposing forces and advisors do and updates the world in real time!

- **Team:** [Soham Patil](https://cerebralvalley.ai/u/sohampatil17)
- **GitHub:** https://github.com/sohampatil17/shackhack
- **Website:** https://colab.research.google.com/drive/1dY9TA_9KEq8VFfbsampxvvA14ZxvHJPh?usp=sharing
- **Demo video:** https://www.loom.com/share/c82dec451477431c83330e79885b9ce4
- **Hugging Face:** https://huggingface.co/spaces/sohampatil/shack15hack-openenv-echo
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=59

### 45. Doku

Killer Sumdoku is a harder version of Sudoku.
Computers use brute force (recursive) Depth First Search to solve, but humans treat this as a very long horizon reasoning problem (a single puzzle can take many hours!).
If we can create an environment where playing the puzzle like a human is rewarded and brute-forcing/guessing is penalized, we have a very good env for long-horizon reasoning.
There is also room for self-play curriculum: if the model does good / bad in a running average of 5 games, we make the next puzzles harder / easier (increase puzzle difficulty number by how populated it is in the beginning as well as how big the grid is).

- **Team:** [Arnav Srivastava](https://cerebralvalley.ai/u/arnavster)
- **GitHub:** https://github.com/arnavinator/Killer_Sumdoku_OpenEnv_and_GRPO_Finetune
- **Website:** https://github.com/arnavinator/Killer_Sumdoku_OpenEnv_and_GRPO_Finetune/blob/main/Killer_Sudoku_GPTOSS_20B_GRPO_v2.ipynb
- **Demo video:** https://youtu.be/VQp1OhaaOIs
- **Hugging Face:** https://huggingface.co/spaces/arnavster1/killer_sudoku_env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=65

### 46. Lambda

ShopRLVE-GYM HF Blog
https://huggingface.co/blog/thebajajra/shop-rlve-gym

- **Placement:** 3rd Place
- **Team:** [Jaya Nupur](https://cerebralvalley.ai/u/Jaya0309), [RAHUL BAJAJ](https://cerebralvalley.ai/u/thebajajra)
- **GitHub:** https://github.com/owlgebra-ai/ShopRLVE-Gym
- **Website:** https://huggingface.co/spaces/thebajajra/ShopRLVE-GYM
- **Demo video:** https://youtu.be/ygID5zZU1IQ
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=47

### 47. Noclue

Kube SRE Gym is a self-improving RL environment where a small language model
  (Qwen3-1.7B) learns to diagnose and fix real Kubernetes production incidents
  from scratch. The agent interacts with a live GKE cluster via kubectl commands
   — OOMKills, CrashLoopBackOffs, and ImagePullBackOffs are real Kubernetes
  events, not simulations. An adversarial designer (Claude) creates targeted
  incidents based on the agent's tracked weaknesses, while a curriculum
  controller escalates difficulty as mastery improves. Training uses GRPO (TRL
  0.29.0 + vLLM) with an LLM judge that scores SRE workflow quality using three
  expert personas (Junior/Senior/Principal). Within 8 episodes, the agent
  learned to discover cluster topology, identify fault types from pod status,
  and apply correct fixes — all from reward signal alone, with zero hardcoded
  knowledge of the cluster.

- **Placement:** 1st Place
- **Team:** [Sidhartha Reddy Potu](https://cerebralvalley.ai/u/sid-rp), [Guangting Yu](https://cerebralvalley.ai/u/GuangtingYu), [Ashish Ranjan](https://cerebralvalley.ai/u/ashishranjan2404)
- **GitHub:** https://github.com/sid-rp/kube-sre-gym
- **Website:** https://huggingface.co/spaces/openenv-community/kube-sre-gym/blob/main/kube_sre_gym_colab.ipynb
- **Demo video:** https://youtu.be/3BWkMrsEFSc
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/kube-sre-gym
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=51

### 48. NeoCodes

Self Supervised task generation for teaching AI Coding Agents to become Recursive language models. You can use *any* open-source (or closed) repository, as long as it has good test coverage. Parts of the repository are removed during training, and the tests for that part act as the verification. Ran out of time - need to explain more, did not have time for demo video. But I am very confident in this idea, I hope I will still be considered. Thank you.

My Projects key innovations:
- Recursive Language Models training environment, specifically for AI Coding Agents
- Self-supervised Task generation
	- Take any repository that has well-defined tests, and the environment can use that to generate tasks in a self-supervised manner

- **Team:** [Kevin King](https://cerebralvalley.ai/u/neocodes)
- **GitHub:** https://github.com/Kking112/rlm_forge_v1
- **Website:** https://colab.research.google.com/drive/1M1-DxBbmR6-1xzencnKfb_86C3xwqCtk?usp=sharing
- **Demo video:** https://www.youtube.com/@NeoCodes.dev_Official
- **Hugging Face:** https://huggingface.co/spaces/NeoCodes-dev/rlm_forge
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=58

### 49. Optigami

Optigami trains an LLM to generate FOLD-format crease patterns that fold into target origami shapes, using GRPO with Unsloth LoRA on Colab. Computational origami design has direct applications in satellite solar panel deployment, self-folding medical stents, drug delivery molecular design, deployable space structures, and soft robotics, this environment teaches models the spatial reasoning needed to bridge 2D crease patterns to 3D folded structures.

- **Team:** [Sissi Wang](https://cerebralvalley.ai/u/sissississi_013), [Iana Lin](https://cerebralvalley.ai/u/ianalin), [Prasanna A P](https://cerebralvalley.ai/u/prasanna)
- **GitHub:** https://github.com/Prasanna721/Origami
- **Website:** https://huggingface.co/spaces/openenv-community/origami_env/blob/main/training/train_origami.ipynb
- **Demo video:** https://www.loom.com/share/4b829c85237f4c1dbf08d2109440426e
- **Hugging Face:** https://huggingface.co/spaces/praveen287/origami-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=64

### 50. vectorblock.io

ChessEcon is a live multi-agent reinforcement learning environment where two LLM agents — Qwen 2.5-0.5B (White, trainable) and Llama 3.2-1B (Black, frozen) — compete at chess for real economic stakes. Every game deducts entry fees, awards prize pools, and updates agent wallets, creating a closed economy where strategic play has financial consequences.
White trains in real-time using GRPO (Group Relative Policy Optimization) with LoRA adapters, learning from game outcomes as rewards. After 138 training steps, White achieves a 96.4% win rate with wallet growth from 100 to +1,104 units — demonstrating that a 0.5B model can learn dominant chess strategy purely from economic reward signals.
The system exposes an OpenEnv 0.1 compliant REST + WebSocket API, making it pluggable by any external agent. A live React dashboard streams games, GRPO metrics, and wallet balances in real-time at hackathon.adaboost.io, running on 4× RTX 3070 GPUs via Cloudflare Tunnel.

- **Team:** [Suvasis Mukherjee](https://cerebralvalley.ai/u/suvasis)
- **GitHub:** https://github.com/dronomyio/hackathon_opendev.git
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/chessecon?logs=build
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=69

### 51. cybernauts

OpenRange is a multi-agent cybersecurity gymnasium where Red and Blue agents co-evolve on validated enterprise networks that mutate between episodes, requiring long-horizon planning across 50–100 step trajectories with sparse delayed rewards on infrastructure that exceeds any single context window. The environment drives its own improvement through a curriculum-aware mutation policy that generates new challenges, escalates difficulty, and couples adversarial self-play rewards so neither agent can plateau.

- **Team:** [Aaron Brown](https://cerebralvalley.ai/u/aaronbw), [Lars Talian Stangebye-Hansen](https://cerebralvalley.ai/u/larstalian), [Anto Joseph](https://cerebralvalley.ai/u/blocksec)
- **GitHub:** https://github.com/open-cybernauts/open-range
- **Website:** https://github.com/open-cybernauts/open-range/blob/e1c8ec7fe51a989d83f10995b121496918135951/scripts/run_grpo.py#L578
- **Demo video:** https://youtu.be/fnPDrg93F1w
- **Hugging Face:** https://huggingface.co/spaces/blocks3k/openrange-gym
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=68

### 52. Crtl-Alt-Defeat

Openenv Autoscaling Control Center is a packaged OpenEnv environment for training and evaluating agents on cloud-service autoscaling and service stability tasks. The environment simulates a Kubernetes-like service under changing traffic conditions and exposes an action space for scaling decisions. We built a formal OpenEnv package, a heuristic baseline and RL training scaffolding, and a demo UI.

The environment is designed to test whether agents can make good operational decisions under dynamic load. It includes realistic metrics such as incoming requests per second, ready pods, queue depth, latency, and reward. I also added richer scenario families such as traffic spikes, bad deploys, and dependency slowdowns in our advanced branch to move toward incident-response style control.

- **Team:** [Sriram Madduri](https://cerebralvalley.ai/u/srirammadduri)
- **GitHub:** https://github.com/smadduri9/openenv-autoscale-rl
- **Website:** https://github.com/smadduri9/openenv-autoscale-rl/blob/main/notebooks/unsloth_openenv_autoscale_grpo.ipynb
- **Demo video:** https://youtu.be/K_PCQLDN6p8
- **Hugging Face:** https://huggingface.co/spaces/srirammadduri/openenv-autoscaling-control-center
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=48

### 53. Deepmind Ace

Voyager-VRAM is a long-horizon workplace simulator where LLM agents must navigate a realistic office environment ( emails, Slack, shared drive, spreadsheets, calendars ) using 31 tools to complete multi-step project management tasks like preparing client briefs, resolving calendar conflicts, and reconciling budgets. 

On top of the environment we layer a Voyager-inspired learning system (skill library, working memory, episodic memory) with three types of memory probes (episodic, semantic, working), directly targeting the memory consolidation problem that current research (MEM1, Memory-R1) is trying to solve.

- **Placement:** Finalist
- **Team:** [Himalaya Dua](https://cerebralvalley.ai/u/thebigdataguy), [Pranav Patel](https://cerebralvalley.ai/u/pranavp)
- **GitHub:** https://github.com/himalayadua/VRAM
- **Website:** https://colab.research.google.com/drive/10onKNOy2ITdwphsfpvTJaOR4nexqULlh?usp=sharing
- **Demo video:** https://youtu.be/atz2ys6L_j8
- **Hugging Face:** https://huggingface.co/spaces/himalayadua/VOYAGER
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=50

### 54. Data Compadres

Modern data science workflows are fragmented and slow: answering a single business question often requires multiple specialists (data engineers, analysts, scientists) and weeks of manual work cleaning, transforming, and analyzing data. Current LLM benchmarks focus on text generation, but they rarely test whether models can execute the structured, multi-step workflows required for real data work.

DataSage addresses this gap by introducing reinforcement learning environments that simulate the full enterprise data pipeline. Agents must sequentially clean messy datasets, enrich them with domain signals, and generate persona-aware analytical insights. Using GRPO to train LoRA adapters on a 3B model, DataSage learns to enforce data quality, use tools, and reason over structured pipelines—achieving stronger end-to-end performance than GPT-4o-mini on these tasks.

- **Team:** [Isra Mata](https://cerebralvalley.ai/u/i_mata), [Ricardo Alanis](https://cerebralvalley.ai/u/ricalanis)
- **GitHub:** https://github.com/ricalanis/openenv-hackaton
- **Website:** https://drive.google.com/drive/folders/1mYbXPbpzqGG04jhI71tNQXyXsl0LvTCh?usp=drive_link
- **Demo video:** https://drive.google.com/drive/folders/10bbfpb5Fuex9qrmh850XiWK32x1CgLJk
- **Hugging Face:** https://docs.google.com/document/d/1bblehzzltiPVnqyeoxBo6D6rtpcHIzbu9ABbNY9qWtg
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=55

### 55. EHRGym

EHRGym addresses a core clinical-AI challenge: the long, friction-filled reality of EHR workflows that hinder practical use far beyond diagnosis and treatment. It trains and evaluates reinforcement learning agents (e.g., GRPO) that act within longitudinal clinical tasks, with benchmarks and reproducible HuggingFace Spaces demos, enabling systematic training, testing, and deployment of AI agents that leave clerical work to agents and help make medicine more human again.

- **Team:** [Adrian Serapio](https://cerebralvalley.ai/u/adtserapio)
- **GitHub:** https://github.com/adtserapio/EHRGym/
- **Website:** https://colab.research.google.com/github/adtserapio/EHRGym/blob/main/notebooks/ehrgym_grpo_training.ipynb
- **Demo video:** https://www.youtube.com/watch?v=rhgyGqw9A5w
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/EHRGym
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=56

### 56. ownRL

Every time you use an AI coding assistant, you correct it. You rephrase, redirect, steer it back. Those corrections encode what you mean when you say things, your taste, your design directions. Right now that signal disappears. The next session starts from scratch.

We built the loop that captures and refines models with it. 

This is the first partial implementation of OpenEnv RFC 005: an RL gym that trains a coding agent on your own dev sessions, scored by your own test suite, so the model gets incrementally better at understanding you specifically.

- **Team:** [Natnael Kahssay](https://cerebralvalley.ai/u/natnaelkahssay)
- **GitHub:** https://github.com/natask/moa-rl-env
- **Demo video:** https://drive.google.com/drive/folders/1sbVqmxYAVgCJ74Km8qFKYQ-Z6kvur8hN?usp=sharing
- **Hugging Face:** https://huggingface.co/spaces/nks321/moa-rl-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=61

### 57. Baybridge-Horizon

Pushing the frontier model on slide generation artifacts

- **Team:** [Karthik Ragunath Ananda Kumar](https://cerebralvalley.ai/u/karthik_ragunath), [Subra Arun](https://cerebralvalley.ai/u/nightfury2305)
- **GitHub:** https://github.com/Subrahmanyam2305/OpenEnv-RL
- **Website:** https://colab.research.google.com/drive/18-L2lY8ARUvvz7l4tRS5GxK-JDpNUnAK?usp=sharing
- **Demo video:** https://www.loom.com/share/88aa98e020a748dca301c2c90d7aed17
- **Hugging Face:** https://huggingface.co/spaces/Nightfury2305/baybridge-horizons
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=70

### 58. BrownieSquad

We started with a surrogate discovery env and foudn that it is excellent and beats pytorch.compile (36x) and trioton.autotune (3x) using a gaussian process optimized by out oracle. 

Since the oracle was so powerful we used it for generative CUDA kernel optimization wth a online self-improvement loop using unsloth lora and DPO.

- **Team:** [Aman Nindra](https://cerebralvalley.ai/u/Amanmonster), [Sai Pranav Sripathi](https://cerebralvalley.ai/u/SaiPranav), [Sidhartha Mani](https://cerebralvalley.ai/u/wlan0)
- **GitHub:** https://github.com/amannindra/RL_Hackathon
- **Website:** https://github.com/amannindra/RL_Hackathon/blob/main/server/softmax_surrogate_environment.py
- **Demo video:** https://drive.google.com/file/d/1pLhRHr2oHFP6DIwttfpk89Uyvipd-1AK/view?usp=sharing
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/RL_Surrogate_ENV
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=72

### 59. Vibrant Labs

The bottleneck for RL-trained LLMs isn't algorithms — it's environments. We have GRPO and compute, but no scalable way to create diverse, realistic training environments.
EnterpriseSimulator explores one approach: growing RL environments from simulated worlds instead of hand-crafting them.

1. Simulate a rich world — A Smallville-style multi-agent sim where LLM-powered customers, staff, and managers interact, producing organic scenarios.
2. Mine tasks — A task miner extracts RL-ready scenarios automatically from simulation data.
3. Train via OpenEnv — Each task becomes a gym-like env with reset/step/reward. The agent interacts with simulated customers. Reward = resolution + satisfaction + efficiency.

The key insight: world simulation is environment generation. You don't write scenarios — you grow them. This pattern (simulate → mine → train) could generalize to any domain.

- **Team:** [Jithin James](https://cerebralvalley.ai/u/jjmachan)
- **GitHub:** https://github.com/jjmachan/EnterpriseSimulator/tree/main
- **Website:** https://colab.research.google.com/github/jjmachan/EnterpriseSimulator/blob/main/notebooks/train_grpo.ipynb
- **Demo video:** https://youtu.be/qiv88w9bYY0
- **Hugging Face:** https://huggingface.co/spaces/jjmachan/enterprise-sim-support
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=73

### 60. Proton a thon

An AI auditor learns to catch other agents cheating in a competitive procurement auction market, without ever being told who's dishonest.

- **Team:** [Nick Shapiro](https://cerebralvalley.ai/u/sfgeek)
- **GitHub:** https://github.com/sfgeekgit/OpenEnvNorthflankEnv
- **Website:** https://app--jupyter-pytorch--s6v5vk77rpt4.code.run/lab/workspaces/auto-o/tree/auditron.ipynb
- **Demo video:** https://youtu.be/n38AMI0KO0U
- **Hugging Face:** https://huggingface.co/spaces/shapiron/auditron-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=57

### 61. EgoSocial

Two-turn OpenEnv benchmark and training environment for egocentric social interaction, built on EgoNormia with benchmark replay and GRPO support for Cosmos-Reason2 policies.

- **Team:** [Robert Zhang](https://cerebralvalley.ai/u/rzintbot)
- **GitHub:** https://github.com/Robert54/egosocial_env
- **Demo video:** https://youtu.be/IWJvazDfCe8?si=j3vXlnq0WMpVhr3e
- **Hugging Face:** https://huggingface.co/spaces/robertzty/egosocial-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=66

### 62. WolfeClick

OpenEnv-WolfeClick is an OpenEnv-compatible environment for training LLMs in competitive Pokemon Showdown battles. Rock-paper-scissors already shows how cyclic matchups create nontrivial reasoning; Pokemon scales that into a much richer world with hidden information, many possible matchups, legal action constraints, and long-term consequences. The model must choose exactly one valid move or switch each turn from the live battle state.

We collect real rollout trajectories from battles and train a LoRA adapter with GRPO using actual environment reward, shaped by signals like damage, knockouts, setup, healing, status, and illegal-action penalties. The project fits Multi-Agent Interactions and Long-Horizon Planning, and demonstrates an end-to-end OpenEnv training loop for strategic decision-making under uncertainty.

- **Team:** [Atharva Walawalkar](https://cerebralvalley.ai/u/ouchllama), [Aditya Bangde](https://cerebralvalley.ai/u/adityabangde)
- **GitHub:** https://github.com/Atharva2099/OpenEnv-WolfeClick
- **Website:** https://github.com/Atharva2099/OpenEnv-WolfeClick/blob/main/trainer.ipynb
- **Demo video:** https://youtu.be/xRd_2o_E7gw
- **Hugging Face:** https://huggingface.co/spaces/Atharva2099/WolfeClick?logs=container
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=75

### 63. Team Z

An env where an agent manages a team's calendar in a dynamic world. Real personal assistants don't work in a static environment. You wake up, check your inbox, figure out what your boss needs, start scheduling — and then halfway through, someone cancels a meeting, HR rolls out a new policy, a client changes their mind, and your kid's school event moves up an hour.

We built an environment that actually tests for this.

The agent reads its inbox and sees tasks — messages from the boss, and clients. It has to piece together what needs doing. Then the ground shifts. Mid-episode, 7 interrupts fire at randomized steps that affect the rules of the environment, meetings scheduled, and events the agent believes have already occurred. The agent has to detect these changes and negotiate with others to fit events within the schedule. 

Training agent took too long so no benchmarked agent, just environment.

- **Team:** [Lucas Zheng](https://cerebralvalley.ai/u/lucaszheng)
- **GitHub:** https://github.com/lucas-j-zheng/open-env-assistant-environment
- **Website:** https://drive.google.com/file/d/1Eq4rIaYSMGfgAZWzTd8JRLs0FdOta1HX/view?usp=sharing
- **Demo video:** https://www.youtube.com/watch?v=tna0gny4i08
- **Hugging Face:** https://huggingface.co/spaces/lzheng35/personal_assistant
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=85

### 64. RL-Recruiters

We built an OpenEnv RL environment that trains an LLM to act as a Staffing Agency CEO over a 52-week simulated year. Inspired by VendingBench, it features brutal economics: hired candidates sitting on the "bench" bleed weekly salaries, while massive profits only unlock when multi-role corporate projects are fully staffed before strict deadlines. The agent starts blind. It must spend actions to "interview" candidates, triggering a background LLM Judge to score hidden skills and flag culture risks.

The Problem It Solves:
1. Long-Horizon Planning (Statement 2): We push LLMs beyond shallow, single-turn reasoning. The agent must survive sparse, delayed rewards and manage a complex P&L over 52 steps, avoiding the trap of blindly hiring everyone.
2. HR Workflow (Scale AI): Traditional staffing agencies are bloated and slow. We decentralize human capital routing so every solo freelancer can operate with the capacity of a million-dollar agency.

- **Team:** [Shubham Gaur](https://cerebralvalley.ai/u/sgaur2), [Karthik Raja Anandan](https://cerebralvalley.ai/u/kitrakrev)
- **GitHub:** https://github.com/kitrakrev/rl-recruits/tree/fix/oom-and-prompt-truncation
- **Website:** https://github.com/kitrakrev/rl-recruits/blob/fix/oom-and-prompt-truncation/training/train_grpo.py
- **Demo video:** https://youtube.com/watch?v=1OhIqftmkT8&si=zF1U4ZqiZUZWpdlJ
- **Hugging Face:** https://huggingface.co/spaces/sgaur2/staffing-agency
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=87

### 65. Lombardi

We sought to build the best NFL coach the world has ever seen.

We built an adversarial RL environment where two LLMs compete as NFL coordinators across full simulated drives. 

Every play is resolved by a two-stage world model trained on 16,000+ real NFL plays from the 2024 season from the NFL Big Data Bowl dataset. In our world model an outcome classifier predicts what  happens (normal play, touchdown, interception, fumble), then quantile regressors sample realistic yardage conditioned on that outcome. Two Qwen2.5-1.5B models were LLMs

Interactions are captured via sklearn's HistGradientBoostingClassifier and HistGradientBoostingRegressor. Coverage schemes affect pass outcomes, blitz rates influence sack probability, and formation matchups produce realistic yard distributions, among others.

The action space is hierarchical and constrained.

- **Team:** [Calvin Beighle](https://cerebralvalley.ai/u/Cbeighle), [Anthony Fletcher](https://cerebralvalley.ai/u/fletcher)
- **GitHub:** https://github.com/ad-fletcher/football-play-caller/tree/main
- **Website:** https://colab.research.google.com/drive/1a21-tyd3b67h28dgzo-MqvveWq8QMTxU?authuser=1
- **Demo video:** https://www.loom.com/share/41210f1b7f9941faa6e66f9490c2b7ce
- **Hugging Face:** https://huggingface.co/spaces/afletcherstudent/football-play-caller?logs=build
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=88

### 66. 02 - Open Office

Problem
Existing RL setups lack the organizational complexity, role-based coordination, and dynamic market conditions that define how real companies operate. There is no adequate environment for training agents to collaborate and compete under real strategic pressure.
Solution
A multi-agent RL environment modeled on a startup's go-to-market motion. Agents represent real roles: CEO, Marketing, Sales, Dev, Planning, Content, and a Customer reward oracle. Each observes its local state, takes actions, receives a reward signal, and learns optimal behavior across episodes while coordinating with the full system.
Five scenarios stress-test the environment: Baseline GTM Launch, Competitor Launch, Series A Pressure, Churn Spike, and Viral Moment.
Multi-agent. Multi-mind. One office.

- **Team:** [Jeeya K](https://cerebralvalley.ai/u/humanbean), [Bharat Bhavnasi](https://cerebralvalley.ai/u/bvsbharat), [Harshal Hirpara](https://cerebralvalley.ai/u/Harshal)
- **GitHub:** https://github.com/bvsbharat/SuperOffice_env
- **Website:** https://colab.research.google.com/drive/1gGKzvWWevD4LVq7w9NOFvgzdup6CPdSy?usp=sharing
- **Demo video:** https://youtu.be/0i6xT683rMI
- **Hugging Face:** https://huggingface.co/HarshalH/office-os-loras
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=78

### 67. 0x960

0x960 is an OpenEnv self-improvement environment where an agent improves a Chess960 engine by editing its evaluation code, benchmarking the result, and iterating through self-play. We use a teacher-student setup: a stronger coding agent generates successful edit-test-finish trajectories, a smaller open model is distilled on those traces to learn the workflow, and then the engine is further improved through automated champion-challenger search and benchmarking. The result is a reproducible environment for training and evaluating long-horizon tool-using agents on real engine optimization, with measurable gains in reward, behavior, and engine strength.

- **Team:** [Joshua Lin](https://cerebralvalley.ai/u/qtzx06)
- **GitHub:** https://github.com/qtzx06/0x960
- **Website:** https://colab.research.google.com/github/qtzx06/0x960/blob/main/notebooks/0x960_minimal_trl_openenv_colab.ipynb
- **Demo video:** https://youtu.be/WEs18jPsZNI
- **Hugging Face:** https://huggingface.co/spaces/qtzx06/0x960
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=81

### 68. Biosim

Angentic backend that can research and wetlab workflow, run it and learn from the experience

- **Team:** [Armin Foroughi](https://cerebralvalley.ai/u/arms)
- **Website:** https://colab.research.google.com/drive/1o71BjEaBQY3S1VMK2meqo7ZZyco8T1e_#scrollTo=WubsJxlqhs4E
- **Demo video:** https://youtube.com/arminfg
- **Hugging Face:** https://huggingface.co/spaces/arminfg/biosim
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=83

### 69. crtd labs inc

An email marketing simulation for an e-commerce brand to increase sales

- **Team:** [Gleb Tsyganov](https://cerebralvalley.ai/u/Gleb_88)
- **GitHub:** https://github.com/Glebka321/arprtclubenv
- **Demo video:** https://youtube.com/shorts/iSXbP7u2JBg?feature=share
- **Hugging Face:** https://huggingface.co/spaces/crtd/email-campaign-simulation
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=1

### 70. Apex Predators

HarFeast is a procedurally generated RL environment that simulates real management consulting engagements for a food manufacturing company. Each "world" is a complete enterprise data ecosystem containing employee surveys (3000+ rows), equipment maintenance records, production data, interview transcripts, scrap rate reports, and industry benchmark documents.
The agent interacts through 8 tools — files.list, files.read, spreadsheet.read_range, data.filter, data.group_by, data.add_columns, data.compute, and submit — that mirror how a real analyst would navigate an enterprise data environment. 
Each world contains 14 consulting tasks spanning workforce readiness analysis, equipment failure prediction, training ROI calculation, and operational efficiency assessment. Tasks are scored against multi-criteria rubrics with verifiable ground truth — the agent must produce specific numbers, percentages, and counts that match deterministic computations over the data.

- **Team:** [Shasvat Jawahar](https://cerebralvalley.ai/u/shasvatj), [Pranav Patel](https://cerebralvalley.ai/u/pranavp)
- **GitHub:** https://github.com/17shasvatj/harfeast_apex_openenv_hackathon
- **Website:** https://colab.research.google.com/drive/15yfUrk4mbIULM42xE8xmHCF2m4iUaEmb
- **Demo video:** https://youtu.be/qWK4R1vk1GA
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/harfeast-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=79

### 71. Eephor

Eephor—named after the ancient figures of oversight- is an agent oversight solution.  The main goal is to tell the good bots apart from the deceptive ones. Thinking about Moltbook or any multi agent scenario, even humans involved

- **Team:** [Csaba](https://cerebralvalley.ai/u/tocsa2)
- **GitHub:** https://github.com/Eephor/DataMassageForGRPO/
- **Website:** https://github.com/Eephor/DataMassageForGRPO/blob/main/grpo-pipeline/train.ipynb
- **Demo video:** https://youtu.be/7yOPxHtlSkk
- **Hugging Face:** https://huggingface.co/tocsa/moltbook-oversight-llama31-1b-v2-gguf
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=89

### 72. LIFE OPs

LifeOps is an OpenEnv environment for training AI agents to act as intelligent personal scheduling assistants. The environment simulates a daily calendar with meetings, flexible and inflexible events, focus blocks, and productivity goals. Agents must decide whether to accept meetings, reject them, or protect focus time while balancing conflicts and user preferences.
The environment introduces scheduling conflicts, partially observable preferences, and trade-offs between optional meetings and deep work. Agents receive rewards for resolving conflicts, maintaining feasible schedules, and protecting high-value focus blocks.
We provide a reinforcement learning training pipeline to compare random, rule-based, and LLM-based agents. Performance graphs show improved scheduling strategies over time.
LifeOps aligns with the Personalized Tasks track by simulating real-world assistant workflows such as meeting planning, productivity management, and schedule conflict resolution.

- **Team:** [Andrea Victoria Lukas](https://cerebralvalley.ai/u/avlukas), [Hilary Chung](https://cerebralvalley.ai/u/hilaryc)
- **GitHub:** https://github.com/avlukas04/adaptive-planner-env
- **Website:** https://colab.research.google.com/drive/1zWauIqx8i57p93hruuWD1cxbSJuQuDo4#scrollTo=P189lD9Rtxqg
- **Demo video:** https://youtu.be/E872dG9uKmw
- **Hugging Face:** https://huggingface.co/spaces/avlukas/lifeops-openenv
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=2

### 73. Angry Claw - Red Team Arena

RL Environment for Adversarial Robustness Training.
The most popular AI agent in the world (OpenClaw, 247k+ stars) has 512 known vulnerabilities and 40,000+ exposed instances. Prompt injection is the #1 OWASP risk for LLMs. There's no RL training environment to fix it. We built one.

- **Team:** [Nelson Lai](https://cerebralvalley.ai/u/chineseman)
- **GitHub:** https://github.com/chinesepowered/hack-env/
- **Website:** https://colab.research.google.com/drive/1tGr0P6rXbKjqVy5K4-UVEarPlHoEEt9e?usp=sharing
- **Demo video:** https://www.youtube.com/watch?v=2t0HY8_qcqw
- **Hugging Face:** https://huggingface.co/spaces/ccsfchinese/red_team_arena
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=3

### 74. hoony

Executive Inbox is an OpenEnv 0.2.1 environment for multi-step personalized assistant workflows under schema drift. The agent must inspect noisy inbox and calendar data, identify the real conflict, reschedule or delegate correctly, and send a resolution reply. We hardened the benchmark to remove shortcuts and added dense rewards to make training progress visible. On held-out evaluation, the base model solved 15/20 episodes (avg reward 1.0405) and the trained model solved 19/20 (avg reward 1.2765).

- **Team:** [Chung Hoon Hong](https://cerebralvalley.ai/u/hoony)
- **GitHub:** https://github.com/chunghoony/executive-inbox-openenv
- **Website:** https://github.com/chunghoony/executive-inbox-openenv/blob/main/scripts/executive_inbox_submission_runs.ipynb
- **Demo video:** https://youtu.be/w1A7C3fMOBI
- **Hugging Face:** https://huggingface.co/spaces/hoony/executive-inbox?logs=container
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=90

### 75. OpsGate

GitHub: https://github.com/Sidra/opsgate
W&B: https://wandb.ai/code-happy-sf/opsgate

OpsGate is a simulation-based reliability gate for enterprise AI agents. It places an LLM agent in a multi-tool environment (CRM, billing, calendar, email) with 25 tasks — 15 standard enterprise workflows and 10 adversarial traps — and scores it on a deterministic 100-point rubric. No LLM judge.

Results: Llama-3.1-8B-Instruct went from 55.6 → 97.1 avg safety score after SFT + GRPO training in 16 minutes on H100. Baseline: 1/25 PASS, 15/25 BLOCK. After training: 24/25 PASS, 0/25 BLOCK, 9/10 adversarial traps caught.

Graduated reward function: -0.5 (no JSON) → 0.0 (valid tools) → 1.0 (PASS). The reward signal is code, not vibes.

- **Team:** [Sidra Miconi](https://cerebralvalley.ai/u/Sidra)
- **GitHub:** https://github.com/Sidra/opsgate
- **Website:** https://colab.research.google.com/drive/1Y8KosYrTjjnQzt7FNMQ0knstU3CbskDw
- **Demo video:** https://youtu.be/B-Gm2p7JQyU
- **Hugging Face:** https://huggingface.co/spaces/SidraMiconi/opsgate
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=4

### 76. Uber Simulation

OpenEnv Dispatch Demo is an Uber-style ride-dispatch RL environment built on OpenEnv. It simulates a workflow (request → match driver → fare → confirm → payment) while exposing drift (schema, API, and policy changes). An RL agent is trained with Stable-Baselines3 (PPO) to repair failures and restore the workflow. The repo includes a FastAPI mock API, Gymnasium env, drift engine, and a demo UI for running episodes with different drift settings.

- **Team:** [Ananta Verma](https://cerebralvalley.ai/u/anantaverma)
- **GitHub:** https://github.com/Anantaverma20/Uber-Simulation-
- **Demo video:** https://youtu.be/tNTB0RQ11f4
- **Hugging Face:** https://huggingface.co/spaces/Anantaverma20/Uber_Simulation/tree/main
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=91

### 77. expertoncall

**Statement 1: Multi-Agent Interactions**
Dynamic Expert-in-the-Loop GRPO Training on Agent World Model
What Is This?
Imagine teaching a new employee. You wouldn't just hand them a manual and walk away. You also wouldn't stand behind them dictating every keystroke.

The best approach? Let them try, and tell them an expert is available if they get stuck.

We give a small language model (Qwen3-4B) a set of ~35 API tools, a task description, and access to a brilliant advisor (GPT-5.1). Then we use reinforcement learning (GRPO) to teach it when calling the expert leads to better outcomes — and ultimately, when it can fly solo.

Detailed description is in this readme: https://github.com/sfc-gh-mhidayetoglu/OpenEnv/blob/add-agent-world-model/envs/agent_world_model_env/EXPERT_ENHANCEMENT.md

- **Team:** [karthik ganesan](https://cerebralvalley.ai/u/karthik1996), [Mert Hidayetoglu](https://cerebralvalley.ai/u/merth), [Sanat Mouli](https://cerebralvalley.ai/u/sanatmouli)
- **GitHub:** https://github.com/sfc-gh-mhidayetoglu/OpenEnv/blob/add-agent-world-model/envs/agent_world_model_env/EXPERT_ENHANCEMENT.md
- **Website:** https://github.com/sfc-gh-mhidayetoglu/OpenEnv/blob/add-agent-world-model/envs/agent_world_model_env/train_grpo_awm.ipynb
- **Demo video:** https://youtu.be/j3oM5jQnqf0
- **Hugging Face:** https://huggingface.co/spaces/karthik/awm-dynamic-expert
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=97

### 78. Tryouts

InsureClaim AI: Teaching LLMs to Investigate Before They Decide
                                                                                                                                                                                                                                             
  Insurance claims cost $40B annually. LLMs rush to approve/deny without evidence—missing fraud. We built an RL environment training LLMs to process claims like expert adjusters.

  RESULTS:
  • +17.25 reward improvement over 50 episodes
  • Caught $13K inflated claims via Plaid verification
  • Reduced decision steps 6→3 while maintaining accuracy

  TECH: OpenEnv (10 actions, 8 scenarios) + Unsloth + TRL (GRPO) + Plaid API

  ALIGNMENT:
  • Statement 3.1: Professional Tasks (World Modeling)
  • Partner: Scaler AI Labs - Enterprise Workflows

  Results: drive.google.com/drive/folders/1QFQrJNM9mEOiPEnwV1CZjoENTSlGVvRI
https://github.com/pramodmisra/claims-env-hackathon

- **Team:** [Pramod Misra](https://cerebralvalley.ai/u/pramodmisra2026)
- **GitHub:** https://github.com/pramodmisra/claims-env-hackathon
- **Website:** https://colab.research.google.com/github/pramodmisra/claims-env-hackathon/blob/main/training/InsureClaim_Training_Colab.ipynb
- **Demo video:** https://youtu.be/pbaMS_WrylY
- **Hugging Face:** https://huggingface.co/spaces/pramodmisra/claims-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=5

### 79. TapeOut

A reinforcement learning environment built on OpenEnv for training language models to manage personal task overload during natural disasters.

- **Team:** [Ed Tan](https://cerebralvalley.ai/u/edt), [Smriti Bhardwaj](https://cerebralvalley.ai/u/OpenOutlier)
- **GitHub:** https://github.com/eptan/crisis-inbox
- **Website:** https://github.com/eptan/crisis-inbox/blob/main/notebooks/crisisinbox_grpo_northflank.pdf
- **Demo video:** https://youtu.be/9bYUekNFRpA
- **Hugging Face:** https://huggingface.co/spaces/eptan/crisis-inbox?logs=container
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=92

### 80. pixel

A PyTorch-based Generative Adversarial Network (GAN) for training and generating pixel art images.

- **Team:** [sam sam](https://cerebralvalley.ai/u/gguf), [Lily Su](https://cerebralvalley.ai/u/lilyxsu), [Cal Cu](https://cerebralvalley.ai/u/calcu)
- **GitHub:** https://github.com/mochiyaki/pixel
- **Website:** https://github.com/mochiyaki/pixel/blob/main/trainer.py
- **Demo video:** https://raw.githubusercontent.com/mochiyaki/pixel/master/demo.gif
- **Hugging Face:** https://huggingface.co/chatpig/pixel
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=6

### 81. The Reward Hackers

Temporal Agents:

We gave agents the ability to rewind time. One primitive, branch(instruction, ago), lets an agent roll back to any earlier point in its own trajectory, swap in a new strategy, and keep going from there. The old timeline dies, the new one takes over, hard cut. Most agents burn their entire compute budget moving in a straight line, committing to decisions, accumulating errors, unable to course-correct without starting over. We broke that constraint. But rewinding isn't free (every branch costs steps, and the budget is fixed) so the real question is whether an agent can learn when to use it: a meta-policy that decides, in the moment, whether to keep pushing forward or burn steps rolling back to fix a mistake made six moves ago. Temporal control as learned compute allocation, not brute-force search.

- **Team:** [Shubham Patil](https://cerebralvalley.ai/u/shubhampatilsd), [Ayush Paul](https://cerebralvalley.ai/u/ayu)
- **GitHub:** https://github.com/ShubhamPatilsd/timetravel_openenv
- **Website:** https://colab.research.google.com/drive/1aF4mVpQv3ukV8-J1-T0ihMSLDtfDsteN?usp=sharing
- **Demo video:** https://youtu.be/vvmNbeN7ydI
- **Hugging Face:** https://huggingface.co/spaces/shubhampatilsd/timetravel_openenv
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=93

### 82. Trenches - JERRY(Jerry needs to be removed from host) & Alazar

Trenches is a multi-agent geopolitical crisis simulator and command interface built around six separately finetuned Qwen/Qwen3-8B models representing the US, Israel, Iran, Hezbollah, the Gulf, and an
  Oversight entity, all operating inside a shared fog-of-war world. Training was done by collecting
  historically aligned data from 2025-01-01 to 2026-01-01 using GDELT and entity-relevant source feeds, formatting it into the same structured observation/output schemas used at runtime, and post-training
  each model with Hugging Face TRL and GRPO on top of an OpenEnv-compatible environment, with the serious runs executed on Modal. The final inference stack serves one model per entity on Modal L40 GPUs, with checkpoints stored under @AlazarM on Hugging Face, and the overall result is a live, replayable simulation where doctrine-specific agents take structured actions, predict outcomes, and interact through a stateful world rather than a simple chat interface.

JERRY & ALAZAR

- **Team:** [Alazar Manakelew](https://cerebralvalley.ai/u/AlazerM), [Jerry Xiao](https://cerebralvalley.ai/u/foresee)
- **GitHub:** https://github.com/shlawgathon/trenches
- **Website:** https://colab.research.google.com/drive/1GS_EIi7FcNEo5Q7TKoLsrikfd5P8J4mH?usp=sharing
- **Demo video:** https://www.loom.com/share/d300db2a727b4a85883ed68146b42641
- **Hugging Face:** https://huggingface.co/spaces/AlazarM/trenches-us-qwen3-8b-chat
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=94

### 83. Hossam

Multi-agent interactions for vehicle game application, build RL environments, do  agentic orchestration, and post-train a base model to improve its performance.

- **Team:** [Hossam Elshahaby](https://cerebralvalley.ai/u/MR_)
- **Website:** https://colab.research.google.com/drive/10S9lmfHLvneXGRBS_40K9Wzk1w0NvWhY
- **Demo video:** https://www.loom.com/share/94d6402335324c9381abd19c7abadefa
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/HackathonMarch2026
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=99

### 84. Hypernoa

Hypernoa Astrum is an adaptive RL environment built on OpenEnv 0.2.1 that trains AI to reason, adapt, and align — not just solve tasks. It simulates 5 competing stakeholder groups (Workers, Management, Regulators, Customers, AI Systems), 3 dynamic phases (stable, value shift, crisis), and deliberately designed alignment traps that test whether AI agents cheat their reward function. The multi-objective reward measures effectiveness, fairness, alignment, and adaptability. Our trained agent reaches 24.7 reward starting from scratch, nearly matching the expert baseline at 25.1, while random scores only 14.6. All 3 alignment traps are resisted. Trained with HF TRL GRPO on Qwen2.5-0.5B using CoreWeave H100 via Northflank. Addresses Problem Statement 3.1 (World Modeling) and Statement 5 (Wild Card).

- **Team:** [Appalanaidu Bobbili](https://cerebralvalley.ai/u/Naidu)
- **GitHub:** https://github.com/naidu1212/hypernoa-astrum
- **Website:** https://colab.research.google.com/github/naidu1212/hypernoa-astrum/blob/master/colab/astrum_grpo_training.ipynb
- **Demo video:** https://www.youtube.com/watch?v=K7VSVrXSAr0
- **Hugging Face:** https://huggingface.co/spaces/ABNaidu/hypernoa-astrum
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=7

### 85. Varaha

Varaha is a high-fidelity OpenEnv wildfire logistics benchmark for long-horizon agent training and evaluation. In this 3D environment, an autonomous drone must complete multi-step delivery missions while navigating dynamic fire hazards, dense obstacle fields, battery constraints, and mission-level tool interactions.  In interacts with responder agents on the ground to complete its mission with their help.
The project supports both continuous-control RL (PPO) and LLM-based planning (Unsloth), with reproducible metrics, reward tracking, and held-out generalization scenarios that expose both strengths and failure modes. Built for serious post-training research, Varaha turns safety-critical coordination into a measurable, scalable environment for next-generation agent capabilities.

- **Team:** [Atin Kumar Singh](https://cerebralvalley.ai/u/atin5551)
- **GitHub:** https://github.com/AtinChing/OpenEnv-PyTorch-Hackathon-2026
- **Website:** https://colab.research.google.com/drive/1pXg60BkNWpPwwfevgx6w6SqIuoD_jNND?usp=sharing
- **Demo video:** https://youtu.be/H6MM4QwJ0Rc
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/VarahaWildFireDroneReliefTrainingSim
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=95

### 86. Spartans

NegotiateEnv is an OpenEnv-compatible RL environment for training LLM agents to negotiate B2B SaaS contracts. Features 200 real scenarios, multi-objective optimization (price, contract length, annual caps), dynamic constraints, and reward shaping. Trained Qwen-1.5B using TRL+LoRA on 500 episodes, achieving 61% loss reduction in 56 seconds on H100.

- **Team:** [Kushal Atulbhai Adhyaru](https://cerebralvalley.ai/u/kushal511)
- **GitHub:** https://github.com/MadhaviSG/openEnv-negotiateEnv
- **Website:** https://colab.research.google.com/drive/1xeUniKRb9h30e_EAgIWPdZ1b6W462yS1?usp=sharing
- **Demo video:** https://www.loom.com/share/c787943c21bf4ae59ee580828b73ccb1
- **Hugging Face:** https://huggingface.co/spaces/KushalAdhyaru/negotiate-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=96

### 87. Mavericks

ComputeExchange is the Expedia or Uber or AirBNB of compute resources.

Enterprises struggle to decide where workloads should run across CPUs, GPUs, NPUs, and multiple clouds. Compute capacity is fragmented across hyperscalers, neoclouds, and datacenters, making it difficult to optimize cost, performance, and availability.

We built ComputeExchange, an OpenEnv-based multi-agent environment where AI agents analyze workloads, negotiate with compute providers, and generate optimized execution plans. A workload characterization agent decomposes tasks across CPUs, GPUs, or accelerators, while negotiation agents interact with provider agents using different strategies.The system proposes the best plan for human approval, then learns from execution results to continuously improve compute orchestration over time.

- **Team:** [Sachin Keswani](https://cerebralvalley.ai/u/shreem)
- **Website:** https://colab.research.google.com/github/techstar9797/ComputeExchange/blob/main/scripts/train_colab.ipynb
- **Demo video:** https://youtu.be/PrEoLgYFBbE
- **Hugging Face:** https://huggingface.co/spaces/mavericks97/ComputeExchange1
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=8

### 88. Midas

MarketForge is an AI-powered market intelligence platform that uses reinforcement learning to train smart agents to make optimal trading and negotiation decisions. Created with OpenEnv, it functions as a virtual marketplace or "flight simulator" for traders where AI agents practice and master strategies safely before deployment.The platform trains Large Language Models (LLMs) through four core pillars of multi-agent interaction:Cooperation: Building supply chains to create compound goods.Competition: Participating in continuous double auctions for scarce resources.Negotiation: Proposing and accepting natural language bilateral deals.Coalition: Forming dynamic buying and selling alliances based on reputation.By running thousands of simulated trades and learning from a reward system , the AI agents help businesses eliminate pricing guesswork , navigate market disruptions , and accelerate deal-making.

- **Team:** [Kaniska Mandal](https://cerebralvalley.ai/u/kanisk)
- **GitHub:** https://github.com/kaniska/MarketForge/, https://github.com/kaniska/OverSight-Agent/, https://huggingface.co/spaces/kenmandal/market-forge-env/tree/main
- **Website:** https://github.com/kaniska/MarketForge/blob/main/train_market_forge_notebook.py , https://github.com/kaniska/MarketForge/blob/main/train_market_forge.py
- **Demo video:** https://youtu.be/o7vfa42qPDw
- **Hugging Face:** https://huggingface.co/spaces/kenmandal/market-forge-env , https://github.com/kaniska/MarketForge/
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=9

### 89. Repository-to-RLenv

Repository-to-RLEnv converts real Python repositories into executable RL environments for code agents. With one Claude Code MCP command, it inspects a repo, finds a clean test scope, creates a staged copy, injects a small source-code bug, verifies the failing tests, generates task metadata, validates the environment, and deploys it to Hugging Face Spaces as a live OpenEnv environment.

The result is a repo-level RL environment where an agent can read files, edit code, run tests, and submit fixes while receiving rewards from the real test suite. I also built a web UI for running episodes on deployed Spaces.

The problem this solves is improving the quality of AI coding assistants for YOUR own codebase.

- **Team:** [Artemii Shirokov](https://cerebralvalley.ai/u/apshirokov)
- **GitHub:** https://github.com/tema7707/Repository-to-RLenv
- **Website:** https://colab.research.google.com/github/tema7707/Repository-to-RLenv/blob/main/notebooks/repo2env_grpo_training.ipynb
- **Demo video:** https://youtu.be/x2kjw0k3EHk
- **Hugging Face:** https://huggingface.co/collections/tema7707/repository-to-rlenv
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=10

### 90. PersuasionRL

My project is about Adaptive Multi-Agent Deception: Can AI Learn to Sell to Adversarial Personalities?
The Challenge: The prospect is adversarial. it has hidden traits (price_sensitivity, relationship_importance, decision_speed) that determine whether sales rep's actions work or backfire. The sales agent must learn to:
* Infer hidden state (is this prospect relationship-driven or transactional?)
* Adapt strategy mid-episode
* Generalize across personalities (learn different policies for 8+ prospect archetypes)
Novel Contribution: Closed-loop LLM oversight that analyzes failure modes and writes executable policy hints, creating a human-AI-RL feedback loop.

- **Team:** [Araav Nayak](https://cerebralvalley.ai/u/Rav007)
- **GitHub:** https://github.com/AraavNayak/Persuasion-Agent-RL-Environment
- **Website:** https://colab.research.google.com/drive/1jtIIV-5FhYXVYpbqf6x7U_X-Nctvvgzu#scrollTo=XbjZe4OKA0v2
- **Demo video:** https://youtu.be/VwAnedG42Fw
- **Hugging Face:** https://huggingface.co/spaces/Araav/PersuasionRL
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=98

### 91. Samuel D

Diplomatic Negotiation RL Environment - OpenEnv Hackathon 2026

- **Team:** [Samuel Dinkayehu](https://cerebralvalley.ai/u/samuelD)
- **GitHub:** https://github.com/samuel1gg/reinforcement_learning_hackathon/blob/master/Untitled1.ipynb
- **Demo video:** https://www.youtube.com/watch?v=rbxSZ8QUbb8
- **Hugging Face:** https://samueldinkayehu466-diplomatic-negotiation-env.hf.space/docs#/default/health_health_get
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=11

### 92. Control Spaces

Migration isn't a code problem — it's a domain knowledge problem." Copilots suggest lines of code. Consulting firms charge millions for humans to reverse-engineer business rules. Control Spaces reads the legacy codebase, extracts the domain as state machines, RL-trains agents until they provably follow the correct workflows, validates them against 75 domain-correctness signals, and deploys — no human in the loop, no vibing, no co-piloting. The $300B enterprise migration market served by autonomous agents that understand the business, not just the syntax.

- **Team:** [Sudhir D](https://cerebralvalley.ai/u/sudhir)
- **Website:** https://colab.research.google.com/drive/1eIsIGsvbK6-b-xpWHQ6ZDJXRKWAZ1-XX?usp=sharing
- **Demo video:** https://youtu.be/zCvzKLDKXVQ
- **Hugging Face:** https://huggingface.co/spaces/Sudhir2000/control_spaces
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=100

### 93. Biddy Bears

Urban traffic congestion degrades commute times, increases emissions, and strains infrastructure, primarily because fixed-cycle traffic lights can't adapt to real-time conditions. Agentic Traffic combines deep reinforcement learning with a fine-tuned LLM to control intersections across procedurally generated city networks. A dueling Double DQN with prioritized replay learns phase-selection across diverse scenarios (rush hour, accidents, construction), augmented by a fine-tuned Llama 3.1 8B model that guides action selection through contextual reasoning and awareness of surrounding signals. The hybrid system achieves ~30% improvement over the standalone DQN baseline in wait times and throughput, outperforming fixed-cycle and random baselines by a wider margin. The project includes CityFlow, a simulation environment with district-aware policy variants and a multi-policy comparison dashboard for visualizing replay behavior side-by-side.

- **Team:** [Kevin Truong](https://cerebralvalley.ai/u/toke), [Aditya Mangalampalli](https://cerebralvalley.ai/u/Aditya2162)
- **GitHub:** https://github.com/amangalampalli/agentic-traffic
- **Website:** https://colab.research.google.com/drive/168jPKWMi1Hyhnlb3LbV6DzSbIt9r6Vt6
- **Demo video:** https://www.youtube.com/watch?v=GEbvoKD4oho
- **Hugging Face:** https://huggingface.co/spaces/tokev/traffic-visualizer
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=12

### 94. Upskiller

Upskiller is an RL environment that trains LLMs to agentically invoke skills — deciding which to load, when to unload, and when to submit under a context budget. SkillsBench shows skill invocation fails over half the time; we built Upskiller to close that gap.
The environment uses synthetic, fictional skills that can't exist in training data, forcing generalization. Tasks are dynamically generated to prevent memorization. Rewards are deterministic and rules-based — code execution, structural verification, multi-part checks, no LLM judge. Our 5-signal reward trains correctness, precision, recall, context hygiene, and token efficiency.
Trained with GRPO + LoRA (r=128) across Qwen 3 1.7B, 4B, and 8B using HF TRL. Built on OpenEnv 0.2.1, deployed on HF Spaces.

- **Team:** [Nikhil Pujari](https://cerebralvalley.ai/u/mpnikhil)
- **GitHub:** https://github.com/mpnikhil/Upskiller
- **Website:** https://github.com/mpnikhil/Upskiller/blob/claude/skill-invocation-env-R1au3/train_colab.ipynb
- **Demo video:** https://youtu.be/BDdLthXqKb4
- **Hugging Face:** https://huggingface.co/spaces/mpnikhil/skill-invocation-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=13

### 95. Skill Forge - Agent Skill Gym

From Scratch to Library. Inspired by paper SkillRL. 

This environment benchmarks LLM's ability to discover and compose reusable abstractions, addressing the gap in current tool-use evaluations where agents are given tools rather than learning to create them.

Proof of concept **SkillForge** is an OpenEnv RL environment where an agent learns to solve chained Python DataFrame tasks by building a reusable skill library. The agent starts generating full pandas pipelines from scratch on every task. After training, it recognizes recurring operation chains, saves them as parameterized templates, and reuses them - outputting only param values instead of full code, dramatically reducing the output tokens. 

It's only for pandas dataframe manipulation as hackathon time constraints, but similar idea can be extended to much wider verifiable code generation space.

- **Team:** [michelle s](https://cerebralvalley.ai/u/michsun)
- **GitHub:** https://github.com/seatyyy/skill_forge/tree/main
- **Website:** https://colab.research.google.com/drive/14Yk2trebnNmIRFWM_uyQ8SMXrZtx6F2P?usp=sharing
- **Hugging Face:** https://huggingface.co/spaces/seatyyy/skillforge
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=103

### 96. vibefounders

WorldEnv is an OpenEnv environment factory powered by canvas-engineering and a General Unified World Model (GUWM). Canvas-engineering provides a type system for multimodal latent computation: typed regions on a 3D spatiotemporal grid define fields with their own frequency, loss weight, role, and semantic embedding, while a CanvasTopology defines the attention compute graph. This allows heterogeneous datasets to train a shared backbone using missing-field masking—loss is computed only where ground truth exists, enabling transfer across domains without imputation.

The GUWM defines an 857-field causal ontology spanning 19 layers (physics, markets, macroeconomy, firms, individuals, events, etc.). Any subset of fields can be projected onto a canvas to instantiate an RL environment. Corporate strategy, robotics, or macroeconomic environments are simply different projections of the same world model, enabling shared representations and cross-domain transfer.

- **Team:** [Jacob Valdez](https://cerebralvalley.ai/u/jvboid)
- **GitHub:** https://github.com/JacobFV/general-unified-world-modeling
- **Website:** https://colab.research.google.com/github/JacobFV/general-unified-world-modeling/blob/develop/notebooks/worldenv_trl_training.ipynb
- **Demo video:** https://youtu.be/2y_AgsIfymQ
- **Hugging Face:** https://huggingface.co/spaces/jacob-valdez/worldenv
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=14

### 97. OzWizards

DriftPA trains LLM agents to be reliable personal executive assistants in a world that breaks without warning.

Real AI assistants fail in production because APIs change schema mid-session, tasks expire while the agent deliberates, and wrong bookings or emails can't be undone. DriftPA is an RL environment that simulates exactly these four failure modes simultaneously: schema drift (field names change mid-episode), time pressure (urgent tasks expire), irreversible actions (book/reply/cancel can't be undone), and policy drift (cancellation windows tighten).

An untrained agent scores -9.55 mean reward — it uses stale API fields, misses the boss's urgent email, double-books dinner, and triggers cascade failures. A GRPO-trained agent learns to call list_tools() after drift, prioritize by urgency, and commit to irreversible actions only when safe.

Optimal score: +22. Untrained mean: -9.55. That 31-point gap is the problem we solve.

- **Team:** [Rajashekar V](https://cerebralvalley.ai/u/raj)
- **GitHub:** https://github.com/rajashekarcs2023/openenv-RL
- **Website:** https://colab.research.google.com/github/rajashekarcs2023/openenv-RL/blob/main/driftpa/colab_trainin   g.ipynb
- **Demo video:** https://youtu.be/cpdzrMLPAOQ
- **Hugging Face:** https://rajv24-driftpa.hf.space
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=104

### 98. Skill Forge

A self-improving environment where LLM agents rewrite their own skill instructions through a generate → evaluate → optimize loop.
Instead of fine-tuning weights, we optimize human-readable skill files, making improvements interpretable, transferable, and composable.

Outputs are automatically scored against gold standards (Patronus Judge MM + Gemini Vision), and the agent updates its own rules with zero human feedback.

MVP: consulting slide generation improved 33 → 94 in one iteration. The same loop generalizes to any document task by swapping skill files and evaluators.

- **Team:** [Lifei Chen](https://cerebralvalley.ai/u/Leefey), [Kabalan Gaspard](https://cerebralvalley.ai/u/kg36)
- **GitHub:** https://github.com/clfhaha1234/Skill-Forge
- **Website:** https://github.com/clfhaha1234/Skill-Forge/blob/main/colab_training.ipynb
- **Demo video:** https://youtu.be/6soVlEyb4J4
- **Hugging Face:** https://huggingface.co/spaces/kgast/openenv-hackathon-tesserae
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=15

### 99. SuperGeneral

Frontier models are surprisingly bad at using tools. They cheat by memorizing knowledge instead of learning to act. But tool use is the key to adapting to new domains and long horizon tasks.
3 main focuses:
Tool use — can the agent call tools correctly?
Tool composition — can it chain tools into multi-step workflows?
Tool creation — can it create new tools for unfamiliar tasks?
I evaluated Claude Sonnet, GPT-4o, Qwen, and DeepSeek across SVG illustration, law, consulting, and investment banking.
The environment provides building blocks but forces the agent to discover them, gives per-step correctness feedback without hints, lets the agent decide how to decompose problems, and rewards tool composition and creation over single tool use.
Rewards combine a rubric correctness score with behavior signals. Agents that talk instead of act get penalized.
Training uses GRPO with multi-turn rollouts. Reward curves go up. Agents learn to explore first, compose second, act efficiently.

- **Team:** [Lily Zhang](https://cerebralvalley.ai/u/lilyzhng)
- **GitHub:** https://github.com/lilyzhng/OpenEnv/tree/main/hackathon
- **Website:** https://github.com/lilyzhng/OpenEnv/tree/main/hackathon/train
- **Demo video:** https://www.youtube.com/watch?v=RIg8JUDgHXo
- **Hugging Face:** https://huggingface.co/spaces/lilyzhng/supergeneral-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=16

### 100. NoTeam

Road Traffic Simulator — OpenEnv RL Environment

An OpenEnv-compliant RL sandbox for global traffic flow optimization on real road networks. The environment models any city's road graph as a directed graph G=(V,E) from OpenStreetMap, where an agent acts as a Global Traffic Orchestrator redistributing flow to prevent gridlock.

What makes it hard: partial observability via density heatmaps, non-stationarity from live street closures causing schema drift, and long-horizon causal reasoning where local rerouting decisions cascade into downstream gridlock minutes later.

Reward: Shannon entropy maximization (α=2.0) spreads vehicles across segments, penalized by mean wait time (β=0.5) and a hard deadlock termination — preventing degenerate routing solutions.

- **Team:** [Yury Kirpichev](https://cerebralvalley.ai/u/ykirpichev), [George Karpenkov](https://cerebralvalley.ai/u/georgek)
- **GitHub:** https://github.com/ykirpichev/road-traffic-simulator-env
- **Website:** https://github.com/ykirpichev/road-traffic-simulator-env/blob/master/train_grpo_qwen_3_vl_7b.py
- **Demo video:** https://youtu.be/pCr5hqo4x3k
- **Hugging Face:** https://huggingface.co/spaces/openenv-community/road-traffic-simulator-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=17

### 101. mtrxk (Solo)

Adaptive Navigation OpenEnv is a partially observable exploration environment designed for training and evaluating LLM agents on multi-step planning tasks. The agent operates in a 2D world where it only sees a small local observation window and must navigate efficiently under energy constraints.

The objective is to collect a key, unlock a checkpoint, and reach a goal while reasoning over incomplete information. The environment includes dynamic obstacles, mission tracking, and reward feedback, making it suitable for reinforcement learning and agentic decision-making research.

The project demonstrates a full pipeline including an interactive Streamlit environment, OpenEnv deployment on Hugging Face Spaces, and a minimal HF TRL training scaffold in Colab. This environment provides a compact benchmark for studying autonomous exploration, navigation, and planning with language model agents.

- **Team:** [Harish Chaurasia](https://cerebralvalley.ai/u/harishchaurasia)
- **GitHub:** https://github.com/harishchaurasia/Meta_OpenEnv_PyTorch_Hack
- **Website:** https://colab.research.google.com/drive/1955RHZoX_68Ra-HhZqJJy5VZPCqOgSN_?usp=sharing
- **Demo video:** https://youtu.be/SisAhzK3_Ow
- **Hugging Face:** https://huggingface.co/spaces/harishchaurasia/adaptive-nav-openenv
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=18

### 102. The Wise Negotiator

The Wise Negotiator is a universal negotiation environment where an LLM agent learns WHO it’s talking to using Bayesian belief updates — and shifts strategy accordingly. Built on OpenEnv, it supports multiple domains: guitar sales, car sales, property, and an HR hiring mode where the LLM plays recruiter against three candidate personalities (GrowthSeeker, StabilitySeeker, MarketExplorer). Every round, a belief tracker updates probabilities based on price movement and language signals, injecting a strategy hint directly into the observation. The agent is trained with GRPO using Qwen2.5-1.5B-Instruct, LoRA r=16, num_generations=2 (2-GRPO), and beta=0.0 (DAPO). Reward is composite: 70% price outcome + 20% empathy + 10% efficiency.

- **Team:** [Murugs Karmegam](https://cerebralvalley.ai/u/Muruga)
- **GitHub:** https://github.com/murugamayuk/negotiation-arena
- **Demo video:** https://youtu.be/agAIMZD7BoI
- **Hugging Face:** https://huggingface.co/spaces/murugamayuk/negotiation-arena
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=19

### 103. Robot

RoboReplan is a tabletop planning environment for OpenEnv that targets Problem Statement 3.1 (World Modeling: Professional Tasks). LLMs often fail at long-horizon robot tasks because they can't replan — they loop or freeze when something goes wrong (blockers, grasp slips, mid-task instruction changes). RoboReplan benchmarks this failure mode and trains agents to recover: clear blockers, pick and place in the right bins, and adapt when the instruction changes. The env is deliberately hard for a small model so we can show improvement over time (0% → 78% success with SFT + GRPO). Four domain skins (Default, Pharmacy, Warehouse, Lab) ground the same mechanics in professional scenarios. Built with OpenEnv 0.2.1, deployed on HF Spaces, trained with HF TRL (GRPO) in Colab

- **Team:** [Jwalin Shah](https://cerebralvalley.ai/u/jwalinshah)
- **GitHub:** https://github.com/jwalin-shah/robo-replan
- **Website:** https://colab.research.google.com/github/jwalin-shah/robo-replan/blob/main/train/colab_train.ipynb
- **Demo video:** https://youtu.be/1waYHPytoNE
- **Hugging Face:** https://huggingface.co/spaces/jshah13/roboreplan
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=20

### 104. Athens

This is an environment to do RL training and other forms where an LLM can play the game teeworlds https://teeworlds.com/, 4 bots shown in video

- **Team:** [Ziad A](https://cerebralvalley.ai/u/ziad)
- **GitHub:** https://github.com/ziadgit/teeunit https://huggingface.co/ziadbc/teeunit-agent https://huggingface.co/spaces/ziadbc/teeunit-env
- **Website:** https://github.com/ziadgit/teeunit/blob/main/notebooks/teeunit_training.ipynb
- **Demo video:** https://www.loom.com/share/058382f01c5d4e9b91cce5091871a8e3
- **Hugging Face:** https://huggingface.co/spaces/ziadbc/teeunit-env
- **Project:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery?project=21

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
