# expertoncall

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Team:** [karthik ganesan](https://cerebralvalley.ai/u/karthik1996), [Sanat Mouli](https://cerebralvalley.ai/u/sanatmouli), [Mert Hidayetoglu](https://cerebralvalley.ai/u/merth)
- **GitHub:** https://github.com/sfc-gh-mhidayetoglu/OpenEnv/blob/add-agent-world-model/envs/agent_world_model_env/EXPERT_ENHANCEMENT.md
- **Website:** https://github.com/sfc-gh-mhidayetoglu/OpenEnv/blob/add-agent-world-model/envs/agent_world_model_env/train_grpo_awm.ipynb
- **Demo video:** https://youtu.be/j3oM5jQnqf0
- **Hugging Face:** https://huggingface.co/spaces/karthik/awm-dynamic-expert
- **Gallery:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/97

**Statement 1: Multi-Agent Interactions**
Dynamic Expert-in-the-Loop GRPO Training on Agent World Model
What Is This?
Imagine teaching a new employee. You wouldn't just hand them a manual and walk away. You also wouldn't stand behind them dictating every keystroke.

The best approach? Let them try, and tell them an expert is available if they get stuck.

We give a small language model (Qwen3-4B) a set of ~35 API tools, a task description, and access to a brilliant advisor (GPT-5.1). Then we use reinforcement learning (GRPO) to teach it when calling the expert leads to better outcomes — and ultimately, when it can fly solo.

Detailed description is in this readme: https://github.com/sfc-gh-mhidayetoglu/OpenEnv/blob/add-agent-world-model/envs/agent_world_model_env/EXPERT_ENHANCEMENT.md

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/97. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
