# Play-gent

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Placement:** Finalist
- **Team:** [Abraham Bhatti](https://cerebralvalley.ai/u/abhatti)
- **GitHub:** https://github.com/AbeBhatti/Play-gent
- **Website:** https://colab.research.google.com/github/AbeBhatti/Play-gent/blob/main/training/arbitragent_colab.ipynb#scrollTo=FmGOks6oquiy
- **Demo video:** https://youtu.be/bsKZgyDsuDc
- **Hugging Face:** https://huggingface.co/spaces/Abeee32t/ArbitrAgent?logs=container
- **Gallery:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/27

Our agent uses OpenEnv to build a curriculum of video game environments that train a TinyLlama 1.1B agent via GRPO reinforcement learning. We implement three OpenEnv-compliant environments — Diplomacy (coalition tactics), webDiplomacy human gameplay (211k real states), and IRC poker (bluff detection) — each with dense reward signals that teach the core primitives of strategic negotiation. The agent learns through RL across this curriculum: Phase 1 trains coalition pressure in Diplomacy, Phase 2 grounds it in human gameplay patterns, Phase 3 unifies all three reward signals in a live arbitrage environment. The result is an agent that transfers video game negotiation skills to real economic interactions — starting with $20, it compounds capital to $80 (4x return) against Groq Llama 3.1 8B adversarial sellers, detecting bluffs at 97% confidence and reaching actual price floors 100% of the time.

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/27. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
