# Hypernoa

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Team:** [Appalanaidu Bobbili](https://cerebralvalley.ai/u/Naidu)
- **GitHub:** https://github.com/naidu1212/hypernoa-astrum
- **Website:** https://colab.research.google.com/github/naidu1212/hypernoa-astrum/blob/master/colab/astrum_grpo_training.ipynb
- **Demo video:** https://www.youtube.com/watch?v=K7VSVrXSAr0
- **Hugging Face:** https://huggingface.co/spaces/ABNaidu/hypernoa-astrum
- **Gallery:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/7

Hypernoa Astrum is an adaptive RL environment built on OpenEnv 0.2.1 that trains AI to reason, adapt, and align — not just solve tasks. It simulates 5 competing stakeholder groups (Workers, Management, Regulators, Customers, AI Systems), 3 dynamic phases (stable, value shift, crisis), and deliberately designed alignment traps that test whether AI agents cheat their reward function. The multi-objective reward measures effectiveness, fairness, alignment, and adaptability. Our trained agent reaches 24.7 reward starting from scratch, nearly matching the expert baseline at 25.1, while random scores only 14.6. All 3 alignment traps are resisted. Trained with HF TRL GRPO on Qwen2.5-0.5B using CoreWeave H100 via Northflank. Addresses Problem Statement 3.1 (World Modeling) and Statement 5 (Wild Card).

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/7. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
