Hypernoa
Built at OpenEnv Hackathon SF · Mar 7, 2026 · San Francisco, CA

Hypernoa Astrum is an adaptive RL environment built on OpenEnv 0.2.1 that trains AI to reason, adapt, and align — not just solve tasks. It simulates 5 competing stakeholder groups (Workers, Management, Regulators, Customers, AI Systems), 3 dynamic phases (stable, value shift, crisis), and deliberately designed alignment traps that test whether AI agents cheat their reward function. The multi-objective reward measures effectiveness, fairness, alignment, and adaptability. Our trained agent reaches 24.7 reward starting from scratch, nearly matching the expert baseline at 25.1, while random scores only 14.6. All 3 alignment traps are resisted. Trained with HF TRL GRPO on Qwen2.5-0.5B using CoreWeave H100 via Northflank. Addresses Problem Statement 3.1 (World Modeling) and Statement 5 (Wild Card).