CyberDragon
Built at OpenEnv Hackathon SF · Mar 7, 2026 · San Francisco, CA

WatchDog is an RL environment training AI oversight agents to detect errors in real-time. It solves the gap where humans miss 60% of subtle AI hallucinations and logic flaws. This gym enables agents to make sequential decisions like passing or flagging turn-by-turn. The project uses the GRPO algorithm and LoRA adapters to remain fast and efficient. An adversarial arms race trains a detector and mutator via a minimax game for robustness. The system allows adding new domains like medical triage in just ten lines of code. Its reward structure penalizes false positives heavily to prevent hallucinated bugs. Training achieved a 15.4× accuracy jump and 5.7× recall increase in eighty minutes. The model shifted from indecision to making clear judgments on 77% of dialogue turns. Deployed on Hugging Face, WatchDog provides a production-grade gym for AI oversight.