# mgsa

- **Event:** [Built with Opus 4.7: a Claude Code hackathon](https://cerebralvalley.ai/e/built-with-4-7-hackathon)
- **When:** Apr 21 at 12:00 PM – Apr 27 at 2:00 AM (EDT)
- **Where:** Online
- **Team:** [matthieu gsa](https://cerebralvalley.ai/u/mgsa1)
- **GitHub:** https://github.com/mgsa1/PolicyGrader/tree/main
- **Demo video:** https://www.youtube.com/watch?v=uYHvZM8eQWY
- **Gallery:** https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery/31

PolicyGrader is an agentic system that helps make robots safer to operate alongside humans. It automatically stress-test and analyse robot control policies, end-to-end. It compresses a process that costs robotics teams weeks of frame-by-frame video review into minutes. You describe an evaluation goal in plain English; then a swarm of specialized Claude Opus 4.7 Managed Agents — a planner, a rollout worker, K-parallel vision judges, and a reporter — design the test suite, execute it in simulation, point at each failure with pixel accuracy (or honestly abstain), and cluster the deployment findings into actionable patterns using Opus 4.7's 1M-context window.
The differentiator is trust: every run includes a calibration cohort whose ground truth comes from a human labeling a sampled subset, and the judge's measured per-label precision is attached as a confidence chip to every deployment finding. No vibes with safety. The judge is auditable from the dashboard's runtime.json and findings.jsonl. Submitted to the Anthropic Opus 4.7 Hackathon.

---

Markdown version of https://cerebralvalley.ai/e/built-with-4-7-hackathon/hackathon/gallery/31. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
