# GPU Sitter

- **Event:** [AI Engineer World's Fair Hackathon 2026](https://cerebralvalley.ai/e/aiewf-hackathon-2026)
- **When:** Jun 27 at 9:00 AM – Jun 28 at 5:00 PM (PDT)
- **Where:** San Francisco, CA
- **Team:** [Jeff Miller](https://cerebralvalley.ai/u/Look), [Mohid Butt](https://cerebralvalley.ai/u/mohid), [Octavian Cosmin](https://cerebralvalley.ai/u/octavianc), [John Coleman](https://cerebralvalley.ai/u/johncoleman)
- **GitHub:** https://github.com/jcmiller/hackathon-datacenter-agent
- **Demo video:** https://youtu.be/BhjBohJsYrY
- **Gallery:** https://cerebralvalley.ai/e/aiewf-hackathon-2026/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/aiewf-hackathon-2026/hackathon/gallery/31

GPUSitter is an on call engineer that responds to failures in GPU datacenter clusters. \
We built tools for the agent to inspect NVIDIA sensor data, and power, temperature, and utilization data. 
From that it reasons about the root issue and informs about resolution.
Self improvement: After its finished reasoning, with its expert knowledge, it updates its own small ML prediction model for failures. If it scores better than the existing ML model replaces it.
PS: The we data (schema) used it real LLM training cluster logs from the AcmeTrace paper. So you can actually use our setup for GPU operations :)

---

Markdown version of https://cerebralvalley.ai/e/aiewf-hackathon-2026/hackathon/gallery/31. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
