Skip to Main Content

Replay

Built at The Agent Arena Hackathon · Sep 26, 2026 · San Francisco, CA

Demo video · replay-so101.vercel.app/…

ChatGPT learned by reading the internet. Robots can't: every robot skill today is recorded by hand, one task at a time, and the new robot models (VLAs) are starving for that data. Replay turns the open web into robot skills from one sentence. How it works: an agent turns the prompt into a search plan. Scrapling scrapes inside a throwaway Vultr VM per robot (Creative Commons YouTube video and web pages, licences checked), a vision model checks each clip, and MediaPipe tracks the person's hand and body. The motion is retargeted onto the SO-101's joints and replayed on hundreds of new table layouts in MuJoCo, keeping only the tries that work (the same idea as NVIDIA's MimicGen). A small policy then learns by imitation with ACT-style action chunking; for rule-based tasks the moves are planned instead, and every move is rehearsed in physics. Nothing reaches the robot until it has worked in simulation. Results (simulated SO-101 in MuJoCo): block in the bowl 19/20 unseen layouts (policy trained in 21 s on a CPU), block tower 18/20, a dumbbell curl copied 1:1 from Creative Commons video (elbow within 1.25 degrees), Tower of Hanoi in the optimal 7 moves, a six-cup pyramid, and Deep Blue vs. Kasparov 1997 Game 6 played on both sides (37 moves, worst placement 4.9 mm). Every skill can be downloaded. Next: a real SO-101 hardware test and fine-tuning SmolVLA on Replay's data.

Team