Airlock
Built at The Agent Arena Hackathon · Sep 26, 2026 · San Francisco, CA

Letting an AI agent run code from a random bug report is a security problem. The report is untrusted, the code the agent writes is untrusted, and so is the agent's claim that it fixed anything. Airlock handles that. You paste a real issue for a supported library. The agent reproduces the bug, writes a fix and hands back a patch you can download. All of it runs in throwaway Kata sandboxes on a Vultr VX1 host, with no network, no secrets and hard limits on CPU, memory and time. The agent can edit its own fix. It can't touch the tests that judge it, and it can't decide that it passed. A separate comparator reruns the frozen checks against a sealed copy of the patch and gives the verdict. A fake "all tests passed" log from the agent gets a fail. It also runs general tasks: web research in a sandboxed Chromium behind an allowlisted proxy, offline data analysis, and form submissions that need a human to approve them first. The model is glm-5.3 on Vultr Serverless Inference. We ran it on the real python-tabulate #365 bug and it fixed it in 3 of 3 fresh live runs. Paste `rm -rf /` into the hostile input panel and it kills one sandbox. The control plane, the other tasks and the host keep running.