Skip to Main Content
2nd Place

Nusa

Built at The Agent Arena Hackathon · Sep 26, 2026 · San Francisco, CA

Demo video · www.loom.com/…

AI models on CPUs are held back by generic kernels. Nusa is an agent that writes faster ones for your exact hardware, and it's safe to leave running on its own. Pick a model from Hugging Face (for example AMD's own int4 models), an engine (vLLM + AMD ZenDNN, or llama.cpp) and a Vultr VX1 (AMD EPYC Turin), then say what you want. The agent runs on Vultr Serverless Inference. It measures the baseline, finds the slowest kernel, and writes a new one in C. Every attempt compiles and runs in its own throwaway Microsandbox microVM: no network, no keys, unprivileged, with memory and time caps. A judge with a hidden reference rejects wrong answers. Faster, correct kernels are kept as git commits. Measured end to end, stock vs agent, served in microVMs: - AMD Llama-3.1-8B w4a16 (vLLM + ZenDNN): 5.79 → 11.52 tok/s (1.99×) - Qwen3-4B w4a16: 10.03 → 19.20 tok/s (1.91×) - AMD Phi-4-mini w4a16: 10.06 → 17.42 tok/s (1.73×) - Bonsai 8B / 27B 1-bit (llama.cpp): 1.70× / 1.61× The agent rewrote the int4 matrix kernel that AMD's ZenDNN stack calls (built from source: vLLM, zentorch, ZenDNN, AOCL-DLP): it unpacks 4-bit weights in registers and uses AVX-512 bf16 dot products. Containment first: a red-team mode fires rm -rf /, a fork bomb, an infinite loop, a memory bomb, API-key theft and phone-home through the same microVM path. All six are contained, and the host checks this from outside the sandbox. The app is served through NetBird with no open ports and a password gate, and each run gets a public link that expires when the run ends.

Team