Skip to Main Content

44872

Built at The Harness Engineering & Model Wrangling Hackathon · Sep 26, 2026 · New York, NY

Demo video · streamable.com/…

Harness Architect: an autonomous infrastructure engineer for GPU inference Owning GPUs shouldn't require becoming an inference expert. Describe your hardware and workload, and Harness plans an open-source serving stack, then keeps improving it as traffic changes, proving every change first. Plan: sizes the model to your GPUs, picks from six open-source projects (KServe, llm-d, vLLM, SGLang, Dynamo, Envoy AI Gateway), and outputs schema-validated YAML plus an install.sh. Diagnose: turns incident metrics into one config fix, then confirms it worked. Evolve: a long-running campaign tests each change against the current setup on identical replayed traffic; deterministic gates, not the model, pick the winner. In a long-prompt surge, p95 was 20.9s, doubling the GPUs gave 9.7s, and the agent's fix gave 2.2s. MongoDB Atlas: change streams hot-reload each promoted architecture, time-series collections record every request, and Vector Search recalls what worked before. Checkpoints and a lease make runs crash-safe, and context manifests make every decision auditable. Stack: Python, FastAPI, MongoDB Atlas, Strands Agents, OpenRouter, Kubernetes, KServe, llm-d, NVIDIA Dynamo, vLLM, SGLang, Envoy AI Gateway.

Team