AgentLab
Built at Google I/O Hackathon · May 23, 2026 · San Francisco, CA
When a multi-agent pipeline produces a confident wrong answer — which step broke, and how do you know? AgentLab is an evaluation system for multi-step agent pipelines. It watches production traces, detects quality drift, and does the part nobody else automates: it localizes the regression to the specific step that caused it. From the real failing traces it drafts targeted evaluation cases; a human approves them; the fleet is re-scored against the enlarged suite; the degraded step is rerouted to a healthy agent