croissants
Built at RAISE Summit Hackathon · Jul 4, 2026 · Paris, France

- A FastAPI runner that queues a run to Postgres, claims it, and drives the ADK pipeline one step at a time — invoking the Runner per step with a state_delta carrying current_step and each dependency's status — while translating every ADK event into a raw run_logs trace. - A conditional workflow graph (TestRunner → FixWriter → Verifier → PRAgent) where the needs_fix edge is emitted by an after_agent_callback only on a medium/high-confidence failure; passing steps and low-confidence failures terminate early. - The feature-gate, deterministic evidence-escalation trigger, and confidence-scoring callbacks — the gating and observability layer that keeps dependency and diagnosis logic in code, out of the LLM's discretion. - A shared-Chromium design: one Playwright-launched Chromium at startup that both Computer Use (`connect_over_cdp`) and chrome-devtools-mcp (`--browserUrl`) attach to — never launched or torn down per tool call. - The Python ↔ web DB seam: Python reads/writes Prisma-owned tables via raw asyncpg and never migrates; the Next.js app owns schema and migrations. ## Challenges - Two agent surfaces that cannot be combined. Computer Use and the Antigravity agent are separate surfaces that can't share one call, so state is handed off explicitly (evidence bundle out of TestRunner, environment_id into FixWriter) rather than passed as tools. - Reliable escalation on stage. Trusting the LLM to decide "should I check the console now?" was too flaky for a live demo, so escalation was made a deterministic callback on Computer Use's own result instead of a prompt instruction. - Browser sharing. Computer Use and chrome-devtools-mcp both need the *same* live page; solved with a single keep-alive Chromium on CDP :9222 that both attach to. ## Results QA Sentinel runs live, end to end against the bundled cambria-shop demo app, which carries a deliberate bug: checking out ceramic-mug hits a missing shipping-cost key and raises an unhandled KeyError → bare HTTP 500. On a run, the pipeline: 1. drives the checkout via Computer Use and observes the failure, 2. escalates to chrome-devtools-mcp *because* that step failed, capturing the 500 and the console/network trace, 3. scores the evidence, writes an evidence-grounded fix, 4. re-verifies the error is gone, and 5. opens a GitHub PR whose body is the full evidence bundle. What it demonstrates: an agent that reasons from root cause, not symptom, knows what it *can't* diagnose (and defers those to a human), and produces a **verified, reviewable deliverable** — the PR + evidence trail — rather than a confident guess.