Ghost Walk
Built at RAISE Summit Hackathon · Jul 4, 2026 · Paris, France

Inspiration Site engineers in bandwidth-denied environments — oil fields, mines, remote farms — walk inspection rounds every day, and what they noticed yesterday evaporates. The rattle heard Monday and the tilt photographed Tuesday never meet. Cloud AI can't help where there is no cloud. We were inspired by research on memory consolidation in LLMs ("LLM sleep"): what if the agent didn't try to think in real time at all, but slept on it — like a brain? What it does GHOST-WALK turns a browser tab into a self-contained reasoning agent with a day/night cycle: The Walk — the engineer photographs fixed checkpoints (or imports a photo/video frame) and dictates voice notes, transcribed on-device by Whisper-tiny. Each observation is logged locally. The Sleep — reasoning on every photo with a 2B-parameter model would drain a field device mid-round, so GHOST-WALK defers all heavy inference to the night dock: wall power, idle GPU. Overnight, Gemma 4 E2B captions each photo, recalls that checkpoint's baseline from the previous pass, scores structural drift on an anchored 0–10 rubric, and correlates the worker's spoken symptoms with visual change. Each result commits atomically and today's observation becomes tomorrow's baseline — memory consolidation, literally. The Briefing — by morning: a triaged worklist. Critical drift (≥7) surfaces as one imperative action: "Check the motor mount bolts on Pump A — the rattle you reported correlates with visible housing tilt." Everything — images, voice, model weights, reasoning — stays on the device. After one online initialization, the entire loop runs in airplane mode. Domain profiles make the agent generic: a ~20-line config swaps checkpoints, persona, and vision vocabulary — industrial sites and farm/agronomy ship today. How we built it 100% client-side, zero backend: Next.js 16 static export · Gemma 4 E2B (onnx-community/gemma-4-E2B-it-ONNX, q4f16) on WebGPU via transformers.js v4, with WASM fallback · Whisper-tiny for offline speech-to-text (the Web Speech API is cloud-backed and fails offline) · all inference in a Web Worker with per-request timeouts · Dexie.js/IndexedDB for logs, image blobs, and per-checkpoint memory · PWA service worker for the app shell; transformers.js caches weights in the Cache API. Reliability engineering for a 2B model on the edge: a JSON-hardening ladder (tolerant parse → one retry → keyword heuristic, honestly tagged HEURISTIC in the UI), per-log atomic commits so an aborted night pass loses nothing, and a deterministic mock brain behind a flag so the demo loop was testable end-to-end from hour 8. Challenges we ran into Gemma 4's browser API had sharp edges: the processor's (text, images, audio, options) signature silently ate our options object as audio input, and TextStreamer assumes stdout exists. Both found by running the real pipeline early. Real-photo calibration taught us the model is honest: photograph two different scenes and it correctly refuses to correlate a reported rattle with an unrelated visual change. We redesigned the test protocol (same subject, same framing, one visible change) and added an anchored scoring rubric so worker-reported symptoms floor the drift score. Offline speech: Chrome's SpeechRecognition streams to Google's servers — useless in airplane mode. Whisper-tiny on WebGPU replaced it. Accomplishments we're proud of Real Gemma 4 vision + reasoning on WebGPU, verified end-to-end: ~30 seconds per consolidated observation on a consumer laptop. A full walk→sleep→briefing loop that survives airplane mode, browser restarts, and mid-batch aborts. Built solo, entirely during the event — the repo's gate-tagged commit history (gate-1 → gate-4-core) documents every milestone. What we learned Edge AI isn't a smaller cloud — it's a different design language. Deferring inference to charging time, flooring scores on human testimony, and refusing to infer equipment identity from vision (it's declared via checkpoint selection — deterministic beats probabilistic in safety workflows) all came from taking the constraint seriously instead of fighting it. What's next Asset-tag scanning (QR/NFC) for one-gesture checkpoint identity · grounding drift scores in equipment documentation stored offline ("tilt exceeds the manufacturer's 0.5° alignment tolerance") · multi-day trend memory · Gemma 4's native audio input to replace Whisper — one model for everything · thermal-camera frames through the same pipeline. Built with gemma-4-e2b · transformers.js · webgpu · onnx · next.js · typescript · dexie · indexeddb · tailwindcss · whisper · pwa · web-workers · vitest