Skip to Main Content

Digisafe

Built at Google DeepMind Bangalore Hackathon · Jul 11, 2026 · Marathahalli, Marathahalli Main Road

Digisafe — Demo video

DigiSafe: Your On-Device Acoustic Bodyguard The Problem Every year, billions of dollars are lost to sophisticated AI phone scams. Bad actors use deepfakes to impersonate loved ones, or pretend to be authorities like the CBI or tax agencies claiming you are under "digital arrest." Current scam-blocking apps rely on checking known spam phone numbers, which fails against spoofed numbers or deepfakes. Alternatively, they require sending your audio to the cloud for analysis, which completely compromises your privacy. We shouldn't have to sacrifice privacy for security. What is DigiSafe? DigiSafe is an always-on, local-first acoustic defense system. When you engage Defense Mode, DigiSafe continuously listens to your phone calls or ambient environment. It instantly detects deepfakes, phishing attempts, and digital arrests without a single byte of your audio ever leaving your device. Everything—from transcribing speech to recognizing voices and detecting threats—happens 100% locally on your machine. How It Works: The Multi-Agent Pipeline Under the hood, DigiSafe replaces traditional cloud APIs with a parallel Multi-Agent Pipeline. When audio is captured, it is fanned out simultaneously to four specialized AI agents running on-device: 1. The STT Agent (Whisper): Transcribes the audio locally to understand what is being said. 2. The Identity Agent (Librosa): Extracts industry-standard 120-dimensional MFCCs (Mel-Frequency Cepstral Coefficients), Delta, and Delta-Delta features. This creates a highly accurate, mathematical fingerprint of the speaker to verify if the caller is actually your enrolled family member, or a deepfake. 3. The Acoustic Stress Agent: Analyzes the raw audio waveform for RMS Energy and Zero-Crossing Rate to detect aggressive vocal strain and tension, which are massive red flags for coercion. 4. The Threat Classifier Agent (Gemma 4): The star of the pipeline. Powered by the incredibly fast Gemma 4 model running locally on the device, this agent semantically evaluates the transcript for phishing, generic scams, or digital arrest patterns. Because we use asynchronous processing, all four agents analyze the audio at the exact same time. The results are instantly merged by our Orchestrator, allowing DigiSafe to make a lightning-fast, highly accurate threat assessment. Key Features 1. 100% Local-First: Privacy is paramount. We use local LLMs (Gemma 4) and local transcription (Whisper). No internet connection required. 2. Deepfake Resistance: By analyzing 120-dim acoustic features, DigiSafe doesn't just listen to what is said, it mathematically verifies who is saying it. 3. Semantic Threat Detection: Gemma 4 acts as a private detective, instantly understanding the context of the conversation and flagging coercion techniques like "your account is frozen" or "do not hang up." 4. Comprehensive Audit Log: Every intercepted threat is saved in an on-device audit log, complete with the audio recording, transcript, and a detailed Agent Breakdown explaining exactly why it was flagged as a scam. Best Use of Gemma 4 - Local-First Agents DigiSafe was built aggressively for this special prize category. By leveraging Gemma 4, we were able to bring cutting-edge semantic threat detection directly to the edge. Gemma 4 provides the perfect balance of small footprint and high intelligence, allowing our local Threat Agent to understand complex scam context—such as digital arrests and phishing attempts—without the extreme latency or privacy violations of a cloud API. DigiSafe proves that with Gemma, local-first AI can protect our most vulnerable populations.

Team