Skip to Main Content
Events
SearchSign In

Google Gemini Hackathon

Oct 18, 2025 · San Francisco, CA

Event Page
Event Page

Story Time

Draw a character, write a prompt, select narrator, theme, and art style. Then it's Story Time! This project was inspired by how my niece would love watching Coco Melon or Baby Fin. I figured having a way for them to enjoy drawing as well as getting the parents involved in a fun way to have actual engagement and parent bonding.

Story Time project preview
0

#SoulBits

Upbeat – AI Music Tutor for Everyone Upbeat - Accessible Music Education For All An AI-powered music learning platform that makes music education accessible to everyone, anywhere, anytime with a camera and internet. Built on Gemini AI with real-time vision and audio analysis, it watches your hands, listens to your playing, and gives instant feedback—like a personal teacher available 24/7. It solves the high cost ($50–100/hr), limited access, rigid schedules, and lack of feedback in traditional lessons. Learners get adaptive instruction based on skill level, visual/audio corrections, and engaging avatars inspired by pop culture. Practice anytime, anywhere—no geography, income, or time constraints. Upbeat democratizes music education for rural communities, low-income families, and beginners worldwide. It scales infinitely, works in any language, and proves how AI can enhance—not replace—creativity and learning.

drive.google.com/…
0

MODO

AI Automation workflow powered by visual and audio Will put demo vid in github readme, youtube not letting me upload

github.com/…
0

DrystAI

Never forget a face or conversation again. DrystAI ("Dryst" means vision in Sanskrit), turns your smart glasses into a real-world memory. It recognizes faces, pulls up past conversations, and shows you names, companies, and context right when you need it. No awkward pauses. No “sorry, what was your name again.” It’s LinkedIn, your notes, and your brain working together in real life. Built for people who actually want to remember.

DrystAI project preview
0

Sports Clips

Our platform captures the best moments from sports games instantly, delivering them in a TikTok-style iOS feed. Watch foreign games with automatic AI commentary in your language. Built with Swift (iOS), Kotlin/Ktor (middleware handling profiles, likes, comments via MongoDB), and Python (Gemini Flash API for highlight detection). Features vector-based recommendations using Voyager AI embeddings and Cloudflare R2 storage for clips.

Sports Clips project preview
0
+1

Syntra

We built the first Cursor for Designers, a platform that enables designers to autonomously launch multimodal agents through Google's cloud resources. This saves designers like our team member, Carmah, an estimated 7.23 hours each week because it has the ability to autonomously make decisions on what design tasks can be automated instead of the designer performing repetitive, monotonous tasks. Many companies have integrated with Figma's MCP for context retrieval, but until now, companies have been unable to directly edit Figma files through multimodal agents, something that's exceptionally valuable to top designers.

www.veed.io/…
0
+1

Ultron

AI security system that secures your property, challenges suspects, and terminates intruders with high tech lasers.

Ultron project preview
0

Mother Lover Hawk Tuah

Create comic books and visual maps based on your repo.

www.loom.com/…
0

FindMe

FindMe - AI Vision for Precise Location Problem: Google Maps says "500 feet away" - useless in crowds of 10,000. GPS fails when precision matters: lost children, friends at concerts, conference coordination. Solution: FindMe adds precision to Google Maps. Gemini 2.5 Pro transforms vague GPS into AR guidance. Point camera, AI detects target, shows arrows + distance: "Turn left 15°, walk 23m." Impact: Search 15min→2min. Accuracy 500ft→15ft. Works where GPS fails. Applications: Find friends at festivals, locate lost children, emergency response, crowd management, security surveillance via natural language. Tech: 100% Google - Gemini 2.5 Pro + Maps API. <2sec processing, multi-modal, scalable. Why It Wins: We complete Google Maps by adding computer vision. "Nearby" becomes "exactly here." FindMe: Because "nearby" isn't close enough.

www.loom.com/…
0

13th floor Fireguard.ai

Traditional fire alarm systems in apartment buildings are archaic. They are loud, impersonal, and often terrifying, causing panic rather than providing clear, actionable guidance. 13th Floor is an application that redefines the resident experience during a fire emergency. It moves beyond the simple alarm and provides a comprehensive, intelligent insight, and interactive safety net for every resident. Residents will be alerted of a fire and the floor its taking place on, then be given instructions on the best route to evacuate the building safely. If stuck on a floor above the fire, residents can send a distress signal which will be sent to fire fighters so that they get an aggregate of how many residents are still within the apartment building.

13th floor Fireguard.ai project preview
0

Nathan Burg

Protocol Copilot transforms written lab procedures into multimodal, interactive assistants that guide researchers step-by-step through experiments using real-time speech and vision. It solves the reproducibility crisis in experimental science by standardizing execution, capturing deviations, and linking every run to structured data for transparent, verifiable research.

Nathan Burg project preview
0

Liber-T

Truthy - Real-Time AI Fact-Checker for Live Debates THE PROBLEM: Misinformation spreads faster than fact-checkers can verify it. Traditional fact-checking happens hours or days after debates, when false claims have already influenced millions. Live debates lack real-time accountability, allowing unchallenged misinformation to shape public opinion. THE SOLUTION: Truthy uses Gemini's Realtime API to fact-check spoken claims in under 5 seconds during live debates. It: • Transcribes two speakers simultaneously using Gemini Live sessions • Detects check-worthy factual claims (numbers, dates, comparisons) • Searches Google for authoritative evidence (.gov, .edu, official sources) • Displays color-coded verdicts (True → False → Unverifiable) with confidence scores • Shows clickable sources with publication dates for transparency

www.loom.com/…
0

5-Dee

Skillshare: Marketplace to Share & Monetize Skills & Agents

ai.studio/…
0

Vulcan Eyes

Problem: Traditional retail surveillance lacks real-time intelligence and actionable insights. Customers wandering without assistance, inefficient queue management, no behavioral analytics. Manual review is time-consuming and doesn't help in real-time. Helping sales reps, managers, owners better service and understand their customers using AI

drive.google.com/…
0

ShiftLeft

Shift-Left Command Centre is an AI-powered platform that automatically detects security vulnerabilities and compliance issues in developer workflows before code reaches production. 🛡️ It solves the problem of traditional security checks happening too late by using Google Gemini and Fetch.ai agents to analyse code commits and screenshots in real-time, providing immediate feedback and automated alerts (Jira, Slack, GitHub) to catch risks early. ✨

drive.google.com/…
0

Slate.ai

Education is the foundation of growth, yet as we age, learning becomes impersonal. We move from guided learning on a slate with parents or teachers to self-learning through books and screens. Many students hesitate to ask questions due to fear or self-doubt, losing the confidence to clarify concepts. The lack of personalized support means when they struggle, no one is there to guide them hand-in-hand — leading to gaps in understanding and limited academic progress. Our Solution: Slate.ai is an AI-powered assisted learning platform that revives the childhood slate experience — where learning felt personal and guided. Acting as a virtual teacher, it provides real-time help, feedback, and motivation in an interactive, canvas-like environment. By blending nostalgia with cutting-edge AI, Slate.ai makes learning intuitive, engaging, and fear-free — helping every learner grow confidently at their own pace.

Slate.ai project preview
0

ScreenMate

Wise assistant that watches your screen and suggests smart tasks for email, calendar invites with a built in chat and voice mode.

www.loom.com/…
0

Agentic Citizens

An AI powered civic assistant which helps humans report infrastructurre issues in their cities easily with Gemini's multimodal AI and computer use. Users can report evidence of issues like potholes, broken sidewalks, trash, and Gemini helps them file a report on the app and the SF gov website. Further, Gemini can auto-create volunteering clean-up like events based on the leading posts with most votes.

www.loom.com/…
0

UI Context

Context augmentation for your AI agents via copy and paste. Click on elements via Chrome Extension to either extract text or component screenshot to your context library. Access your context library on the web where you can copy the context to your clipboard to paste it into your AI agent.

UI Context project preview
0

Desktop Agent

Desktop Agent

Desktop Agent project preview
0

Guardian Agent

Guardian Agent is a personal safety system which uses Gemini 2.5 Flash to process audio transcripts and determine live if a situation someone is in is unsafe and if so creates a phone call to get them out of an uncomfortable situation, recording information about the situation to be used to file a report if needed after.

Guardian Agent project preview
0

CodeWise

CodeWise 🎯 AI-Powered Algorithm Leetcode Practice with Adaptive Learning - Master coding interview patterns through intelligent problem generation and personalized feedback. Select the pattern you would like to practice and CodeWise powered by Gemini will generate a problem for that pattern and chosen difficulty. There is also an added feature to upload a picture of your notes or screen if you are stuck and Code Wise will give you few hints without revealing too much Since problems are generated by gemini there are endless problems to solve and practice it also tracks stats and which patterns you are strong or weak at. 🌟 Features 🤖 AI-Powered Learning Google Gemini AI integration for dynamic problem generation Intelligent Code Evaluation with detailed feedback and scoring Adaptive Learning that tracks weak patterns and provides personalized recommendations Image Hint System - Upload handwritten notes for AI analysis and hints

CodeWise project preview
0

No Ticket

No Ticket is the app that helps you find free parking in SF and keep your sanity (and wallet) intact.

www.loom.com/…
0

Travel Dreaming

Project name: Travel dreamer Description: Imagine dropping a pin on Google Maps and instantly seeing a video of yourself walking around that spot. Travel planning can be exhausting. Mapping out attractions on Maps still makes it hard to feel what being there is like. Travel Dreamer solves this by generating realistic videos of you or your group in any Street View location. Upload a picture of yourself, pin a spot on Google Maps, and the app uses Nano Banana to merge you with the scene into a single image, then Veo 3 turns it into a high-res video. As compute gets faster, users could select start and end points to see a continuous video of themselves walking through entire routes. In the future, Travel Dreamer could even integrate Google Photos with Maps, letting users view dreams beside memories.

Travel Dreaming project preview
0

Watch & Learn

Watch & Learn transforms screen recordings into executable "skills" for AI agents. Instead of writing tedious documentation, simply record yourself performing a task with narration. W&L automatically extracts step-by-step instructions, automation scripts, templates, and reference assets—converting tacit knowledge into structured formats AI agents can follow. Deploy skills as downloadable ZIP files, MCP servers, or run via Gemini computer-use models in-app. The core insight: "showing" captures nuance that written instructions miss, making on-the-job agentic training faster and more reliable.

www.loom.com/…
0

Pixel Throw

Pixel Throw is an interactive app that transforms real-world objects into playable, pixel-art projectiles in a physics-based throwing game. Using the power of generative AI, the app seamlessly blends computer vision, creative content generation, and game mechanics into a single, engaging experience.

Pixel Throw project preview
0

Nuvue (new+view: seeing in a new way)

Did you enjoy the foods prepared by Cerebral Valley? How did you choose between turkey, tofu and ham burritos? What if you had a slight touch with menu decision making to help you in real-time for your wellness life? Nuvue is a real-time food decision solver, using multimodal generative AI to solve everyday life problem, approximately 221 food decision makings per day. We redefined and solved food decision problem from nutrition history record analysis to helping decision making on demand from diverse eating situation (types of foods). When the user uploads or records video by phone and asks for food decision making advice, the service will help you choose. This project can expand to using google glasses and to various users including children education and patients as a personalized problem solver. *Nuvue means new view, seeing food and health habit in a new way.

Nuvue (new+view: seeing in a new way) project preview
0

Regarded Artist

Google God's Eye: AI Flythrough of terrain to help drivers, cyclists, etc preview terrain. Extremely accurate waypoint integration to preview waypoints from every angle.

Regarded Artist project preview
0

Tripverse

Travel planning is time consuming, people are indecisive, and have to consistently switch between services to create the perfect plan. Tripverse uses a swarm of agents to help a user plan trips and also answers any questions they have concerning them regarding the trip, as the conversation is saved. And if the user wants, it can even out a full itinerary for the user upon their request!

www.loom.com/…
0

Visper

Visper is an AI-powered storytelling pipeline that transforms the contents of a GitHub repository extracted and analyzed through Fetch.ai’s uAgents into elegant slides, natural-sounding narration, and a final MP4 video in one automated flow. Built during the Ted AI Hackathon, it combines Gemini models for narration, image generation, and optional TTS with Vectara AI for retrieval-augmented grounding. Visper also integrates Google Cloud TTS and GCS for audio synthesis and cloud delivery. By converting raw repository content into clean, narrated visual summaries, Visper makes technical projects easier to understand and more accessible, especially for blind and low-vision users.

Visper project preview
0
+1

AthleTech.ai - Personalized AI Sports Commentary

Traditional sports broadcasts are one-size-fits-all, failing to engage diverse fans. AthleTech.ai transforms passive viewing into a personalized, interactive experience. Our motivation is to make every fan feel the broadcast was made just for them. Onboarding sets your expertise (beginner-expert), favorite F1 team (for biased commentary), and preferred voice, style, & language (24 options). Our AI generates real-time, context-aware narration, adjusting complexity and bias while maintaining continuity. Ask questions via voice/text and get instant answers on rules or strategy. A Gemini 2.5 Pro/Flash & TTS pipeline powers the experience (Analysis, Generation, Synthesis). It detects key moments, integrates race context, and adapts to preferences live. Our current focus is F1, tailoring technical terms and storylines to each user.

AthleTech.ai - Personalized AI Sports Commentary project preview
0

Declare Agent

Declare Agent saves AI developer from prompt engineering and focuses on what Agent should do. It automatically generates input and output samples and create best agent prompt reflects the developer preference.

Declare Agent project preview
0

Activation Engineering Research

OK, we decided to implement and then think of ideas for applying recent research to gemini! Specifically, we read these 3 papers on representation engineering and activation patches-- steering model behavior by changing their internal representations https://www.nature.com/articles/s44387-025-00031-9 https://arxiv.org/abs/2310.01405 https://arxiv.org/abs/2312.06681 One of us (sevi) did ml work to get this working in a way noone in research has yet, a sliding adjustable steering vector

docs.google.com/…
0

Jarvisos

JarvisOS is an ai agent that lives inside your vr and ar headset and give you intellegence in a way you can interact with it, compared to the old prompt box you can analayze anything using it

www.loom.com/…
0

Monica

An AI-powered WhatsApp bot that automates San Francisco 311 service requests. Users report issues (potholes, graffiti, illegal dumping) via WhatsApp messages, and the system uses Gemini AI to analyze the conversation, extract location and issue details, and automatically submit requests to SF 311. The project includes: - Multi-turn conversational AI for clarifying issue details and location - Multimodal support: analyzes text, images, and video to understand what's being reported - Automated browser-based form submission using Playwright - Real-time dashboard showing message processing and request status - Integration with Twilio WhatsApp API The bot simplifies civic engagement by allowing residents to report issues through natural conversation instead of complex web forms.

Monica project preview
0

Nova

We built Nova - a personal Voice Assistant for macOS. What Siri should have been - powered fully by Google Gemini Family of models for multimodal capabilities.

drive.google.com/…
0

Fuyu

Video generation exiting platform

Fuyu project preview
0

OnCue

Your AI-powered interview wingman & application mentor. OnCue listens to interview questions in real-time and surfaces relevant memories from your experience vault exactly when you need them. OnCue also gives you great suggestions on what to write about in your written applications through an MCP vector search tool.

OnCue project preview
0

VeoQuest

Interactive choose your own adventure games powered by Gemini and Veo 3

screen.studio/…
0

Team Seed

Building Meaningful Connections https://seed.platform.h2mai.com

www.loom.com/…
0
+1

Silicon Valhalla

Silicon Valhalla - Fast-Track Unpredicted Robot Control with Gemini (Gemini Robotics 1.5) This project accelerates robot programming by fusing simulation, vision, and language reasoning into one system. Using Gemini 2.5 Flash, a multimodal model that interprets both images and text, robots in simulation can be guided through natural-language commands instead of code or retraining. Each iteration processes the robot’s camera view, reasons about the scene, and outputs refined motion commands: forming a trial-and-feedback loop that builds precise hardcoded trajectories without datasets or model training. By leveraging Isaac Sim or PyBullet, Gemini’s reasoning bridges perception and control, enabling fast, zero-training robotic demos and radically shortening development cycles.

Silicon Valhalla project preview
0

RememberMe

At the end of events like this one, we all leave with business cards, phone numbers, or LinkedIn contacts. But a few days later, we forget who we actually met. "Remember Me" turns those fleeting encounters into meaningful, lasting connections. Here’s how it works: You meet someone, you take a quick selfie together. Instantly, a QR code appears on your screen. The other person scans it, and it will connect his LinkedIn account to our app and fetch his main information. Then, Gemini AI becomes your memory. You can ask: “Who did I meet in San Francisco last month?” or “Show me everyone I met at the last conference.” Gemini uses metadata, location, date, and encounter context, to rebuild your personal “memory map.” A real-time map make the experience fluid and human-centered. Every action (like opening LinkedIn or adding a contact) is confirmed through a Human-in-the-Loop step for trust and control. And that's why you should Remember Me!

www.youtube.com/…
0
+1

Scroll Scout - Multimodal research assistant

Content marketers spend several hours doomscrolling social media to understand trends, but yep miss out. But, here is the thing doomscrolling isn't the real value add of content marketing but its their ability to be creative based on their understanding of the market. ScrollScout doomscrolls video/image platforms such as instagram/youtube and uses a combination of browser agents and multimodal AI to understand what winning posts/videos at scale like never before. Marketers/market research and brand managers can study markets and people at scale like never before. Last, but not the least but can also use the analysis to create winning content.

Scroll Scout - Multimodal research assistant project preview
0

Stumbler - rabbit holes

Dig deep into rabbit holes. progressively. and responsibly. Find out more about any new topic. or just stumble upon existing ones.

cap.so/…
0

GeminiNow

GeminiNow is an innovative browser extension that combines Google's Gemini AI with real-time voice interaction to create an intelligent assistant that can help users fill out forms, navigate websites, and interact with web content using natural speech.

www.youtube.com/…
0

QuoteScout

AI Agent powered by gemini to identify home service problems, analyze and scope, then find contractors for you

www.loom.com/…
0

TalkBack

TalkBack is a mischievous macOS floating AI companion that acts as your sassy productivity coach. It's a native Swift desktop application featuring real-time voice interaction where users can click-and-hold to speak, and the avatar responds with attitude-filled voice feedback using AI-powered conversation. Key Problem Solved: Developer productivity and mental wellness during coding sessions. TalkBack monitors your terminal in real-time, provides sassy commentary on compilation errors, celebrates build successes, and keeps developers engaged with personality-driven interaction. It combines entertainment with utility - making debugging less frustrating and coding more fun. The app features conversational memory, multilingual voice recognition, and a unique "trash can quit" mechanism. With the new terminal monitoring feature, it watches compilation errors, test results, and command failures, providing real-time snarky but helpful feedback to keep developers motivated.

youtube.com/…
0

Absurdly Visual

Absurdly Visual is an AI-powered social card game that fuses classic party gameplay with modern meme and short-form video culture. Players join themed sessions (like Sports or Politics), where AI generates live, trend-based cards. Participants respond to prompts, an AI player (powered by Gemini) joins the fun, and an AI judge selects the funniest combo. The winning entry is instantly turned into a narrated, TikTok-style video using Gemini-2.5, Gemini TTS, and Veo-3.1, then uploaded to Supabase. Players can also create or monetize their own card packs. Essentially, it’s Cards Against Humanity meets TikTok — turning group humor into shareable AI-generated videos in real time.

Absurdly Visual project preview
0

Mesh

An AI assistant that takes in voice notes, stores information about people you've met, brings up information during meetings, incrementally adds new faces to the detection based on who is present, gives you reminders about previous conversations, and sends action items and summaries after meetings end. It also does live translation, live facial detection, live voice analysis, and has a self-hosted conference calling mechanism and a large backend database to collect information about your network, as well as finds relevant agents in the fetch.ai agentverse through ASI

Mesh project preview
0

Visor

A graduate level engineering tutor that lives in your Meta Ray-Bans and helps solve multi-model engineering problems

Visor project preview
0

ArXini

ArXini is an AI agent that bridges the gap between academic research and practical implementation. It uses Google Gemini to analyze any ArXiv paper and its GitHub repository, allowing users to interactively explore how theoretical concepts are translated into code. ArXini transforms a complex research task into a simple, on-demand service and allows you to get onboarded / start working on existing implementations much faster.

ArXini project preview
0

Kursor

Kursor AI is a multimodal desktop copilot that lives on your mouse, powered by Google Gemini. It sees your screen, reads text, and listens to your voice to offer instant, context-aware help without breaking your flow. With one hotkey, it works anywhere — Gmail, Docs, Figma, VS Code, or your browser. It captures what’s under your cursor and sends it to Gemini for smart actions like rewriting text, summarizing content, adding calendar events, drafting emails, or explaining on-screen code and charts. Kursor AI supports an open plugin ecosystem, so anyone can build actions (e.g., send to Notion, query a database, analyze data). Built with: Gemini 2.5 Flash, Python + Electron, PyAutoGUI, clipboard + screen capture, open plugin manifest. Vision: A privacy-first platform where users design their own AI workflows, making your mouse an intelligent, context-aware companion.

Kursor project preview
0

GeminiDesk

Modern knowledge workers students, freelancers, small business owners spend too much time moving information around. GeminiDesk eliminates this friction by letting users drop any file or message (text, email, PDF, or image) into a single interface, integrating with notion and google calendar.

GeminiDesk project preview
0

Rizzler.ai

CloutFarm (formerly Rizzler AI) is a content creation tool that analyzes a user's raw photos and audio clips to identify currently trending sounds, filters, and formats on TikTok. It then automatically edits the user's media into a new, polished AI-generated reel that mirrors the viral trend, making it easy for anyone to hop on the bandwagon and create content with high viral potential.

drive.google.com/…
0

sciCore-LUCA

LUCA C is our proof-of-concept: an AI cell that “sees” biological structure and reconstructs it as a living, reasoning simulation. It’s built on Gemini multimodal AI and open-source agentic systems, forming the first member of a lineage of LUCA-based autonomous simulators. SciCore will host these agents — not just as demos, but as a framework for research, digital twins, and data-driven discovery. You can even map this visually on the homepage: LUCA C → LUCA M → LUCA E, forming an evolutionary chain.

copy-of-luca-by-scicore-373075233891.us-west1.run.app/…
0

ConfigGuardian

ConfigGuardian is an AI-powered configuration security agent that analyzes raw configuration files or even screenshots like Kubernetes YAMLs or Dockerfiles to automatically detect misconfigurations and propose secure fixes. Using Google’s multimodal Gemini models, ConfigGuardian can parse text or extract it directly from images, reason about risky settings such as running as root, unpinned images, or exposed ports, and then generate secure auto-fix code diffs. Finally, it produces a cryptographic provenance record anchored via a Fetch.ai agent, providing tamper-proof, verifiable audit trails for every scan. In short, ConfigGuardian turns raw or visual configuration data into AI-audited, auto-fixed, and verifiably trustworthy security insights — bridging the gap between developer convenience and compliance-grade integrity.

ConfigGuardian project preview
0

NY

MEMORYSTREAM AI is a multimodal AI platform that transforms spoken memories into cinematic videos within minutes. It addresses a universal problem — every day, priceless family stories and cultural heritage are lost forever because preserving them is too time-consuming, expensive, or left until it’s too late. With MEMORYSTREAM AI, users simply talk about a memory while the AI listens empathetically, asks intelligent follow-up questions, and optionally analyzes old photos for context. Using Google Gemini 2.5 for conversation and vision, Google Veo for cinematic video generation, and Pipecat for real-time audio processing, the platform fuses voice, visuals, and storytelling into a seamless experience. In just a few minutes, users can create stunning, lifelike memory videos—preserving their stories, emotions, and heritage for generations to come.

drive.google.com/…
0

Rodela

Trials Intel automates clinical trial intelligence for pharma investors using Gemini's multimodal capabilities. The agent autonomously monitors ClinicalTrials.gov, classifies trials by therapeutic area and commercial potential, then uses Gemini Vision to extract quantitative data from survival curves and adverse event tables in published trial results. It synthesizes insights across multiple sources to generate market trends, investment opportunities, and competitive intelligence. What currently requires teams of analysts at $100-500K/year is automated in minutes. Built entirely with Google Gemini 2.0 for true text + vision multimodal analysis.

www.loom.com/…
0

Gemini & I

A platform to create AI generated video podcasts (like in LM Notebook but only video) and ability to call in and talk to the hosts

Gemini & I project preview
0