# Elevate

- **Event:** [Google DeepMind Bangalore Hackathon](https://cerebralvalley.ai/e/google-deepmind-bangalore-hackathon)
- **When:** Sat, Jul 11 at 9:00 AM – 10:00 PM (GMT+5:30)
- **Where:** Marathahalli, Marathahalli Main Road
- **Team:** [Pranav Gajjala](https://cerebralvalley.ai/u/thegdpranavl), [Inchara P](https://cerebralvalley.ai/u/inchara_p)
- **GitHub:** https://github.com/gdpranavl/omni-studio
- **Demo video:** https://youtu.be/x2OamT3YHWU
- **Gallery:** https://cerebralvalley.ai/e/google-deepmind-bangalore-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/google-deepmind-bangalore-hackathon/hackathon/gallery/20

Omni Studio

Direct a video without touching a camera.

Omni Studio turns a single typed idea into a finished short-form reel - no filming, no editing software, no crew. You describe the concept, an AI creative director works out the shot list with you, and the pipeline carries it the rest of the way to a scrubbable, editable timeline ready to export.

What it does

1. Ideate - Chat with a Gemini-powered creative director. Upload a reference photo (your face, a product, anything) and Gemini clarifies the concept into a scene-by-scene plan - talking-head beats and B-roll beats, mapped out with dialogue, shot direction, transitions, and text treatments.
2. Production Sheet - Gemini turns that conversation into a structured, fully editable shot list: what's said, how it's delivered, what the camera does, what music/SFX play under it.
3. Scene Generation - Each reference photo is expanded into a full turnaround - multiple angles generated from one shot - so every scene has consistent, on-model frames to build from.
4. Clip Generation with Omni Flash - Those frames are handed to Omni Flash, which turns still images into motion: a talking-face frame becomes a delivered line, a B-roll frame becomes a moving shot. This is the step that actually gets you from *photo* to *footage*.
5. Assemble - A real Remotion-powered timeline stitches the generated clips together with transitions, text overlays, and animation - drag to reorder, trim in place, scrub the playhead, cycle transitions - then export.

Why Gemini + Omni Flash

The two models split the job the way a real production splits it: Gemini is the director - it interprets your idea, asks the right clarifying questions, and writes the shot list, the same reasoning a human creative director would do before a single frame exists. Omni Flash is the camera operator - once the direction and reference frames exist, it's the model that actually produces the moving footage, image-to-video, shot by shot. Neither one alone gets you a finished reel: Gemini can plan a video but can't shoot it, and a video model can't plan a video without direction. Chaining them - plan with Gemini, shoot with Omni Flash, assemble with Remotion - is what makes going from a text idea to an exportable reel possible without a camera.

Built with

Next.js (frontend) + Express (backend) monorepo, Gemini 2.5 (ideation + production-sheet generation, streamed live), Omni Flash (image-to-video clip generation), Nano Banana for reference-frame turnarounds, and Remotion for the real, interactive assembly timeline.

---

Markdown version of https://cerebralvalley.ai/e/google-deepmind-bangalore-hackathon/hackathon/gallery/20. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
