Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Gemini Omni Flash 1.1 for Reference to Video

Turn a reference image into a short video with synchronized audio using Gemini Omni 1.1 Flash. Upload an image, describe the motion, and hit run.

Gemini Omni Flash 1.1
r2v
Video

90

Gen time: -- secs

Nodes & Models

GeminiOmniFlash11ReferenceToVideo_floyo
VideoToFrames
LoadImage
CreateVideo
SaveVideo

ABOUT THE WORKFLOW

Turn a Reference Image Into Video
Upload a reference image and describe how the scene should move. The model generates a short video clip with synchronized audio in one pass. Add more reference images or short video clips to guide the look.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • Gemini Omni 1.1 Flash by Google DeepMind. A multimodal video generation model that takes images, video clips, and text as input and produces video with synchronized audio. Strong at holding subject identity from a reference image.


HOW IT WORKS

Step 1. Upload your reference image
The image that sets the subject, scene, and look for your video.
Works great with: portraits · product shots · landscapes · concept art

Step 2. Describe the motion
Write what should happen in the scene. Be specific about what moves, how the camera behaves, and the mood. Example: "The camera slowly zooms out to reveal the full scene. A gentle breeze moves the trees. Warm golden-hour lighting."

Step 3. Add more references (optional)
Connect up to 9 more reference images to guide identity, style, or composition. You can also add up to 3 short video clips (3 seconds each) to guide the motion style.

Step 4. Hit run and download
The model builds the video and audio together in one pass, then returns a finished clip. Preview it in the workflow, then download.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any NLE

First time? Leave every setting as-is. The defaults (8 seconds · 16:9 · 720p) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard use (most people) — 8 seconds · 16:9 · 720p. The right starting point for almost everyone.

  • Quick preview or test — 360p · 8 seconds. Fastest and cheapest way to check motion and composition before committing to a full-res run.

  • Vertical video for social — 9:16 · 720p · 8 seconds. Use this for Instagram Reels, TikTok, or YouTube Shorts.

  • Longer clip — Increase duration up to 10 seconds. Longer clips cost more per run.

  • Sharper output for delivery — 1080p or 4K. Note that 1080p and 4K are upscaled from 720p native, not true higher resolution. Use 720p for drafts and reviews.

  • Guide identity or style with multiple references — Add more reference images to lock in a specific look, character, or environment across the clip.

  • Motion is not matching the prompt — Rewrite the prompt before changing settings. Describe what moves, how it moves, and where the camera goes. Short, direct sentences work better than long paragraphs.

Prompt: Describe the motion, camera behavior, and atmosphere. "The camera slowly pulls back, revealing the full room. Soft ambient light, shallow depth of field" is stronger than "make a nice video of the room." Keep sentences short and physical.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Short-Form Content Creators
Turn a single product shot or scene photo into a short video with sound for Instagram Reels, TikTok, or YouTube Shorts.

🎨 Concept Artists & Art Directors
Bring a still concept frame to life to pitch a scene, test motion direction, or preview how a mood board image reads in motion.

🛍️ Product & E-Commerce Teams
Animate a product photo into a short hero clip for a landing page or ad campaign without booking a video shoot.

🎥 Video Editors & Motion Designers
Generate reference motion clips or placeholder footage from a still frame to block out an edit or pitch a sequence before production.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-lit reference images with a single subject

  • Short, specific motion prompts (what moves, how, where the camera goes)

  • Portraits, product shots, landscapes, interiors

  • Multiple reference images to lock in a consistent look

⚠️ May produce softer results

  • Complex multi-step actions described in one prompt

  • Heavily composited or collaged reference images

  • Extreme low-light or high-contrast source photos

  • Prompts longer than 3-4 sentences (shorter is better)


FAQ

What is Gemini Omni 1.1 Flash and what does it do?
Gemini Omni 1.1 Flash is a video generation model from Google DeepMind, released on August 27, 2026. It takes text, images (up to 10), and short video references (up to 3 seconds each) as input and produces video clips with synchronized audio. It is not a chatbot or general-purpose assistant. It is built to generate, edit, and extend video.

Does Gemini Omni 1.1 Flash generate audio with the video?
Yes. The model produces video and audio together in a single generation pass. Sound is synchronized to the scene, so dialogue, ambient noise, effects, and music can be directed from inside the prompt. There is no separate audio step or extra charge for sound.

What resolution does Gemini Omni 1.1 Flash output?
The native output resolution is 720p at 24 fps. The model also supports 360p (for fast, low-cost previews), 1080p, and 4K. However, 1080p and 4K are upscaled from the native 720p render, not true higher-resolution generation. Use 720p for most work and only step up when the final deliverable requires it.

How is this different from Veo 3.1 or other Google video models?
Veo 3.1 is Google's dedicated high-fidelity video generation model, built for cinematic output. Gemini Omni 1.1 Flash is designed around a reference-based workflow: you provide images and short clips as input, and the model generates video that holds the identity, lighting, and composition of those references. If you already have a reference photo that defines the look, Gemini Omni 1.1 Flash is the more direct path.

Can I use Gemini Omni 1.1 Flash video commercially?
Gemini Omni 1.1 Flash is a proprietary Google model. Commercial use is allowed under Google's terms of service. All outputs carry an invisible SynthID watermark for provenance detection. Review Google's current usage policies for your specific use case, and make sure you hold rights to any reference images you upload.

What is SynthID and does it affect the video?
SynthID is Google's invisible watermark embedded in all Gemini Omni outputs. It marks the video as AI-generated. The watermark does not change how the video looks or sounds to viewers, but it can be detected programmatically for provenance verification.

How to run Gemini Omni 1.1 Flash online?
You can run Gemini Omni 1.1 Flash online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your reference image, describe the motion, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a reference image, describe the motion, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N