Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Gemini Omni Flash 1.1 for Image to Video

Turn a still image into a short video clip with synchronized audio using Google's Gemini Omni Flash 1.1. Upload a photo, describe the motion, and hit run.

Gemini Omni Flash 1.1
Image to Video
Video

163

Gen time: ~1 min 40 secs

Nodes & Models

GeminiOmniFlash11ImageToVideo_floyo
VideoToFrames
LoadImage
CreateVideo
SaveVideo

ABOUT THE WORKFLOW

Animate a Photo into Video
Upload an image and describe the motion you want. Gemini Omni Flash 1.1 generates a short video clip with synchronized audio in one pass. Add an optional end frame to control where the motion lands. That's it.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • Gemini Omni Flash 1.1 by Google. A multimodal video model that generates video with native synchronized audio from a single image and prompt. Strong at cinematic camera movement, natural motion, and coherent audio matched to the scene.


HOW IT WORKS

Step 1. Upload your image
The photo your video starts from. Any clear, well-lit shot works.
Works great with: portraits · landscapes · product shots · illustrations

Step 2. Describe the motion
Write what moves and how. Be specific about camera, subject action, and timing. Example: "Slow push-in, the woman turns and smiles, petals drift in a gentle breeze."

Step 3. Add an end frame (optional)
Upload a second image for the video to finish on. The model interpolates between the two frames. Leave it empty to let the motion play out freely.

Step 4. Hit run and download
Gemini Omni Flash 1.1 generates the video with matched audio. Preview it in the workflow, then download.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any editor

First time? Leave every setting as-is. The defaults (8 seconds · 16:9 · 720p) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard video (most people) — 8 seconds · 16:9 · 720p. The right starting point for almost everyone.

  • Quick test before committing — 3 seconds · 360p. Fastest way to check framing and motion before a full-resolution run. Costs roughly a third of a 720p generation.

  • Delivery-quality output — 8 seconds · 1080p or 4K. Use this when the clip goes into a final edit or client delivery. Note: 1080p and 4K are upscaled from the native 720p render, so lock the shot at 720p first, then rerun the keeper at higher resolution.

  • Portrait or vertical video — 8 seconds · 9:16 · 720p. For Reels, TikTok, or Shorts. Framing a subject for a tall frame works better than cropping a landscape clip later.

  • Control start and end frames — Upload both an image and an end frame. The model interpolates between the two, giving you precise control over where the motion begins and ends.

  • The motion looks wrong or choppy — Rewrite the prompt with more specific camera and timing cues. "Slow dolly-in, shallow depth of field" lands better than "cinematic video."

Prompt: Describe the motion, camera movement, and mood in one clear paragraph. Name the subject action, the camera move, and the lighting. "Slow push-in, the man lifts a coffee cup, warm morning light, shallow depth of field" works. "Make a nice video" does not. Add audio cues if you want specific sounds: "quiet café ambience, no music."


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Short-form Content Creators
Turn a single product photo or scene into a ready-to-post video clip with sound for Reels, TikTok, or Shorts.

🎨 Concept Artists & Storyboard Artists
Animate a keyframe to pitch motion and timing to a director before committing to full production.

🛍️ E-commerce & Product Marketing
Bring a static product shot to life with camera movement and ambient audio for ads, landing pages, or social posts.

🎮 Game Developers & 3D Artists
Generate animated reference footage or in-engine cutscene prototypes from concept art or renders.

🎵 Music & Audio-Visual Projects
Pair a cover image or artwork with generated audio and motion for visualizers, lyric videos, or social teasers.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-lit single-subject photos

  • Specific camera movement in the prompt (dolly, push-in, pan)

  • Landscapes, portraits, and product shots

  • Audio cues described in the prompt (ambience, dialogue, effects)

⚠️ May produce softer results

  • Vague prompts like "make it move" or "cinematic video"

  • Heavy bokeh or shallow depth of field in the source image

  • Extreme low-light or noisy source photos

  • Expecting native 4K sharpness (1080p and 4K are upscaled from 720p)


FAQ

What is Gemini Omni Flash 1.1?
Gemini Omni Flash 1.1 is Google's multimodal video generation model, released in August 2026. It takes text, images, and video references as input and produces video clips with synchronized audio. Unlike models that handle video and audio separately, Gemini Omni Flash 1.1 generates both in a single pass, so the sound matches the scene without post-production syncing.

Does Gemini Omni Flash 1.1 generate audio with the video?
Yes. Audio is generated natively alongside the video. The model produces sound effects, ambient audio, and even dialogue matched to the visuals. You can guide the audio by adding cues in your prompt, like "quiet room tone, the sound of footsteps on wood" or "upbeat background music." The audio track downloads as part of the video file.

What resolution and duration does Gemini Omni Flash 1.1 support?
Each clip runs 3 to 10 seconds at 24 FPS, with 8 seconds as the default. Resolution options are 360p, 720p, 1080p, and 4K. The model renders natively at 360p and 720p. The 1080p and 4K options are upscaled from the 720p generation. For the best workflow: draft and iterate at 720p, then rerun the final take at 1080p or 4K for delivery.

How is Gemini Omni Flash 1.1 different from Sora, Kling, or Wan 2.2?
Gemini Omni Flash 1.1 generates video and audio together in one pass, so the sound is matched to the scene from the start. Most other models output silent video that needs a separate audio step. It also supports start-and-end frame control, where you provide two images and the model interpolates between them. The trade-off: individual clips cap at 10 seconds, and higher resolutions are upscaled rather than natively rendered.

Can I control where the video ends with a second image?
Yes. Upload an end frame alongside your start image, and the model interpolates the motion between them. This gives you precise control over the start and end poses, which is useful for transitions, looping clips, or matching a storyboard.

Do the outputs have a watermark?
Google applies an invisible SynthID watermark to all video generated by Gemini Omni Flash 1.1. It does not change how the video looks or sounds, but it stays embedded in the file for provenance verification.

Can I use the output commercially?
Gemini Omni Flash 1.1 is a proprietary Google model, so commercial use is governed by Google's terms rather than an open license. Commercial use is generally allowed, but review Google's current terms and note that outputs carry the SynthID watermark.

How to run Gemini Omni Flash 1.1 online?
You can run Gemini Omni Flash 1.1 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your image, describe the motion, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a photo, describe the motion, and run it. The prompt and settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N