Gemini Omni Flash 1.1 for Text to Video
Generate video with sound from a text prompt using Google's Gemini Omni Flash 1.1. Describe a scene, set the duration and aspect ratio, and hit run.
Video
63
Nodes & Models
GeminiOmniFlash11TextToVideo_floyo
VideoToFrames
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Generate a Video from Text Describe a scene and get a short video clip with synchronized audio. The model generates picture and sound in a single pass, so what you download is a complete clip. No stitching, no separate audio step.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
Gemini Omni 1.1 Flash by Google DeepMind. A multimodal video model that generates video with native audio from a text prompt, with cinematic camera control and real-world physics understanding.
HOW IT WORKS
Step 1. Write your prompt Describe the scene you want. Name the subject, the action, the camera movement, and the lighting. The more specific, the better the result. Works great with: cinematic shots · landscapes · product scenes · motion sequences
Step 2. Set the duration Choose how long the clip should be. The range is 3 to 10 seconds.
Step 3. Pick the aspect ratio and resolution Choose 16:9 for landscape or 9:16 for portrait. Resolution defaults to 720p.
Step 4. Hit run and download The model generates a video with matching audio in one pass. Preview it in the workflow, then download. Ready for: Premiere Pro · DaVinci Resolve · After Effects · any video editor
First time? Leave every setting as-is. The defaults (8 seconds · 16:9 · 720p) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard video clip (most people) — 8 seconds · 16:9 · 720p. The right starting point for almost everyone.
Quick concept test — 3 seconds · 16:9 · 720p. Fastest way to check if your prompt reads well before committing to a longer clip.
Vertical video for social — 8 seconds · 9:16 · 720p. Use this for Instagram Reels, TikTok, or YouTube Shorts format.
Shorter clip for a loop or transition — 3 to 5 seconds · 16:9 · 720p. Good for B-roll, intros, or loopable background footage.
Maximum length per run — 10 seconds · 16:9 · 720p. The longest clip you can generate in a single pass.
The motion or subject is off — Rewrite the prompt before changing a setting. Name the subject, action, camera angle, and lighting. Specific, visual language works best.
Prompt: Describe one continuous shot. Name the subject, what it does, how the camera moves, and what the light looks like. "A slow aerial drone shot gliding over a misty valley at sunrise, golden light through fog, cinematic" is strong. "Cool nature video" is too vague.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Filmmakers & Video Editors Generate B-roll, establishing shots, or concept clips from a text description. Get video with matching audio in one pass, ready to drop into a timeline.
📱 Social Media Creators Create short-form vertical or landscape video for Reels, TikTok, or Shorts without a camera, a location, or a crew.
🎨 Concept Artists & Art Directors Pitch a scene before it exists. Write the shot, generate it, and share a moving reference instead of a static mood board.
🎵 Music & Audio Projects The model generates synchronized audio alongside the picture. Use it for music video concepts, ambient clips, or sound-matched visual sequences.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Single continuous camera shots
Cinematic, well-described scenes
Landscape and environment shots
Specific lighting and mood descriptions
⚠️ May produce softer results
Vague prompts like "cool video" or "something interesting"
Multi-cut sequences in one prompt
Precise text or signage in the scene
Complex multi-character interactions
FAQ
What is Gemini Omni Flash 1.1? Gemini Omni 1.1 Flash is a video generation model from Google DeepMind. It takes a text prompt and produces a video clip with synchronized audio in a single pass. It supports durations from 3 to 10 seconds per run, at up to 720p native resolution, with 16:9 or 9:16 aspect ratios.
Does Gemini Omni Flash 1.1 generate audio with the video? Yes. The model generates audio in the same pass as the picture. The output is a complete video file with synchronized sound, including ambient noise, effects, or music that matches the scene. There is no separate audio generation step.
What resolution does Gemini Omni Flash 1.1 support? This workflow supports 720p output. The model can also produce 360p, 1080p, and 4K output in other configurations, though Google describes 1080p and 4K as upscaled rather than natively generated. For this text-to-video workflow, 720p is the standard quality tier.
How long can a Gemini Omni Flash 1.1 video be? Each run generates a clip between 3 and 10 seconds. For longer sequences, Google supports scene extension up to 40 seconds through chained calls in separate workflows. This workflow generates a single clip per run.
Do Gemini Omni Flash 1.1 outputs have a watermark? Yes. Google applies an invisible SynthID watermark to all video generated by Gemini Omni models. It does not change how the video looks or sounds, but it stays embedded in the file to mark it as AI-generated.
Can I use Gemini Omni Flash 1.1 video commercially? Gemini Omni 1.1 Flash is a proprietary Google model, so commercial use is governed by Google's terms rather than an open license. Commercial use is generally allowed, but review Google's current terms for your specific use case. Note that all outputs carry the SynthID watermark.
How to run Gemini Omni Flash 1.1 online? You can run Gemini Omni Flash 1.1 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, write your prompt, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Write your first prompt and hit run. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more




