Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

FLUX 3 · Text to Video

FLUX 3 is Black Forest Labs' multimodal video model that turns a text prompt into an HD clip up to 20 seconds, with multi-shot scenes. Type a prompt and run.

59

Gen time: ~2 min 22 secs

Nodes & Models

Flux3TextToVideo_floyo
VideoToFrames
FloyoStickyNote
CreateVideo
SaveVideo

ABOUT THE WORKFLOW

Turn text into video
Type a prompt describing the scene, action, and camera, and the model generates an HD video clip up to 20 seconds. It handles multi-shot sequences and readable text in the scene, and it draws on broad world knowledge, so a short, clear description often works better than a long shot list.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • FLUX 3 by Black Forest Labs. The lab's first multimodal video model, trained jointly on images, video, and audio in one architecture. It generates HD or Full HD clips up to 20 seconds from a text prompt, with multi-shot sequences, in-scene text, and lip-synced dialogue in more than 14 languages.


HOW IT WORKS

Step 1. Write your prompt
Describe the scene, action, camera move, and any dialogue or sounds you want. A short, clear description often beats a long shot list.
Works great with: cinematic scenes · multi-shot sequences · documentary-style clips

Step 2. Hit run
The model generates the video and audio together in one pass. This is a large model, so give it longer to finish.

Step 3. Download your video
You get one HD clip. The file saved here is picture only, even though the model generates audio.
Ready for: editing timelines · social · ads

First time? Leave every setting as-is. The defaults (auto ratio · 720p · 5 seconds) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard use (most people) — auto aspect ratio, 720p, 5 seconds. The right starting point for almost everyone.

  • Making vertical content — set the aspect ratio to 9:16 for phones and short-form feeds.

  • Want higher quality — switch the resolution to Full HD. It sharpens the clip and costs more per second.

  • Want a longer scene — raise the duration up to the 20-second maximum. Anything past 20 seconds needs a continuation step.

  • Keeping cost down — lower the duration and the resolution, since both drive the per-run cost. A 20-second Full HD clip costs much more than a 5-second HD one.

  • Want multiple shots in one clip — describe the sequence in the prompt, like "wide shot of the market, then a close-up of the vendor," and the model can switch scenes within one generation.

Prompt: Write the idea at the level you think about it and let the model fill in the details. "A street food vendor grilling skewers at night, neon signs, camera pushes in, then cuts to a close-up of the flames" gives it a scene and a shot change to work with. The model has strong world knowledge, so you rarely need a long shot list.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Ad & Promo Clips
Generate short branded spots from a written brief.

📽️ Multi-shot Scenes
Get several connected shots and camera angles in a single clip.

📚 Documentary & Explainer
Use the model's world knowledge to visualise a topic from a short prompt.

📱 Social & Short-form
Make vertical clips for feeds without a shoot.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Short, clear prompts with a defined scene

  • Multi-shot sequences and camera moves

  • In-scene text and typography

  • Clips of 20 seconds or less

⚠️ May produce softer results

  • Long, over-specified shot lists

  • Clips pushed past 20 seconds in one run

  • Many unrelated actions in one prompt

  • Fine detail that needs a still image instead


FAQ

What is FLUX 3?
FLUX 3 is Black Forest Labs' first multimodal video model, announced in July 2026 and made generally available in August 2026. It is trained jointly on images, video, and audio, and generates HD video up to 20 seconds from a text prompt, with multi-shot sequences, in-scene text, and lip-synced dialogue in more than 14 languages. This workflow runs it in text-to-video mode.

Does the saved video have sound?
No. FLUX 3 generates native audio alongside the picture, but this workflow saves the clip as picture only, because it rebuilds the video into frames without the sound track. You get a silent HD clip, so plan to add or replace audio in your editor if you need it.

How long can a FLUX 3 clip be?
A single generation runs up to 20 seconds, which is longer than the roughly 10-second ceiling on many video models. Clips are HD by default, with Full HD available, and anything longer than 20 seconds needs a continuation step rather than one pass.

Is FLUX 3 open source?
Not yet. FLUX 3 is available through an API at launch, and Black Forest Labs has said an open-weights FLUX 3 Dev version is planned but has not released it. For now you run it as a hosted model rather than local weights.

Can I use FLUX 3 videos commercially?
FLUX 3 is a proprietary model accessed through Black Forest Labs' API, so commercial use is governed by BFL's terms and its pay-as-you-go pricing rather than an open license. Commercial use is generally supported for ads and content, but review BFL's current terms and make sure you have the rights to anything you describe or reference.

How is FLUX 3 different from FLUX.2?
FLUX.2 is an image model, while FLUX 3 is Black Forest Labs' first video model and its first multimodal foundation model, handling image, video, and audio in one architecture. In practice, FLUX.2 is what you reach for to generate stills, and FLUX 3 is what you use to generate video with motion and sound.

How to run FLUX 3 online?
You can run FLUX 3 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, type a prompt, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A creator generates a clip and likes it. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Type a prompt and run it. You get an HD clip up to 20 seconds.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N
FloYo: FLUX 3 · Text to Video