MiniMax H3 Max Turbo · Text to Video
Generate short video clips with sound from a text prompt using MiniMax H3 Max Turbo, fal's speed-optimized video model. Describe a scene, hit run, and get a video with audio back in seconds.
ai video
minimax h3 max turbo
text to video
video generation
97
Nodes & Models
MiniMaxH3MaxTurboTextToVideo_floyo
VideoToFrames
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Generate a Video from Text
Describe a scene and get a short video with stereo audio back. MiniMax H3 Max Turbo generates the picture and sound in one pass. Clips run 5 to 15 seconds at up to 768P.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
MiniMax H3 Max Turbo, post-trained by fal from MiniMax H3. The speed tier of the H3 Max family, generating a 5-second clip in about 1 to 2 seconds. Strong at prompt adherence and generates synchronized audio alongside the picture.
HOW IT WORKS
Step 1. Write your prompt
Describe the scene you want. Subject, action, camera move, lighting, mood, and the sound you want to hear. The prompt field is empty by default.
Works great with: cinematic shots · product clips · social content · motion design
Step 2. Pick duration and ratio
Choose 5, 10, or 15 seconds for the clip length, and an aspect ratio for the frame. Longer clips cost more.
Step 3. Hit run and download
The model generates the video with audio and returns it as a video file. Preview it in the workflow, then download.
Ready for: Premiere · DaVinci Resolve · After Effects · any NLE
First time? Leave every setting as-is. The defaults (768P · 5 seconds · 16:9 · balanced prompt expansion · random seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard generation (most people) — 768P · 5 seconds · 16:9 · balanced · random seed. The right starting point for almost everyone.
Want a longer clip — Raise duration to 10 or 15 seconds. Each step costs more per run.
Want a quick test before committing — Drop resolution to 480P and keep duration at 5 seconds. Fastest way to check a prompt before running at full settings.
Want a vertical video — Switch ratio to 9:16 for stories, reels, or portrait content.
Want to reproduce a result — Lock the seed to a specific number. The same seed with the same prompt returns the same clip.
The video does not match the prompt — Rewrite the prompt before changing settings. Clear scene direction matters more than resolution. Include what the camera does, how the subject moves, and what sounds are in the scene.
Prompt: Write it like a shot brief. "A woman walks through a rainy Tokyo street at night, neon reflections on wet pavement, slow tracking shot, city ambience" works better than "woman in the rain." Include sound cues. The model generates audio alongside the picture, so describing what you want to hear shapes the result.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Short-Form Social Content
Generate 5 to 15 second clips with audio for reels, stories, or TikToks. Vertical 9:16 ratio is one dropdown away.
🛍️ Product & E-commerce
Create product reveal clips or lifestyle shots with ambient sound, without a video shoot or stock footage license.
🎨 Motion Moodboards
Test a scene direction before committing to a full production. Run several takes with different prompts and compare them side by side.
🔊 Audio-Visual Prototyping
The model generates stereo audio in the same pass as the video. Describe both the picture and the sound to get a clip that carries its own atmosphere.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Shot-brief style prompts with subject, action, camera, and sound
5-second clips for fast iteration
Cinematic scenes with clear lighting and mood direction
Prompts that describe what the camera does
⚠️ May produce softer results
Vague one-line prompts with no camera or action direction
Expecting sharp detail past 768P (use base H3 for 2K)
Long clips (15 seconds) with many scene changes in one prompt
Precise character likeness or text rendering in the video
FAQ
What is MiniMax H3 Max Turbo?
MiniMax H3 Max Turbo is a speed-optimized video generation model post-trained by fal from MiniMax's open-weight H3. Released as a preview in September 2026, it generates 5 to 15 second clips at up to 768P with stereo audio in one pass. A 5-second clip returns in about 1 to 2 seconds.
Does MiniMax H3 Max Turbo generate audio?
Yes. The model generates synchronized stereo audio alongside the video in the same pass. Include sound cues in your prompt to shape what you hear. There is no separate audio step or model.
What is the difference between H3 Max Turbo, H3 Max, and base MiniMax H3?
MiniMax open-sourced H3 in August 2026. fal post-trained those weights into H3 Max for stronger prompt adherence and better aesthetics, then distilled H3 Max into Turbo for speed. Turbo is about twice as fast as H3 Max and costs less per second. H3 Max is the quality tier. Base H3 supports 2K output if you need higher resolution. This workflow uses Turbo.
What resolution and duration does H3 Max Turbo support?
Output goes up to 768P at 24 fps. Clips can be 5, 10, or 15 seconds. For 2K output, use the base MiniMax H3 model instead. Duration is the main cost driver.
Can I use MiniMax H3 Max Turbo outputs commercially?
The base MiniMax H3 model is open source under Apache 2.0, but the fal post-training on H3 Max Turbo is proprietary. Review fal's current terms of service for commercial use of outputs generated through their API.
How is this different from image-to-video workflows?
This workflow creates video from a text prompt only. There is no image input. If you have a starting frame you want to animate, use an image-to-video workflow instead. H3 Max Turbo also supports image-to-video in a separate workflow.
How to run MiniMax H3 Max Turbo online?
You can run MiniMax H3 Max Turbo online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, write a prompt, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs a clip and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Write a scene brief and run it. Five seconds of video with audio, back in seconds.
Questions? Watch the free course or check the FAQ above.
Read more
%20(1)_1775891000879.webp?width=400&height=300&quality=80&resize=contain&format=origin)
%20(3)_1774349172672.webp?width=400&height=300&quality=80&resize=contain&format=origin)


