Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

MiniMax H3 Open Weights TI2V with Style Embedding

Generate video with native stereo audio using MiniMax H3 open weights. Write a prompt, add an optional start image and a style embedding, and hit run.

I2V
Minimax H3
Open Source
Open Weights
T2V
Video

81

Gen time: ~4 min 36 secs

Nodes & Models

ResolutionSelector
VAELoader
KSamplerSelect
UNETLoader
CLIPLoader
RandomNoise
PrimitiveFloat
LoadImage
Note
BasicScheduler
ComfyMathExpression
MiniMaxH3ImageToVideo
BasicGuider
SamplerCustomAdvanced
VAEDecode
VAEDecodeAudio
CreateVideo
SaveVideo

ABOUT THE WORKFLOW

Generate Video with Sound Write a prompt and get a video with synchronized stereo audio in one pass. Add a start image to anchor the first frame. A style embedding shapes the visual look without rewriting your scene. That's it.

Model

  • MiniMax H3 (Hailuo 3.0) by MiniMax. A 33B open-weight video model that generates 768p video at 24 fps with native stereo audio from text, images, or both. Released August 2026 under the MiniMax H3 Community License.


HOW IT WORKS

Step 1. Write your prompt Describe the scene, the action, the camera move, and the mood. A style embedding is prepended by default, so write your scene after it. Works great with: cinematic scenes · action sequences · atmospheric shots

Step 2. Add a start image (optional) Upload an image to set the first frame. The model animates outward from there. Leave it empty for text-only generation.

Step 3. Pick a style embedding (optional) The default is bullet_time. Swap it for a different embedding tag in the prompt to change the visual style without rewriting the scene. Ten community embeddings are available, including dark_magic, fire_breath, art_is_explosion, and four_seasons.

Step 4. Hit run and download MiniMax H3 builds the video and audio together in one pass, then saves a video file with stereo sound. Ready for: Premiere Pro · DaVinci Resolve · After Effects · any NLE

First time? Leave every setting as-is. The defaults (5 seconds · 16:9 · 20 steps · bullet_time embedding · random seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard generation (most people) — 5 seconds · 16:9 · 20 steps · bullet_time embedding · random seed. The right starting point for almost everyone.

  • Quick test before a long render — Drop steps to 10 or 12. The clip will be rougher, but you can check framing and motion before committing to a full run.

  • Different visual style — Swap the embedding:minimaxh3_bullet_time tag in the prompt for another embedding name. dark_magic, fire_breath, storm_magic, art_is_explosion, truman_show, blooming_flowers, four_seasons, spiral_ascent, and kiss_camera are all available.

  • Start from a specific frame — Upload a start image. The model reads the lighting, colors, and composition from the image and animates outward.

  • Portrait or square output — Change the aspect ratio in the resolution selector. 9:16, 1:1, and other ratios work at 0.9 megapixels.

  • Reproduce a result you liked — Lock the seed to a fixed number instead of random. Same seed, same prompt, same output.

  • The motion is not matching the prompt — Rewrite the prompt first. Motion words matter more than steps or resolution. Describe what moves, in what direction, and how fast.

Prompt: Describe the scene and the motion, not the style. The embedding handles style. "A drone flies low over a calm lake at sunrise, mist lifting off the surface" is better than "cinematic beautiful misty lake drone shot." Be specific about what moves and where the camera goes.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Short Film & Music Video Pre-vis Write a scene, pick a style embedding, and get a clip with stereo audio to pitch a sequence or block a shot before committing to a full production.

🎮 Game Trailers & Cinematics Generate cinematic clips from concept art or in-engine screenshots. Drop a start frame in and let H3 animate the camera move and the soundscape.

📱 Social Media & Ads Turn a product photo or a brand scene into a short video with matching audio. Swap aspect ratios for Instagram Reels, TikTok, or YouTube Shorts.

🎨 Motion Concept Art Explore how a scene feels in motion. Upload a still concept painting, describe the action, and see it play out with sound before handing it to a motion team.

🔊 Audio-Synced Prototypes The native stereo audio means dialogue, effects, and ambience generate alongside the picture. No separate sound design pass for early-stage work.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Cinematic scenes with clear camera direction

  • Single-subject action (a person walking, a car driving, water flowing)

  • Atmospheric and environmental shots

  • Start images with strong lighting and composition

⚠️ May produce softer results

  • Clips longer than 10 seconds (quality drops past that)

  • Multi-character dialogue scenes

  • Fine text or legible signage in the video

  • Rapid scene cuts within a single generation


FAQ

What is MiniMax H3 and how is it different from Hailuo 2? MiniMax H3 is the third generation of MiniMax's Hailuo video line, released July 31, 2026. The biggest change: it generates synchronized stereo audio alongside the video in a single pass, so dialogue, effects, and ambience come out of the same run. It also supports omni-modal reference inputs and outputs 768p at 24 fps with open weights.

What are MiniMax H3 style embeddings? Style embeddings are pre-encoded tensors that inject a visual effect into the prompt. Instead of describing a style in words, you prepend a tag like embedding:minimaxh3_bullet_time and the model applies that look to whatever scene you describe. Ten community embeddings ship with the model, covering effects from slow-motion to fire breath to seasonal time-lapse. You can combine multiple embeddings in one prompt.

Does MiniMax H3 generate audio automatically? Yes. MiniMax H3 generates native stereo audio in the same diffusion pass as the video. The audio includes dialogue, ambient sound, and effects that match the visual content. There is no separate audio model or post-processing step. The workflow decodes both the video and audio latents and packages them into one file.

What resolution and frame rate does this workflow output? The default is 1344x768 (16:9) at 24 frames per second, which is the native resolution for the open-weight model. The resolution selector supports other aspect ratios at 0.9 megapixels. 2K output is available through the MiniMax API only, not through the open weights.

Is MiniMax H3 open source? Can I use the outputs commercially? The weights are released under the MiniMax H3 Community License, not a standard open-source license. Commercial use is permitted for organizations under $20M annual revenue with attribution required. The license restricts deployment in the United States, EU, United Kingdom, and South Korea unless MiniMax grants separate authorization. Review the full license terms before commercial deployment.

What hardware does MiniMax H3 need to run locally? The 33B model is heavy. Running locally requires high-end GPU hardware with large VRAM. Quantized versions (INT8, NVFP4) reduce the footprint. On a cloud platform with pre-loaded models and compatible hardware, you skip the setup entirely.

How to run MiniMax H3 online? You can run MiniMax H3 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, write your prompt, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A director runs a generation and likes the result. A motion artist opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Write a prompt, pick a style embedding, and generate your first video with sound.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N