MiniMax Music 3 for Text to Music
Generate complete songs with vocals and full instrumentation using MiniMax Music 3. Write a style caption, add lyrics, and hit run. Up to five minutes.
Audio
Minimax Music 3
Music
Text to Music
123
Nodes & Models
CLIPLoader
minimax_music3_text_encoder_pruned_int8_convrot.safetensors
UNETLoader
minimax_music3_dit_fp16.safetensors
VAELoader
minimax_music3_dav.safetensors
KSampler
ConditioningZeroOut
VAEDecodeAudio
PreviewAudio
MiniMaxMusic3TextEncode
EmptyMiniMaxMusic3LatentAudio
SeedNode
VAEDecodeAudioTiled
ComfySwitchNode
ABOUT THE WORKFLOW
Generate a Full Song from Text Write a caption describing genre, tempo, key, and instruments. Add lyrics with section tags like [verse] and [chorus]. The model composes, arranges, and renders the full track with vocals in one pass.
Model
MiniMax Music 3 by MiniMax. An open-weights music generation model that produces complete songs with vocals, instrumentation, and song structure up to five minutes long.
HOW IT WORKS
Step 1. Write a caption Describe the music you want: genre, BPM, key, instruments, and vocal style. The more structured the caption, the more precise the result. Works great with: pop · rock · EDM · R&B · ambient
Step 2. Add lyrics (optional) Write your song text with section tags on their own lines: [verse], [chorus], [pre-chorus], [bridge], [outro]. Leave this empty to generate an instrumental track.
Step 3. Hit run and preview The model generates the full track in one pass. Audio plays back in the workflow for preview. Ready for: video editors · DAWs · podcasts · any audio player
First time? Leave every setting as-is. The defaults (60 seconds, random seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard song (most people) — 60 seconds, random seed, tiled decode off. The right starting point for almost everyone.
Full-length track — Set duration up to 300 seconds (five minutes). The model holds structure, vocal identity, and arrangement across the full length.
Quick test before committing — Set duration to 15 or 30 seconds. Hear the vibe before generating a full track.
Instrumental only — Leave the lyrics field empty. The model produces a full arrangement without vocals.
Reproduce a result you liked — Lock the seed to the number from your previous run. Same caption and lyrics with the same seed returns the same track.
Running on a lower-VRAM GPU — Turn tiled decode on. It splits the decoding step to reduce memory use at the cost of some speed.
The track sounds vague or unfocused — Write a structured caption: genre first, then BPM, key, instruments section by section, then vocal character. Vague prompts produce vague tracks.
Prompt: The caption is the most important input. Structure it as a music brief: genre, BPM, key, then instruments and vocal direction. "Progressive house, 124 BPM, G minor. Breathy male tenor, pulsing synth chords, deep sub-bass" works. "Make a cool song" does not.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Video Creators & Filmmakers Generate background music or full songs for YouTube videos, short films, and social content without licensing a track.
🎵 Songwriters & Musicians Demo a song idea fast. Write lyrics and a style description, hear the arrangement, and decide what to take into a full production.
🎙️ Podcasters & Content Producers Create intro music, outro stings, or segment transitions that match a specific mood and tempo.
🎮 Game Developers Produce background music for levels, menus, or trailers. Set the genre and mood in the caption and get a track ready to drop into an engine.
📢 Marketing & Brand Teams Generate on-brand music for ads, product videos, or presentations without hiring a composer or clearing a license.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
English-language vocals with clear lyrics
Pop, rock, EDM, R&B, and ambient genres
Structured captions with BPM, key, and instruments
Lyrics tagged with [verse], [chorus], [bridge] sections
⚠️ May produce softer results
Non-English vocals and non-Western genres
Vague captions without tempo, key, or instrumentation
Tracks longer than three minutes (quality can drift)
Complex instrumental textures that need fine layering
FAQ
What is MiniMax Music 3? MiniMax Music 3 is an open-weights music generation model by MiniMax, the lab behind Hailuo AI. It takes a text caption describing genre, tempo, and instruments, plus optional tagged lyrics, and generates complete songs with vocals and full arrangement up to five minutes long. Output is 32 kHz, 16-bit stereo audio.
How does MiniMax Music 3 compare to Suno? Suno is a closed, cloud-only service with broader genre support and more polished instrumental richness on complex tracks. MiniMax Music 3 is open-weights, which means it can run locally or on platforms like this one without a Suno subscription. It handles English-language pop, rock, EDM, and R&B well, with strong vocal timing and lyric accuracy. The tradeoff: non-Western genres and dense instrumental layering are weaker than Suno's top tier.
Can I generate instrumental music without vocals? Yes. Leave the lyrics field empty and write a caption describing the instruments, genre, tempo, and mood. The model generates a full arrangement without vocals.
What audio format and quality does the output use? The output is 32 kHz, 16-bit stereo audio. That is below CD-quality 44.1 kHz, so if your project needs higher sample rates, plan to upsample the output in a DAW or audio editor.
Is MiniMax Music 3 output licensed for commercial use? The model is released under the MiniMax-Music3 Community License. Commercial use is allowed with attribution: you must display "MiniMax-Music3" in your product interface. Organizations with more than $20 million in annual revenue need a separate written agreement with MiniMax. For most creators and small studios, the license is effectively free for commercial work.
Does MiniMax Music 3 work for non-English songs? The model supports multilingual vocals, with published demos in English and Mandarin. Other languages may work but are not officially documented. English-language output is the strongest, and non-Western musical genres are a known weak point. Test with your target language before committing to a production workflow.
How to run MiniMax Music 3 online? You can run MiniMax Music 3 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, write your caption, add lyrics, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A producer generates a track and likes the result. A teammate opens that exact run from shared history and iterates on the caption. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Write a caption, add your lyrics, and generate your first track.
Questions? Watch the free course or check the FAQ above.
Read more
_1787749481978.png?width=1400&height=620&quality=80&resize=contain)
_1787749481978.png?width=104&height=104&quality=80&resize=cover)




