Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

MMAudio V2: Add Sound Effects to Any Video

Generate synchronized sound effects for any video using MMAudio V2, the open-source video-to-audio model from Sony AI and University of Illinois. Upload a video, describe the sounds, and hit run.

263

Gen time: ~2 min 12 secs

Nodes & Models

MMAudioV2VideoToVideo_floyo
VideoToFrames
LoadVideo
CreateVideo
SaveVideo
FloyoStickyNote

ABOUT THE WORKFLOW

Add Sound to a Video
Upload a silent video and describe the sounds you want. MMAudio V2 watches the visual content, matches the timing, and generates a synchronized audio track laid over the original footage. You get back a video with sound.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • MMAudio V2 by University of Illinois, Sony AI, and Sony Group. An open-source video-to-audio model that generates sound effects synced to what happens on screen. Published at CVPR 2025.


HOW IT WORKS

Step 1. Upload your video
The video you want to add sound to. Works with silent footage or AI-generated clips that need a soundtrack.
Works great with: AI-generated video · silent footage · screen recordings · game captures

Step 2. Write a prompt
Describe the sounds you want. Be specific about what you hear, not what you see. "Footsteps on gravel, distant traffic, wind through trees" is better than "a person walking outside."

Step 3. Hit run and download
MMAudio V2 analyzes the motion and timing in your video, generates matching audio, and returns the video with sound baked in. Download the result or keep iterating.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any video editor

First time? Leave every setting as-is. The defaults (25 steps, 8 seconds, CFG 4.5) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard use (most people) — 25 steps · 8 seconds · CFG 4.5 · random seed. The right starting point for almost everyone.

  • Quick test before committing — Lower the steps to 15. Faster generation, slightly less detail in the audio.

  • Longer video clip — Increase the duration to match your clip length. The model trains on 8-second clips, so results are strongest near that length.

  • More control over the sound — Raise CFG strength above 4.5 to make the model follow your prompt more closely. Lower it to give the model more freedom to interpret the visuals.

  • Unwanted sounds in the output — Write a negative prompt describing what to exclude, like "music, speech, static noise."

  • Reproduce a result you liked — Lock the seed to the number from your previous run. Same seed plus same prompt gives the same output.

Prompt: Describe the sounds, not the scene. "Rain hitting a tin roof, distant thunder, dripping water" gives better results than "a rainy day." Name materials, surfaces, and distances. Two to three specific sound layers are the sweet spot.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 AI Video Post-Production
Generate matching sound effects for AI-generated video clips from tools like Wan, Kling, or Runway, so they ship with audio instead of silence.

🎮 Game Development
Create environment audio, impact sounds, and ambient layers for gameplay footage without recording or licensing stock libraries.

🎥 Filmmakers & Editors
Add foley and atmosphere to rough cuts, behind-the-scenes footage, or silent B-roll before final sound design.

📱 Content Creators
Score short-form video with synchronized sound effects that match the on-screen action, without digging through royalty-free libraries.

🥽 VR & Immersive Media
Generate spatial audio layers for 360 video or VR environments where traditional recording is impractical.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-lit video with distinct motion

  • Environment and ambient sound prompts

  • Sound effects tied to visible actions (footsteps, impacts, water)

  • AI-generated video that needs a soundtrack

⚠️ May produce softer results

  • Realistic human speech or singing

  • Music generation or melodic audio

  • Very long clips far beyond 8 seconds

  • Dark or visually ambiguous footage with no clear motion


FAQ

What is MMAudio V2 and how does it work?
MMAudio V2 is a video-to-audio model that generates synchronized sound effects from video content. It was developed by the University of Illinois, Sony AI, and Sony Group, and published at CVPR 2025. The model extracts visual features from each frame, reads motion and timing cues, then synthesizes audio that matches what happens on screen. You guide the output with a text prompt describing the sounds you want.

Can MMAudio V2 generate speech or music?
MMAudio V2 is trained primarily on sound effects, ambient audio, and environmental sounds. It can produce speech-like sounds and rough musical textures, but it is not designed for clean dialogue or composed music. For speech, use a dedicated text-to-speech model. For music, use a music generation model.

How long can the input video be?
The model trains on 8-second clips, so audio quality is strongest at that length. You can set the duration higher, but the further you go from 8 seconds, the more quality may drop. For longer videos, consider splitting into segments and running each one separately.

What does the negative prompt do?
The negative prompt tells the model what sounds to avoid. If your output has unwanted music, static, or crowd noise, add those terms to the negative prompt. Leave it empty for a first run and add exclusions only if something unwanted appears.

Is MMAudio V2 open source, and can I use the output commercially?
MMAudio V2 is open source, but the model checkpoints are released under the CC-BY-NC 4.0 license (non-commercial). The authors note they do not guarantee the pre-trained models are suitable for commercial use. If you need commercial rights, review the license terms on the official GitHub repository before shipping.

What is the difference between MMAudio V2 and other video-to-audio tools?
Most video-to-audio tools either generate generic ambient sound or require manual sound design. MMAudio V2 reads the visual content frame by frame and syncs the generated audio to on-screen motion, so impacts, footsteps, and environmental changes land on the right beat. The text prompt gives you control over what type of sound is generated, while the model handles the timing.

How to run MMAudio V2 online?
You can run MMAudio V2 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your video, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a video, describe the sounds you want, and hit run.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N
FloYo: MMAudio V2: Add Sound Effects to Any Video