Sonilo v1.1 for Video Sound Effects
Upload a video and Sonilo v1.1 generates synchronized sound effects matched to the on-screen action. Describe the sounds you want or let the model read the scene on its own, then hit run.
Audio
SFX
Sonilo v1.1
V2V
100
Nodes & Models
SoniloV11VideoToVideoSoundEffects_floyo
VideoToFrames
LoadVideo
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Add Sound Effects to Video
Upload a video and get it back with synchronized sound effects mixed in. Describe the sounds you want to hear, or leave the prompt empty and let the model watch the footage and generate effects on its own. For longer clips with distinct scenes, use timed segment prompts to place different sounds at different points.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
Sonilo v1.1 by Sonilo. A proprietary generative audio model strong at reading on-screen actions and producing synchronized, royalty-free sound effects with accurate timing.
HOW IT WORKS
Step 1. Upload your video
The clip you want sound effects added to. Upload an mp4, mov, webm, or gif.
Works great with: AI-generated clips · gameplay captures · silent footage · trailers
Step 2. Describe the sounds (optional)
Write what you want to hear. "Footsteps on gravel, a car door slamming, birds in the background" is better than "add sounds." Leave the prompt empty and the model reads the scene on its own.
Step 3. Set segment prompts (optional)
For clips with distinct scenes, use up to five timed segment slots. Set an end time and a prompt for each one to place different sounds at different points in the video.
Step 4. Hit run and download
Sonilo watches the footage, generates effects matched to visible actions, mixes them in, and returns a finished clip with audio baked in.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any NLE
First time? Leave every setting as-is. The defaults (wav · no segments) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard use (most people) — wav · prompt empty · no segments. Upload the video, hit run, and let the model read the scene.
Want specific sounds — Write a prompt describing each sound and when it happens. "Glass breaking, then footsteps running on pavement" gets closer results than leaving the prompt empty.
Long clip with distinct scenes — Use segment prompts. Set an end time for each segment and write what that section should sound like. One segment per scene is cleaner than one long prompt.
Need a specific audio format — Switch the audio format from wav to mp3, flac, aac, or ogg to match your editing pipeline.
Effects are not matching the action — Add a prompt. The model is more accurate with a hint about what to listen for. Name the objects and the motion.
Static or ambiguous footage — The model works best when there is clear on-screen motion. For still or slow scenes, a prompt helps fill in what the model cannot infer from the visuals alone.
Prompt: Describe the sounds you want, not the visuals. "Rain on a tin roof, distant thunder, wind through a screen door" gives the model clear targets. Vague prompts like "make it sound real" do not help. For most clips, leaving the prompt empty and letting the model read the scene works well on its own.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Filmmakers & Editors
Add foley to silent AI-generated footage or rough cuts without opening a DAW or searching stock libraries.
🎮 Game Developers
Generate synchronized sound effects for gameplay captures, trailers, and demo reels. Upload the clip, describe the audio, and get a finished video back.
📱 Social & Marketing Teams
Turn silent screen recordings, product demos, and short-form content into polished clips with matched sound effects in seconds.
🎥 VFX & Motion Design
Layer realistic sound effects onto animated sequences, composites, and pre-vis footage where recording live audio is not an option.
🎓 Course Creators & Educators
Add contextual audio to tutorial videos, walkthroughs, and explainer content to make them more engaging without manual sound editing.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Video with clear, visible on-screen actions
Clips with distinct motion (doors closing, footsteps, impacts)
Scenes where the prompt names specific sounds
Segmented clips with one scene per segment slot
⚠️ May produce softer results
Static or slow-moving footage with no clear action
Very long clips without segment prompts to guide the model
Ambiguous scenes where multiple interpretations are possible
Footage where the desired sound has no visual cue on screen
FAQ
What is Sonilo v1.1 and what does it do?
Sonilo v1.1 is a proprietary generative audio model made by Sonilo, a San Francisco-based AI audio startup. It watches a video, identifies on-screen actions, and generates synchronized sound effects timed to what is happening in the frame. The output is a finished video with audio mixed in, not a separate audio file you have to sync yourself.
How does AI sound effect generation work compared to stock sound libraries?
Stock libraries require you to search for a matching clip, trim it, and manually align it to the video timeline. Sonilo reads the footage directly, generates effects that match the timing and motion on screen, and returns one finished track. There is no searching, trimming, or manual syncing involved. For clips with distinct scenes, segment prompts give you per-scene control over what the model generates.
Can I control which sounds appear at different points in the video?
Yes. The workflow supports up to five timed segment prompts. Set an end time and a prompt for each segment to place different sounds at different points. For example, segment 1 could cover footsteps for the first three seconds, and segment 2 could cover a door slam from three to five seconds. Segments with an end time of zero are skipped.
Are the generated sound effects royalty-free for commercial use?
Yes. Sonilo outputs are royalty-free and cleared for commercial use. You can use them in shipped products, client work, broadcast, trailers, and published content without additional licensing fees or attribution requirements.
Does Sonilo generate music or speech, or only sound effects?
This workflow generates sound effects only. It does not produce music, dialogue, or voiceover. Sonilo has a separate video-to-music model for soundtrack generation. Use this workflow when you need foley, ambient sounds, and action-synced effects.
What video formats does Sonilo v1.1 accept?
The model accepts mp4, mov, webm, and gif. The generated audio matches the length of your input video automatically. Run cost scales with input video length, so longer clips cost more per run.
How to run Sonilo v1.1 sound effects online?
You can run Sonilo v1.1 sound effects online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your video, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Upload a video and hit run. The model reads the scene and adds sound effects on its own.
Questions? Watch the free course or check the FAQ above.
Read more




