Seedance 2.0: Reference to Video

Composite up to nine reference images into one video clip with Seedance 2.0 by ByteDance. Tag each image in the prompt, describe the scene, and hit run.

ai video
API
multi reference
reference to video
seedance 2.0

1.8k

Gen time: ~4 min 4 secs

Nodes & Models

Seedance20ReferenceToVideo_floyo
VideoToFrames
LoadImage
VHS_VideoCombine

ABOUT THE WORKFLOW

Composite References Into a Video
Upload reference images, tag each one in the prompt, and describe how they combine. The model reads your tags, takes the subject from one image, the environment from another, the props from a third, and composites them into a single video clip.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • Seedance 2.0 Reference to Video by ByteDance. Released February 2026. Takes up to 9 images, 3 videos, and 3 audio references alongside a tagged prompt, and composites them into a single clip with optional sound. The reference-to-video mode of the Seedance family, built for multi-element scene construction.


HOW IT WORKS

Step 1. Upload your reference images
Three loaders are wired by default. Each one gives the model a visual element to work with. Add more by connecting the open slots on the node, up to nine.
Works great with: characters · objects · environments · props · style references

Step 2. Tag each image in the prompt
Start the prompt with a tag block for each reference: [IMAGE 1: female model with sword], [IMAGE 2: glowing red fish], [IMAGE 3: rainy lantern alley]. Then describe the scene, the camera, and the motion below.

Step 3. Set resolution and length
Pick 480p or higher and a duration in seconds. Both settings drive what a run costs.

Step 4. Hit run and download
The model reads the tags, composites the elements, and builds the clip. The result is saved as an MP4 under Seedance2_ref2vid/vid.
Ready for: Premiere · DaVinci Resolve · CapCut · After Effects

First time? Leave every setting as-is. The defaults (480p · 4 seconds · 16:9 · audio off · random seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard composite (most people) — 480p · 4 seconds · 16:9 · audio off · random seed. The right starting point for almost everyone.

  • Need higher resolution — Raise to 720p or above. Resolution and duration together drive cost, so raise one at a time.

  • Need a longer clip — Raise the duration. More time gives the model room to develop the motion and the compositing.

  • Want sound on the clip — Turn audio on. The model writes sound in the same pass as the picture. Describe the sound in the prompt, like "rain on cobblestones, distant thunder."

  • Elements are blending into each other — Sharpen the tags. Name exactly what each image supplies and where it appears in the scene. "IMAGE 1 is the character in the centre" is clearer than "IMAGE 1 is a woman."

  • Fine detail is lost — Fewer references give each one more room. Drop to two images when detail matters more than complexity.

  • Repeat a result you liked — Set a fixed seed number instead of leaving it on random. The same images, prompt, and seed give you the same clip back.

  • Want to add video or audio references — The node has slots for up to 3 video references and 3 audio references alongside the 9 image slots. Connect a video loader or an audio loader to use them.

Prompt: Tag first, describe second. Start with [IMAGE 1: what it is] for every reference, then write the scene below the tags. "The female model from IMAGE 1 stands in the centre of the alley from IMAGE 3, the fish from IMAGE 2 float around her, camera slowly pushes in, rain falls." Giving each image one job and naming where it goes in the frame produces tighter composites than leaving the model to guess.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎭 Character-in-Environment
Place a character from one image into a setting from another with lighting and perspective matched.

🛍️ Product in Scene
Composite a product reference into a lifestyle environment for ads and social content without a photo shoot.

🎬 Multi-Element VFX Shots
Combine a subject, a prop, and a background into one shot where everything moves together.

🎨 Mood Board to Motion
Turn a set of mood board stills into a moving composite that pitches the look and the feeling of a scene.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear tags that name what each image supplies

  • Two to three references with distinct roles (subject, environment, prop)

  • Prompts that say where each element goes in the frame

  • Clips of 4 to 8 seconds to start

⚠️ May produce softer results

  • Many references competing for attention in a short clip

  • Vague tags that leave the model guessing which element goes where

  • References with clashing styles, lighting, or perspective

  • Fine facial detail when several characters share the frame


FAQ

What is Seedance 2.0 Reference to Video?
Seedance 2.0 Reference to Video is the multi-reference mode of ByteDance's Seedance family, released February 2026. It takes up to 9 reference images, 3 video clips, and 3 audio files alongside a tagged prompt, reads the tags, and composites the elements into a single video clip with optional sound. It is built for shots that combine distinct visual elements from separate sources.

How do the image tags work?
Tag each reference at the top of the prompt in square brackets: [IMAGE 1: description]. The number matches the order of the connected image loaders. The model reads the tags and binds each image to the role you described. Without tags, it guesses which element goes where.

How many references can I use?
Up to 9 images, 3 videos, and 3 audio files in a single generation. Three image loaders are wired by default. Connect more by linking additional loaders to the open slots on the node. More references give the model more elements to work with, but compositing gets harder as the count rises.

Does Seedance 2.0 generate audio?
Yes, when you turn audio on. Sound is generated in the same pass as the picture. It ships off by default on this workflow. Describe the sound in the prompt and turn the setting on to hear it.

What is the difference between Seedance 2.0 and Seedance 2.5?
Seedance 2.5 extends clip length from 15 to 30 seconds, adds support for up to 50 multimodal references, and introduces region-level editing. Seedance 2.0 is the earlier model with reference-to-video compositing at up to 9 images. This workflow uses the 2.0 endpoint for reference-to-video work.

Is Seedance 2.0 open source?
No. Seedance 2.0 is a closed model with no published weights. It is reached through the ByteDance Seed API.

How to run Seedance 2.0 reference to video online?
You can run Seedance 2.0 reference to video online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your references, tag them in the prompt, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload your reference images, tag them in the prompt, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N