
COMMUNITY PAGE
Run SCAIL 2 on Floyo
Home / Model / SCAIL 2 on Floyo
AI CHARACTER ANIMATION
Run SCAIL 2 on Floyo
End-to-end character animation without skeleton extraction. Transfer motion from any driving video to any reference character in a single pass. 14B DiT on Wan 2.1 backbone. Handles humans, animals, illustrations, and multi-character scenes. Apache 2.0.
Run Zhipu AI's SCAIL 2 through ComfyUI in your browser. No API key, no installs, no local GPU.
|
Parameters 14B (Wan 2.1 backbone) |
Modes 4 (animate, replace, multi, zero-shot) |
|
Resolution 512p / 704p |
Skeleton Required No (end-to-end) |
No installation. Runs in browser. Updated July 2026.
character animation
character replacement
motion transfer
scail 2
video to video
wan 2.1
Transfer the motion from any video onto your own character with SCAIL 2, Z.ai's end-to-end character animation model built on Wan 2.1. Upload a video and a character photo, then hit run.
Wan2.1 + SCAIL-2 for Character Motion Transfer
Transfer the motion from any video onto your own character with SCAIL 2, Z.ai's end-to-end character animation model built on Wan 2.1. Upload a video and a character photo, then hit run.
What you get?
SCAIL 2 is an end-to-end character animation model from Zhipu AI and Tsinghua University, released June 9, 2026 under Apache 2.0. Built on the Wan 2.1 I2V 14B DiT backbone. It transfers motion from any driving video to any reference character without skeleton extraction, pose estimation, or intermediate representations. Four modes in one model: single character animation, cross-identity replacement, multi-character scenes, and zero-shot generalization to animals and nonstandard figures. Trained on MotionPair-60K (60K synthesized motion pairs). Available as a ComfyUI node on Floyo.
SCAIL 2 WORKFLOWS ON FLOYO









What are SCAIL 2's technical specifications?
SCAIL 2 is a 14B parameter DiT built on the Wan 2.1 I2V backbone with a 3-segment RoPE design (reference, video, pose). Trained on MotionPair-60K (60K synthesized motion pairs). Four modes: animation, cross-identity replacement, multi-character, and zero-shot. Output at 512p and 704p (dimensions divisible by 32). Uses in-context mask conditioning and mode-specific RoPE. Includes a DPO LoRA for detail improvement. Apache 2.0 licensed.
| Spec | Details |
|---|---|
| Developer | Zhipu AI (Z.ai) / Tsinghua University |
| Backbone | Wan 2.1 I2V 14B DiT (modified with 3-segment RoPE) |
| Parameters | 14 billion |
| Training Data | MotionPair-60K (60K synthesized motion pairs from SCAIL-Preview, Wan-Animate, MoCha) |
| Conditioning | In-context mask conditioning + mode-specific RoPE (no skeleton/pose intermediates) |
| Resolution | 512p and 704p (dimensions divisible by 32) |
| Modes | Animation, Cross-Identity Replacement, Multi-Character, Zero-Shot (animals/nonstandard) |
| DPO LoRA | Yes (Bias-Aware DPO for detail improvement, released on HuggingFace) |
| Zero-Shot Support | SAM3D-Body mesh rendering (never seen in training, works zero-shot) |
| Text Encoder | UMT5-XXL |
| VAE | Wan 2.1 VAE |
| VRAM (Full) | 32GB+ (fp8 scaled) / 16GB+ (fp16 with offloading) |
| VRAM (GGUF) | 6-17GB (Q2 lowest, Q8 highest quality) |
| License | Apache 2.0 (full commercial rights) |
| Paper | arXiv 2606.10804 (CVPR 2026 Findings: SCAIL v1) |
| ComfyUI Access | Native support on Floyo (1 workflow) |
| Release Date | June 9, 2026 |
What can you create with SCAIL 2?
SCAIL 2 covers character animation, cross-identity replacement, multi-character group motion, animal and creature animation, dance transfer, gesture and performance transfer, mascot animation, and game character motion. All four modes run from one model with one workflow. No mode switching, no separate pipelines, no skeleton preprocessing.
| Capability | What It Does | Use Case |
|---|---|---|
| Character Animation | Upload a reference character image and a driving video. The model animates the character following the motion from the video. Background is preserved from the reference. | Social content, mascot videos, character promos |
| Cross-Identity Replacement | Replace a character in a driving video with a different reference identity. The motion, timing, and composition transfer while the appearance changes completely. | Talent swaps, brand character insertion, localization |
| Multi-Character Scenes | Animate multiple characters in the same scene with interaction, occlusion handling, and group motion. Each character can have a separate reference identity. | Group performances, ensemble animations, team content |
| Animal and Creature Motion | Zero-shot generalization to nonstandard figures. Drive animal characters, pets, creatures, and fantasy characters from human or animal driving video. No human skeleton dependency. | Pet content, creature animation, fantasy characters |
| Motion Transfer | Transfer dance, gesture, and performance timing from real footage into any character asset. The model preserves timing and energy while changing the performer. | Dance videos, TikTok content, performance capture |
| Pipeline Integration | Chain with image models in ComfyUI. Generate a character with Nano Banana or Ideogram V4, animate with SCAIL 2, add voiceover with Fish Audio S2. All in one workflow. | Concept-to-animation pipelines |
How does SCAIL 2 compare to other character animation models?
SCAIL 2 leads on skeleton-free animation, multi-character scenes, and animal/creature support. Moonvalley Marey leads on commercially safe training data and pose transfer fidelity. Kling Omni leads on 4K resolution and native audio. Wan-Animate leads on ecosystem integration. SCAIL 2's edge: no skeleton bottleneck means it works for any character type, any body shape, and any number of characters in a scene.
| Model | Skeleton-Free | Animals | Multi-Character | License |
|---|---|---|---|---|
| SCAIL 2 | Yes (end-to-end) | Yes (zero-shot) | Yes (native) | Apache 2.0 |
| Moonvalley Marey | No (pose transfer) | No (humans only) | No | Commercial API |
| Wan-Animate | No (skeleton-based) | Limited | Via tracking | Apache 2.0 |
| MimicMotion | No (pose-driven) | No | No | Open source |
Source: SCAIL 2 arXiv paper (2606.10804), HuggingFace model card (zai-org/SCAIL-2), GitHub (zai-org/SCAIL-2), AI FILMS Studio coverage, and third-party ComfyUI workflow documentation as of July 2026.
What are SCAIL 2's key features?
SCAIL 2's feature set centers on removing the skeleton bottleneck from character animation. Every other feature follows from that design choice. Animals work because there is no human skeleton to violate. Multi-character works because there are no overlapping skeletons to misinterpret. Replacement works because the model reads visual identity directly, not through a skeleton intermediary.
End-to-End (No Skeleton)
SCAIL 2 concatenates driving video latents directly to the generation sequence. The model reads movement, appearance, environment, and fine visual details from the raw video. No skeleton estimation, no pose extraction, no intermediate representation at any stage. This is the core architectural difference from every other character animation model. Information that gets lost in a skeleton bottleneck (subtle head tilts, fabric motion, weight shifting, finger positioning) is preserved.
Four Modes in One Model
Single character animation, cross-identity replacement, multi-character scenes, and zero-shot animal/creature driving all run from the same model weights. No mode switching, no separate checkpoints, no different pipelines. The in-context mask conditioning and mode-specific RoPE handle the routing internally. Upload your inputs, and the model figures out what you are asking for.
Cross-Identity Replacement
Replace a character in a driving video with a completely different reference identity. The motion, timing, and scene composition transfer while the appearance changes. This is the Runway-style character replacement workflow, but open-source under Apache 2.0 and running through ComfyUI. Use it for talent swaps, brand character insertion, or creating localized variants of the same ad with different performers.
Multi-Character Group Motion
Animate multiple characters in the same scene. The model handles interaction, occlusion (one character passing behind another), and group motion coordination. Conventional models fail here because overlapping skeleton maps from multiple characters create ambiguous keypoint assignments. SCAIL 2 does not use skeletons, so there is nothing to overlap.
Animal and Creature Zero-Shot
Drive animal characters, pets, creatures, and fantasy figures from human or animal driving video. This works zero-shot because the model never depended on human skeleton anatomy in the first place. A cat, a robot, a slime creature, and a four-legged dragon all process through the same pipeline that handles human characters. No separate animal model needed.
Bias-Aware DPO LoRA
SCAIL 2 includes a DPO (Direct Preference Optimization) LoRA that corrects for bias patterns inherited from the pose-driven teacher models used during training. The DPO LoRA improves fine detail in hands, faces, and fabric. It is released separately on HuggingFace and can be enabled in both the standalone inference code and ComfyUI implementations.
MotionPair-60K Training Recipe
The model was trained on 60K synthesized motion pairs built from SCAIL-Preview, Wan-Animate, and MoCha. A "reverse driving" technique teaches the model to infer motion beyond what its teacher models could produce. This is why SCAIL 2 shows emergent capabilities (animal driving, zero-shot mesh rendering support) that none of its training sources had individually.
How does SCAIL 2 work?
SCAIL 2 is a 14B parameter DiT built on the Wan 2.1 I2V backbone. The key modification is a 3-segment RoPE design that handles reference image tokens, driving video tokens, and pose signal tokens in separate positional encoding segments. Instead of extracting skeletons, the model directly concatenates driving video latents into the generation sequence so it can read all visual information from the raw input.
In-context mask conditioning tells the model what to animate and what to preserve. Black mask regions indicate hidden background. White regions indicate visible background. Colored regions encode correspondence between character body parts and driving motion. This masking system is what makes multi-character and replacement modes work: you can mask different characters independently and assign each one a separate reference identity.
The training pipeline synthesized 60K motion pairs using three teacher models (SCAIL-Preview, Wan-Animate, MoCha). The "reverse driving" trick is critical: the model was trained to infer the original motion from generated output, forcing it to learn bidirectional motion understanding. This is why it develops emergent capabilities (animal driving, SAM3D-Body mesh support) that the teacher models did not have.
On Floyo, SCAIL 2 runs through ComfyUI nodes on H100 NVL GPUs. Upload your reference character image and driving video. The workflow handles mask generation, model loading, and inference. The output is a video of your character performing the driving motion. You can chain it with image models (generate the character first) or audio models (add voiceover after).
Fair warning: SCAIL 2 is a character animation model, not a general video generator. It expects a reference character image and a driving video as inputs. The mask input is critical even in animation mode. Output resolution is 512p or 704p (not 1080p). The model can struggle with extreme perspective changes and very fast rotations. For raw video generation from text, use Wan 2.7, Vidu Q3, or HappyHorse 1.0 instead.
Frequently Asked Questions
Common questions about running SCAIL 2 on Floyo.
You can start with Floyo's free pricing plan. To continue using the service beyond the free tier, upgrade your Floyo pricing plan. SCAIL 2 is open-source under Apache 2.0, so there is no additional API cost beyond your Floyo plan.
Open Floyo in your browser, search "SCAIL" in the template library, and open the "Wan 2.1 SCAIL 2 for Character Motion" workflow. Click Run, upload your reference character and driving video, and generate. Floyo handles the H100 GPU, ComfyUI environment, and 14B model weights. No local install, no Python setup.
Zhipu AI (Z.ai) and Tsinghua University. The team includes Wenhao Yan, Fengjia Guo, Zhuoyi Yang, and Jie Tang. SCAIL 2 was released June 9, 2026. The original SCAIL (v1) was accepted to CVPR 2026 Findings Track. Weights are on HuggingFace (zai-org/SCAIL-2). Code is on GitHub (zai-org/SCAIL-2). Apache 2.0 license.
No skeleton extraction. Most character animation models extract skeleton keypoints from a reference, then map motion onto those keypoints. That two-stage pipeline breaks down with complex motion, non-human characters, and multi-character scenes. SCAIL 2 bypasses it entirely by concatenating driving video latents directly into the model. The result: one model handles humans, animals, creatures, groups, and replacement with no pipeline changes.
Yes. This is a zero-shot capability. The model never depended on human skeleton anatomy, so it generalizes to animals, pets, creatures, and fantasy figures without a separate animal model. Drive a cat character from human walking video, or a robot from dance footage. The motion transfers to whatever body shape your reference character has.
Yes. Floyo runs ComfyUI, which lets you chain multiple models. Generate a character with Nano Banana, Ideogram V4, or Z-Image Turbo, animate with SCAIL 2, add voiceover with Fish Audio S2 or ElevenLabs, upscale with Topaz. All in one pipeline, all in your browser.
Yes. SCAIL 2 is released under the Apache 2.0 license, which grants full commercial usage rights. You can use the generated animation in products, marketing, client work, and any other commercial context.
512p and 704p with dimensions divisible by 32. This is not 1080p. For higher resolution output, chain SCAIL 2 with Topaz Video AI on Floyo to upscale after generation. The model prioritizes motion fidelity and character consistency over raw pixel count.
Try SCAIL 2 on Floyo
End-to-end character animation without skeleton extraction. 14B DiT on Wan 2.1 backbone. Humans, animals, multi-character, and replacement in one model. Run it in your browser.
Try SCAIL 2 Now → Browse All ModelsRelated Reading
Film and Animation Workflows on Floyo
Character and Concept Design on Floyo
Last updated: July 2026. Specs from SCAIL 2 arXiv paper (2606.10804), HuggingFace model card (zai-org/SCAIL-2), GitHub (zai-org/SCAIL-2), SCAIL 2 project page (teal024.github.io/SCAIL-2), AI FILMS Studio coverage, GGUF quantization documentation (realrebelai/SCAIL-2_GGUF), and ComfyUI workflow guides.
Run SCAIL 2 through ComfyUI on Floyo. End-to-end character animation without skeleton extraction. Transfer motion from any driving video to humans, animals, creatures, and multi-character scenes in one pass. 14B DiT on Wan 2.1. Apache 2.0. Free to try.
