Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing
Wan Animate VS Wan2.1+Scail-2 VS Minimax H3 hero

COMMUNITY PAGE

Wan Animate VS Wan2.1+Scail-2 VS Minimax H3

Floyo / Pages / AI motion transfer comparison

MODEL COMPARISON

Wan Animate 2 vs Wan 2.1 + SCAIL-2 vs MiniMax H3

Three character motion transfer workflows, three different bets on what matters. One strips the process down to six sampling steps and no masks. One segments the subject first so your character can step into real footage. One generates an entire scene around the motion, with sound.

All three take the same two inputs — one character image, one driving video — and all three run as open-weight ComfyUI workflows on Floyo's H100 GPUs. No API keys, no per-clip fees. The difference is what survives the trip: the pose, the background, or the whole shot.

Here is how they compare on resolution, clip length, audio, pose control, and what each one is actually built for.

3 Workflows Image + video to video ComfyUI workflows on Floyo Open weights October 2026

Fastest to a usable take

Wan Animate 2

Image plus pose video, six LCM steps, no skeleton extraction and no masks. Two chained passes in the workflow push clips past the single-pass limit. Apache 2.0.

Tradeoff: 480p default. Hands and long-sequence identity are the weak spots.

Best scene preservation

Wan 2.1 + SCAIL-2

The only one that segments first. SAM 3.1 tracks the subject in both the footage and your image, so the model replaces the person and leaves the plate intact.

Tradeoff: heaviest model stack. Longest cold start of the three.

Best finished shot

MiniMax H3

720p by default and up to 2K, 15-second clips in one run, and the only native stereo audio in the group. Omni-modal reference-to-video.

Tradeoff: 20 steps instead of 6. The prompt does real work — vague prompts show.

· Specifications

The technical details, side by side

All three run open weights on Floyo's H100 GPUs. Every value below is read from the shipped workflow, not from a spec sheet.

SpecificationWan Animate 2Wan 2.1 + SCAIL-2MiniMax H3
DeveloperAlibaba Tongyi LabAlibaba (Wan 2.1) + Meta SAM 3.1MiniMax (Hailuo 3.0)
ReleasedAugust 7, 2026Wan 2.1 lineage + SCAIL-2July 31, 2026
LicenseOpen-source (Apache 2.0)Open weightsH3 Community License (open weights)
Parameters14B14B + SAM 3.1 + 2 LoRAsPruned int8 + Qwen3-VL 32B encoder
InputsImage + pose videoImage + driving videoImage + driving video + prompt
Prompt requiredOptionalOptional (feeds SAM, not the video model)Required
BackgroundRebuilt from the reference imageOriginal footage preservedGenerated from the prompt
Default resolution832 × 480512 × 8961280 × 720 (model to 2K)
Dimension ruleMultiples of 16 (hard fail otherwise)Floors to multiples of 32Multiples of 32
Clip length81 frames per pass, 2 passes chained65 frames per pass, 5-frame carryUp to 360 frames / 15 s in one run
AudioCopied from driving videoCopied from driving videoGenerated, native stereo
OutputMP4 + side-by-side comparisonMP4 + side-by-side comparisonMP4 with audio + comparison
Floyo creditsFloTime GPU creditsFloTime GPU creditsFloTime GPU credits

None of these are API nodes. All three run open weights locally on Floyo's H100 GPUs and bill FloTime GPU credits, so cost tracks generation time rather than clip count. Every model file is pre-loaded — no downloads, no VRAM ceiling to plan around.

01 · Strengths

Where each model wins

What each workflow does better than the other two, based on architecture and shipped configuration.

Wan Animate 2Wan 2.1 + SCAIL-2MiniMax H3
Iteration speedBest (6 steps at 480p)Good (6 steps + segmentation pass)Slowest (20 steps at 720p)
Pose fidelityBest (direct pose conditioning at 1.0)Strong (mask-guided at 0.75)Good (prompt-mediated)
Background controlRegenerated from imageBest (original plate survives)Fully generated, prompt-directed
Resolution ceiling480p tested defaultUp to 1280 × 704 testedBest (720p default, 2K capable)
Clip durationShort passes, chainedShort passes, 5-frame carryBest (15 s in one run)
AudioPassthrough onlyPassthrough onlyBest (generated stereo)
Setup simplicityBest (no prompt, no masks)SAM prompts to setStructured prompt required
Model load weightLightest (1 model + 1 LoRA)Heaviest (3 models + 2 LoRAs)Heavy (model + 32B encoder + 2 VAEs)

02 · What we measured

Benchmark criteria

Eight things we looked at across every output. These are the questions you ask when reviewing a motion transfer take.

1

Motion accuracy

Does the character hit the same beats as the driving video? Timing, weight shifts, limb extension.

2

Identity hold

Is it still the same face, hair, and outfit at the last frame as at the first?

3

Hands and extremities

Do fingers stay separate and countable, or melt into mittens during fast motion?

4

Temporal consistency

Does the clip flicker, drift, or shift colour between passes and context windows?

5

Cloth and hair physics

Do fabric and hair react to the motion, or stay welded to the body like paint?

6

Background integrity

Does the environment hold still and stay coherent, or warp as the subject moves through it?

7

Edge quality

Where the character meets the background, is the boundary clean or does it halo and smear?

8

Production readiness

Would you cut this into a timeline without a roto pass or a second take?

03 · Benchmark tests

Same clip, all three models

Same character image and same driving video across all three workflows. No per-model tuning. What you see is what each one did with the exact same reference.

Benchmark 1 · Full-body dance

Fast, continuous motion with arm and hand detail

What to look for

Does the character land on the same beats? Do the hands survive the fast sections? Does the face hold from the first frame to the last?

Wan Animate 2

Wan 2.1 + SCAIL-2

MiniMax H3

Benchmark 2 · Character replacement

Swapping a subject into existing footage

What to look for

Does the original background survive untouched? Is the edge where the character meets the scene clean, or does it halo? Does the lighting match the plate?

Wan Animate 2

Wan 2.1 + SCAIL-2

MiniMax H3

Benchmark 3 · Stylised character

Illustrated or 3D-rendered subject, not a photo

What to look for

Does the art style hold, or does the model drift toward photoreal? Do flat colours stay flat? Does the character keep its proportions under human motion?

Wan Animate 2

Wan 2.1 + SCAIL-2

MiniMax H3

Benchmark 4 · Long take

Ten seconds or more of continuous motion

What to look for

Can you see the seam where passes join? Does identity drift across the clip? Does colour or exposure shift between chained segments?

Wan Animate 2

Wan 2.1 + SCAIL-2

MiniMax H3

04 · Results

Output quality

How each workflow performs across the eight criteria, read from architecture and shipped settings.

Wan Animate 2Wan 2.1 + SCAIL-2MiniMax H3
Motion accuracyBest (direct pose conditioning)Strong (mask-guided)Good (prompt-mediated)
Identity holdDrifts on long sequencesBest (reference mask anchors it)Strong within a single pass
Hands and extremitiesKnown weak pointGoodBest (720p gives more pixels to work with)
Temporal consistencyGood (context windows + continue_motion)Good (5-frame carry)Best (no pass seams in a single run)
Background integrityRegenerated, can wanderBest (original pixels untouched)Coherent but fully synthetic
Edge qualityGoodBest (explicit SAM masks)Good (no compositing boundary)
Production readinessPreviz and blockingBest for compositing into a plateBest for standalone delivery

These are capability reads from the workflow architecture and shipped defaults, not scores from a matched render test. The benchmark clips in section 03 are where measured results come from once each row has been reviewed.

05 · Speed & cost

What drives generation time

All three run on Floyo's H100 GPUs and bill FloTime credits, so cost is generation time. These are the factors that set it.

ModelSampling costPre-processingSpeed lever
Wan Animate 26 steps × 2 passes at 832 × 480NoneDrop the second pass for a short take
Wan 2.1 + SCAIL-26 steps at 512 × 896SAM 3.1 tracks every frame, twiceShorter driving clip cuts the tracking pass
MiniMax H320 steps at 1280 × 720, up to 360 framesAudio VAE decode on top of videoTurbo LoRA drops 20 steps to 4

Wan Animate 2 is the cheapest pass by a clear margin — six LCM steps at 480p. SCAIL-2 matches it on sampling but adds a segmentation stage the others don't have, plus the longest cold start from loading three models and two LoRAs. MiniMax H3 is the expensive one on every axis: more steps, higher resolution, more frames, and a second VAE. The Turbo LoRA is the main lever if you want H3's output without H3's bill.

06 · Model by model

The three workflows

What each one is like to run on Floyo, and when to pick it.

01 Wan Animate 2 Apache 2.0 Fastest take

Pose in, character out. The lowest-friction route from a photo to a moving character on Floyo.

Speed4/4
Pose match4/4
Resolution2/4
Audio1/4

Wan Animate 2 released on August 7, 2026 from Alibaba Tongyi Lab as part of the Wan 2.2 video series. It is a 14B model that takes a single character image and a driving video with no skeleton extraction step in between — the pose is read directly from the footage. A distilled lightx2v LoRA ships with the weights, which is what brings sampling down to six LCM steps at cfg 1.

On Floyo, the workflow runs the generation subgraph twice. The first pass hands its last frames and a video frame offset to the second through continue_motion, so the clip extends past the 81-frame single-pass default without an obvious seam. Context windows are set to 21 with 8 overlap and pyramid fusion, the cache runs int8 on GPU, and pose_strength sits at 1.0 with a full 0–100% window. Audio is lifted straight off the driving video and muxed back in, and both the character clip and a side-by-side stitch are saved.

Reach for it when

Dance and performance clips. Previz and blocking where you want several takes cheaply. Stylised characters, since the prompt field describes character and background to hold the look.

Skip it when

You need the original background kept (SCAIL-2 does that) or a finished 720p shot with sound (MiniMax H3 does that). Also skip for close-ups on hands, which are a known weak point.

02 Wan 2.1 + SCAIL-2 Open weights Best scene preservation

The only workflow here that knows who it is replacing. Segmentation first, generation second.

Scene control4/4
Pose match3/4
Speed3/4
Setup effort2/4

SCAIL-2 is a Wan 2.1 14B fine-tune that takes explicit masks as a conditioning input. The Floyo workflow runs SAM 3.1 twice — once over every frame of the driving video, once over your reference image — then feeds both tracks into SCAIL2ColoredMask, which produces the colored mask pair the video model consumes. Two LoRAs stack on top: a DPO LoRA at full strength and the lightx2v distill at 0.8, which is what keeps it to six euler steps.

One boolean changes the whole job. With replacement_mode on, your character steps into the original footage and the background comes through untouched. With it off, the same graph does plain motion transfer. The workflow runs at 512 × 896 with pose_strength 0.75 and a 5-frame carry between passes; the note node lists 512 × 896, 704 × 1280, 896 × 512 and 1280 × 704 as the tested resolutions, and anything else floors to a multiple of 32. Both mask passes preview on canvas, so you can confirm it tracked the right subject before committing.

Reach for it when

Compositing a character into footage you already shot. Vertical and social video, since the default is portrait-first. Multi-subject shots where you need to target one person out of several.

Skip it when

You want the quickest possible take (Wan Animate 2 skips the whole segmentation stage) or you are building a scene that does not exist yet (MiniMax H3 generates one).

03 MiniMax H3 Open weights Best finished shot

Omni-modal reference-to-video. The only one that builds a scene and scores it at the same time.

Resolution4/4
Audio4/4
Clip length4/4
Speed1/4

MiniMax H3, also known as Hailuo 3.0, released on July 31, 2026 as an omni-modal video model: it generates picture and synchronised audio in one pass from text, images, or reference video. Resolution goes to 2K at 24 fps with clips from 4 to 15 seconds. The Floyo workflow uses the reference-to-video pipeline, driven by a Qwen3-VL 32B text encoder and decoding through two VAEs — one for video, one for audio.

The prompt is doing real work here, not decorating. It uses tagged references — <Subject 1> for the character, <Picture 1> for the image, <Video 1> for the motion — which is how the model binds appearance to movement while ignoring the person in the driving clip. Defaults are 1280 × 720 at 20 res_multistep steps with a beta scheduler and no negative conditioning. The driving video is forced to 24 fps and capped at 360 frames, and clip length snaps to a 17-frame latent boundary, so output can run slightly past the input. A Turbo LoRA ships with the workflow but is off by default; switching it on drops 20 steps to 4.

Reach for it when

Finished shots that need to ship with sound. Clips over ten seconds where chained passes would show. Any case where the scene has to be built rather than borrowed.

Skip it when

You are iterating and want several cheap takes (Wan Animate 2 is a fraction of the cost) or you need the original footage preserved (only SCAIL-2 does that).

Final takeaway

These three do not really compete — they sit at different points in a shot's life. Wan Animate 2 is the iteration choice: six steps, no setup, two chained passes, and the lightest load on the GPU. Wan 2.1 + SCAIL-2 is the compositing choice: SAM 3.1 masks let your character step into footage you already shot, and one boolean flips it between replacement and plain motion transfer. MiniMax H3 is the delivery choice: 720p by default, fifteen seconds in one run, and the only generated audio in the group.

The right pick depends on whether your priority is cheap takes, a preserved background, or a finished shot with sound. In a real pipeline you will likely use two of them — block with Wan Animate 2, finish with whichever of the other two matches the delivery.

All three run as open-weight ComfyUI workflows on Floyo's H100 GPUs. Upload a character image and a driving video, hit Run, and compare the takes side by side in your browser.

Frequently asked questions

Which motion transfer workflow should I start with?

Start with Wan Animate 2 if you want a usable take in the fewest moves — six sampling steps, no masks, no prompt required. Start with Wan 2.1 + SCAIL-2 if you need your character dropped into footage you already shot with the background left intact. Start with MiniMax H3 if the output has to be a finished shot: 720p by default, up to fifteen seconds in one run, and generated audio.

Can I run all three on Floyo right now?

Yes. All three are available as ComfyUI workflows on Floyo. Open a workflow in your browser, upload a character image and a driving video, and click Run. Every model file is pre-loaded on Floyo's H100 GPUs, so there is nothing to download. Each workflow returns two MP4 files: the generated clip and a side-by-side comparison against the driving video. No local install and no API key required.

Are these free to use on Floyo?

None of these three are API nodes, so there is no per-clip fee and no API Wallet charge. All three run open weights locally on Floyo's H100 GPUs and bill FloTime GPU credits, which are metered by compute time. That means cost tracks how long a generation runs rather than how many clips you make, so resolution, step count, and clip length are what move the number.

Can I use the generated videos commercially?

Wan Animate 2 is released under Apache 2.0, which permits commercial use with no attribution required. Wan 2.1 + SCAIL-2 runs on open weights, as does the SAM 3.1 segmentation stage. MiniMax H3 ships under the MiniMax H3 Community License, which carries its own terms — check those before commercial release. Separately, you are responsible for rights to the character image and driving footage you feed in.

Which one keeps the original background?

Only Wan 2.1 + SCAIL-2. It runs SAM 3.1 over every frame of the driving video and over your reference image, producing explicit masks that tell the model exactly which pixels to replace. With replacement_mode on, your character steps into the original footage and the plate comes through untouched. Wan Animate 2 rebuilds the background from your reference image, and MiniMax H3 generates a new scene from your prompt.

How long does each one take to generate?

Wan Animate 2 is the cheapest pass — six LCM steps at 832 × 480, run twice to extend the clip. SCAIL-2 matches that sampling cost at 512 × 896 but adds a segmentation stage, and loads three model files plus two LoRAs, giving it the longest cold start. MiniMax H3 is the slowest on every axis: twenty steps instead of six, 1280 × 720 instead of 480p, up to 360 frames, and a second VAE decode for audio. Its Turbo LoRA drops twenty steps to four if you need the output faster.

Do I need to write a prompt?

Only for MiniMax H3, where the prompt does real work. It uses tagged references — <Subject 1> for the character, <Picture 1> for the image, <Video 1> for the motion — and that is how the model binds appearance to movement while ignoring the person in the driving clip. Keep those tags. Wan Animate 2 takes an optional prompt describing character and background. SCAIL-2's two text fields feed SAM rather than the video model, so they name what to track, such as "person".

How long can the output clip be?

MiniMax H3 generates up to 360 frames at 24 fps — fifteen seconds — in a single pass, with no seams to hide. Wan Animate 2 produces 81 frames per pass and the Floyo workflow chains two of them through continue_motion, so length costs you passes rather than hitting a wall. SCAIL-2 runs 65 frames per pass with a five-frame carry between them. One quirk on H3: clip length snaps to a 17-frame latent boundary, so output can run slightly past your driving video.

Which one generates audio?

MiniMax H3 is the only one. It is an omni-modal model that produces picture and synchronised stereo audio in the same pass, decoding through a separate audio VAE. Wan Animate 2 and SCAIL-2 both lift the audio track off your driving video and mux it back into the output, so what you hear is your source footage rather than anything the model created.

What resolution should I use?

Each has its own rule. Wan Animate 2 requires multiples of 16 and fails the run outright on anything else; its tested default is 832 × 480. SCAIL-2 floors to multiples of 32, and the workflow lists 512 × 896, 704 × 1280, 896 × 512 and 1280 × 704 as the tested sizes — the portrait defaults make it the natural pick for vertical video. MiniMax H3 takes multiples of 32, defaults to 1280 × 720, and the model goes up to 2K.

Can I combine these with other AI models in the same workflow?

Yes. Because these run as ComfyUI nodes on Floyo, you can chain them with any other node in the same workflow. Generate your character with Z-Image or Flux and pipe it straight into the reference image input. Pull a driving clip from a video workflow instead of uploading one. Send the output into an upscaler or a reframing workflow before you export. The SCAIL-2 workflow already chains SAM 3.1 segmentation ahead of generation, which is the same pattern.

Explore More On Floyo

Wan2.1 + SCAIL-2 for Character Motion Transfer

character animation

character replacement

LoRAs

motion transfer

scail 2

Video

video to video

wan 2.1

Transfer the motion from any video onto your own character with SCAIL 2, Z.ai's end-to-end character animation model built on Wan 2.1. Upload a video and a character photo, then hit run.

Wan2.1 + SCAIL-2 for Character Motion Transfer

Transfer the motion from any video onto your own character with SCAIL 2, Z.ai's end-to-end character animation model built on Wan 2.1. Upload a video and a character photo, then hit run.

character animation

motion transfer

wan 2.2

wan animate 2

Transfer the motion from any reference video onto a character photo using Wan Animate 2, Alibaba's open-source 14B video model. Upload an image, upload a video, and hit run.

Wan Animate 2 · Motion Transfer R2V

Transfer the motion from any reference video onto a character photo using Wan Animate 2, Alibaba's open-source 14B video model. Upload an image, upload a video, and hit run.

character motion transfer

hailuo3

image to video

minimax h3

video generation

Upload a character image and a driving video. MiniMax H3 maps the motion onto the character and returns a new video with AI-generated stereo audio.

MiniMax H3 · Character Motion Transfer

Upload a character image and a driving video. MiniMax H3 maps the motion onto the character and returns a new video with AI-generated stereo audio.

LTX-2.3 · Motion Transfer (Pose)

Audio

LoRAs

ltx

motion transfer

Video

video to video

Transfer the movement from any reference video onto your portrait with LTX-2.3 by Lightricks, audio and all. Upload a photo and a clip, then hit run.

LTX-2.3 · Motion Transfer (Pose)

Transfer the movement from any reference video onto your portrait with LTX-2.3 by Lightricks, audio and all. Upload a photo and a clip, then hit run.

Kling 3.0 Pro Motion Control · Image to Video
ashree

ashree

237

image to video

kling 3.0

motion control

motion transfer

Video

Upload a character image and a motion reference video, and Kling 3.0 Pro generates a new video of your character performing the referenced action with the option to keep the original audio.

Kling 3.0 Pro Motion Control · Image to Video

Upload a character image and a motion reference video, and Kling 3.0 Pro generates a new video of your character performing the referenced action with the option to keep the original audio.

Kling 3.0 Pro Motion Control

animation

character design

image to video

kling

Video

video generation

Apply motion from a reference video to a still image with Kling 3.0 Pro.

Kling 3.0 Pro Motion Control

Apply motion from a reference video to a still image with Kling 3.0 Pro.

SCAIL-2 bf16 Loop FullLength Motion Transfer

character-animation

full-length

image-to-video

motion-transfer

pose-transfer

scail-2

Video

video-to-video

wan2.1

Transfers the motion of a driving video onto a reference character image

SCAIL-2 bf16 Loop FullLength Motion Transfer

Transfers the motion of a driving video onto a reference character image

SCAIL2_Motion_Control
agi

agi

5.8k

Motion Control

SCAIL2

Video

SCAIL2を使用したモーションコントロール用のワークフローです。動画をモーション元としてアップロードし、入れ替えたい人物画像をアップロードすることで、動画内の動きをキャラクターへ転送できます。緑色のノードだけ設定すればすぐ使えます。 This is an Image-to-Video motion control workflow using SCAIL2.

SCAIL2_Motion_Control

SCAIL2を使用したモーションコントロール用のワークフローです。動画をモーション元としてアップロードし、入れ替えたい人物画像をアップロードすることで、動画内の動きをキャラクターへ転送できます。緑色のノードだけ設定すればすぐ使えます。 This is an Image-to-Video motion control workflow using SCAIL2.

Wan 2.7 Reference to Video with Motion Control

character design

consistency

film production

image to video

Video

video generation

wan

Wan 2.7 Reference to Video with Motion Control

Wan 2.7 Reference to Video with Motion Control

Wan 2.7 Reference to Video with Motion Control

Ready to test them yourself?

The best model for your workflow depends on your own images, prompts and production requirements.

Start creating on Floyo >
TABLE OF CONTENTS
OVERVIEW

Compare three powerful open-weight character motion transfer workflows — **Wan Animate 2, Wan 2.1 + SCAIL-2, and MiniMax H3**. See how they differ in motion accuracy, identity preservation, background control, resolution, speed, clip length, and audio to find the right model for your workflow.