Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Wan 2.1 FusionX MultiTalk · Image to Video

Upload a portrait and an audio file, and Wan 2.1 FusionX with MultiTalk generates a video of the character speaking or singing with precise lip-sync driven by your audio.

52

Generates in about 12 mins 39 secs

Nodes & Models

LoadAudio
WanVideoBlockSwap
LoadImage
WanVideoTorchCompileSettings
WanVideoTeaCache
WanVideoEnhanceAVideo
Label (rgthree)
WanVideoTextEncodeSingle
DownloadAndLoadWav2VecModel
LoadWanVideoT5TextEncoder
WanVideoVAELoader
CLIPVisionLoader
WanVideoLoraSelect
MultiTalkModelLoader
WanVideoModelLoader
MultiTalkWav2VecEmbeds
WanVideoSampler
WanVideoDecode
AudioCrop
AudioSeparation
ImageResizeKJv2
WanVideoImageToVideoMultiTalk
WanVideoClipVisionEncode
WanVideoApplyNAG
FloyoStickyNote
VHS_VideoCombine

ABOUT THE WORKFLOW

Animate a Face with Audio-Driven Lip-Sync
Upload a portrait image and an audio file of speech or singing. The workflow extracts vocal features from the audio, syncs them to the character's mouth movements, and generates a video where the person speaks or sings along to the audio track. The output is an MP4 with the uploaded audio embedded. No text-to-speech step needed. You bring the voice, the model brings the motion.

Model

  • Wan 2.1 14B FusionX. A community fine-tune of Alibaba's Wan 2.1 video model optimized for visual quality and motion fidelity. Paired with the MultiTalk adapter for audio-driven lip-sync and a Wav2Vec2 encoder for vocal feature extraction.

  • Detailz LoRA (strength 0.8). Enhances fine facial detail, skin texture, and sharpness.

  • LightX2V Distill LoRA (strength 0.8). Reduces inference steps while maintaining quality.


HOW IT WORKS

Step 1. Upload your image
A portrait or headshot of the person you want to animate. Front-facing, well-lit, with a visible mouth and jaw.
Works great with: portraits · headshots · AI-generated faces · character renders

Step 2. Upload your audio
A recording of speech, singing, or voice acting. The model syncs the character's mouth to this audio. The workflow crops audio to 12 seconds.
Works great with: voiceovers · singing clips · dialogue recordings · podcast excerpts

Step 3. Write a short prompt
Describe the action and mood. "A woman calmly singing" or "A man speaking with a confident expression" tells the model what motion style to generate around the lip-sync.

Step 4. Hit run and download
The workflow extracts vocal features, generates the lip-synced video at 25fps, and combines it with your original audio into an MP4.
Ready for: Premiere · DaVinci Resolve · After Effects · TikTok · YouTube · podcasts

First time? Upload a clear front-facing portrait and a short audio clip. Leave every setting as-is.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard talking head video — Default settings (512x512, 169 frames, 25fps, 5 steps). Upload your image and audio and run.

  • Singing video instead of speech — Change the prompt to describe singing. "A woman singing with subtle head movement and emotion" steers the motion toward a musical performance rather than a conversation.

  • Longer clip — Increase the frame count beyond 169. More frames at 25fps extends the video. Adjust the audio crop end time to match. Generation time scales with length.

  • More facial detail — Raise the Detailz LoRA strength from 0.8 toward 1.0. Skin pores, hair strands, and eye reflections sharpen. Trade-off: slight increase in generation time.

  • Lip-sync looks off — Use cleaner audio with less background noise. The Wav2Vec encoder reads vocal features more accurately from isolated speech or singing than from heavily mixed tracks. The workflow includes audio separation, but starting with clean audio helps.

  • Face looks distorted or artifacted — Use a higher-quality, front-facing portrait with even lighting and no occlusion around the mouth and jaw. Extreme angles or low-light photos make lip-sync harder.

Prompt: Keep it short and focused on the performance, not the scene. "A man speaking calmly with slight head nods" is specific. "A person in a room talking about something" adds nothing the model can use.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎤 Audio-Driven Talking Head Videos
Upload a voiceover and a portrait to produce a talking head clip where the character's mouth syncs to your recorded speech.

🎵 Singing and Music Performance Clips
Upload a vocal track and a character image to generate a singing performance video with matched lip movement and musical expression.

🎙️ Podcast and Narration Visuals
Turn a podcast audio clip into a video of a host or character speaking, giving audio-only content a visual presence for social distribution.

👤 Virtual Spokesperson Content
Generate a consistent talking character from a single portrait across multiple audio clips for explainer videos, product walkthroughs, or course content.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Front-facing portraits with visible mouth, jaw, and chin

  • Clean audio with isolated speech or vocals

  • Short, focused prompts describing the performance mood

  • Audio under 12 seconds for best sync quality

⚠️ May produce softer results

  • Profile or three-quarter angle portraits where the mouth is partially hidden

  • Audio with heavy background music, reverb, or overlapping voices

  • Portraits with hands, objects, or hair covering the lower face

  • Very long audio clips that exceed the crop window


FAQ

What is MultiTalk in this workflow?
MultiTalk is an audio-driven lip-sync adapter for Wan 2.1. It takes vocal features extracted from your uploaded audio using a Wav2Vec2 encoder and conditions the video generation so the character's mouth movements match the speech or singing in the audio. It works with any language because it reads acoustic features, not words.

Does this workflow work with singing or only speech?
Both. The Wav2Vec encoder reads vocal features from any audio containing a human voice. Speech, singing, humming, and voice acting all drive the lip-sync. Change the prompt to describe the performance type so the model generates matching head movement and expression.

What is Wan 2.1 FusionX?
FusionX is a community fine-tune of Alibaba's Wan 2.1 14B video model. It improves visual quality, motion smoothness, and prompt adherence over the base model. This workflow pairs it with a Detailz LoRA for sharper facial detail and a LightX2V distillation LoRA that reduces inference to 5 steps.

Can I use my own voice recording as the audio input?
Yes. Record your voice, upload the file, and the model syncs the character's mouth to your speech. The workflow includes an audio separation step that isolates vocals from background noise, but starting with a clean recording produces the best results.

Does the lip-sync work in languages other than English?
Yes. The Wav2Vec encoder reads acoustic features from the audio signal, not language-specific phonemes. Lip-sync works across languages including English, Mandarin, Japanese, Korean, Spanish, and others. The audio drives the motion regardless of language.

Is Wan 2.1 FusionX licensed for commercial use?
Wan 2.1 FusionX is a community derivative of Alibaba's open-source Wan 2.1 model. The MultiTalk adapter and LoRAs each have their own license terms. Check each component's license on its model page for commercial use in your specific project.

How to run audio-driven lip-sync video generation online?
You can run audio-driven lip-sync video generation online through Floyo. No installation, no setup, no local GPU needed. Open the workflow in your browser, upload your image and audio, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a portrait and an audio clip, describe the performance, and hit run.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N