Wan 3.0 · Image to Video With Audio
Animate a still into a 30-second clip with sound using Wan 3.0, Alibaba's latest video model. Upload an image, describe the motion, and hit run. 1080p output.
alibaba
image to video
video with audio
wan 3.0
0
26
Nodes & Models
AlibabaWan30ImageToVideo_floyo
VideoToFrames
LoadImage
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Animate a Still With Sound Upload an image and describe the motion. Wan 3.0 builds the picture and the audio together in one pass and returns a clip with sound. Up to 30 seconds in a single run at up to 1080p. Add an optional last frame to steer where the shot ends.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
Wan 3.0 by Tongyi Lab at Alibaba. Public beta 6 August 2026. The current flagship of the Wan video family. Unifies text-to-video, image-to-video, reference-to-video, and editing into a single model. Generates native 30-second clips at up to 1080p with audio in one pass. Includes intelligent duration control that matches clip length to the prompt. Closed model with no open weights.
HOW IT WORKS
Step 1. Upload your image The frame the clip opens on. It sets the subject, the scene, and the framing. Works great with: portraits · product shots · character art · location photos
Step 2. Upload a last frame (optional) The frame the clip should land on. Helps the model plan the motion path from start to finish.
Step 3. Describe the motion Say what moves and how. Write the prompt as one evolving take: subject motion, camera path, lighting, and sound across the full duration.
Step 4. Set resolution and duration Pick from 480P to 1080P and a duration up to 30 seconds. Both settings drive what a run costs.
Step 5. Hit run and download The model builds the picture and the audio together and saves the clip under video/Wan3.0. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects
First time? Leave every setting as-is. The defaults (1080P · 5 seconds · 16:9 · audio on · watermark off · random seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard clip (most people) — 1080P · 5 seconds · 16:9 · audio on · watermark off · random seed. The right starting point for almost everyone.
Need a longer clip — Raise duration up to 30 seconds. Write the prompt as one continuous take that fills the time. Duration and resolution together drive cost.
Want the shot to land on a specific frame — Upload a second image as the last frame. The model plans the motion between the two.
Need a different shape — Switch the aspect ratio. 16:9, 9:16, and 1:1 are available.
Sound is not matching the action — Describe the sound separately in the prompt after the visual action. Name the room tone, effects, and score.
Want silent output — Turn audio off. The model skips the sound pass.
Repeat a take you liked — Set a fixed seed. The same image, prompt, and seed give you the same clip back.
The clip is generic — Write a longer prompt. Name the camera speed, the lighting direction, the surface textures, and the atmosphere.
Prompt: Write it as one evolving take. "She steps off the porch into light rain, camera follows from behind, raindrops hit her umbrella, she pauses at the gate and looks back over her shoulder, trees sway, puddles reflect the house lights, soft rain pattering and distant thunder." Fill the duration with action. Name the camera, the subject, and the sound as separate threads inside the same paragraph.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 One-Take Character Scenes Animate a character portrait into a 30-second continuous shot with camera moves, action, and ambient sound.
📢 Ads and Product Films Turn a product still into a finished spot at 1080p with room tone and motion, up to 30 seconds from one image.
🎞️ Start-to-End Shots Upload a first frame and a last frame and get smooth motion between the two with matched audio across the full duration.
🎨 Previz and Concept Test how a still moves and sounds at full scene length before committing to production.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clear, well-lit start images with one readable subject
Prompts written as one evolving take across the full duration
Multi-character scenes with named interactions
Prompts that describe both the picture and the sound
⚠️ May produce softer results
Short keyword prompts padded to 30 seconds
Expecting 4K output (the model caps at 1080P)
Fast collisions and extreme physics
Start and last frame images with wildly different framing
FAQ
What is Wan 3.0? Wan 3.0 is the latest video model from Tongyi Lab at Alibaba, in public beta since 6 August 2026. It unifies text-to-video, image-to-video, reference-to-video, and editing into one model. It generates native 30-second clips at up to 1080p with audio in a single pass. This workflow uses the image-to-video mode, animating a still you upload.
How long can a Wan 3.0 clip be? Up to 30 seconds in a single pass. That is double the 15-second ceiling on Wan 2.7. The model includes intelligent duration control that recommends a clip length based on your prompt.
Does Wan 3.0 generate audio with the video? Yes, by default. Sound is generated in the same pass as the picture. Describe the sound in your prompt and it goes onto the track. Turn audio off if you want a silent clip.
What is the difference between this and the Wan 3.0 text-to-video workflow? Text to video generates a clip from a prompt alone, with no input image. This workflow starts from a picture you upload, using it as the opening frame and conditioning the clip on its look. Use text to video to build a scene from scratch and this one to animate a still you already have.
What is the difference between Wan 3.0 and Wan 2.7? Wan 2.7, released April 2026, goes up to 15 seconds and adds a thinking mode for compositional planning. Wan 3.0, beta August 2026, doubles the duration to 30 seconds, unifies four separate Wan 2.7 endpoints into one model, and adds Omni-Reference input that accepts documents and webpages. Neither has open weights.
Is Wan 3.0 open source? No. Wan 3.0 is API-only with no published weights. Open weights in the Wan family stop at Wan 2.2, which is released under Apache 2.0.
How to run Wan 3.0 image to video online? You can run Wan 3.0 image to video online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your image, describe the motion, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Upload an image, describe the motion, and run it. The clip comes back with sound, up to 30 seconds.
Questions? Watch the free course or check the FAQ above.
Read more



