MiniMax H3 Open Weights · Image & Text to Video
Generate 2K video with stereo sound from a start image and a prompt using MiniMax H3 (Hailuo 3.0), the open-weights model. Mute the image to go text-only.
hailuo 3.0
image to video
minimax h3
text to video
video with audio
6
389
Nodes & Models
ResolutionSelector
KSamplerSelect
LoadImage
UNETLoader
minimax_h3_fl2va_pruned_bf16.safetensors
BasicScheduler
VAEDecode
VAELoader
minimax_h3_video_vae_fp16.safetensors
minimax_h3_audio_vae_fp32.safetensors
CLIPLoader
qwen3vl_32b_minimax_h3_bf16.safetensors
SamplerCustomAdvanced
RandomNoise
SaveVideo
BasicGuider
VAEDecodeAudio
CreateVideo
MiniMaxH3ImageToVideo
ComfyMathExpression
PrimitiveFloat
FloyoStickyNote
ABOUT THE WORKFLOW
Image or text to video
Upload a start image and a prompt, and the model animates it into a clip with stereo sound, building the picture and audio together in one pass. No image? Skip it and the model generates the whole clip from your prompt alone.
Model
MiniMax H3 (Hailuo 3.0) by MiniMax. A 33B multimodal video model that generates up to 15 seconds of 2K video at 24 fps with native stereo audio in a single pass. Open weights under the MiniMax H3 Community License.
HOW IT WORKS
Step 1. Upload your start image
The image the video begins from. It sets the subject, scene, and framing.
Works great with: portraits · characters · products · scenes
Step 2. Write your prompt
Describe the motion, camera, lighting, and sound you want, like "a woman walks through a sunlit garden, camera tracks beside her, birds singing."
Step 3. Run text-only instead (optional)
No image? Mute the image node and the model generates the whole clip from your prompt alone.
Step 4. Hit run and download
The model builds the picture and stereo audio together and saves one video file. This is a large model, so give it longer to finish.
Ready for: editing timelines · social · ads
First time? Leave every setting as-is. The defaults (16:9 · scale 0.4 · 5 seconds) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard use (most people) — 16:9, scale 0.4, 5 seconds, random seed. The right starting point for almost everyone.
Want more resolution — raise the scale above 0.4 for more pixels. It sharpens the clip and adds to the generation time.
Want a longer clip — set the duration to 10 or 15 seconds. Anything past 15 seconds needs a separate extend step.
Change the frame shape — pick a different aspect ratio, like 9:16 for vertical or 1:1 for square.
Want the exact same result again — set a fixed seed in place of random. Random gives you a fresh take on each run.
Keeping generation time down — raise the scale or the duration one at a time, since both multiply how long a run takes.
Prompt: Describe the motion, camera, lighting, and any sound you want, since the model generates the audio from the same prompt. "A woman walks through a sunlit garden, camera tracks beside her, birds singing" beats "a garden video." Replace the default prompt with your own before your first run.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Ad & Promo Clips
Generate short branded spots with sound from an image and a brief.
🎞️ Bring a Photo to Life
Animate a still into a short clip with matching ambient sound.
✍️ Text to Video
Mute the image and generate a whole scene from a prompt alone.
📱 Social & Short-form
Make vertical clips with audio for feeds without a shoot.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
A clear, well-lit start image
Prompts that name motion, camera, and sound
A single, defined scene
Clips of 15 seconds or less
⚠️ May produce softer results
Cluttered or low-resolution start images
Vague motion like "make it move"
Clips pushed past 15 seconds in one run
Contradictory camera or motion notes
FAQ
What is MiniMax H3 (Hailuo 3.0)?
MiniMax H3, also called Hailuo 3.0, is MiniMax's 33B multimodal video model. It launched on July 31, 2026, and its weights were released on August 3, 2026. It reads text, images, video, and audio as one input and outputs 2K video at 24 frames per second, up to 15 seconds, with native stereo audio. This workflow runs the open weights.
Does the video come with sound?
Yes. MiniMax H3 generates native stereo audio in the same pass as the picture, and this workflow saves both together in one video file. Dialogue, effects, and ambient sound are timed to the action, so you get a finished clip with audio rather than a silent render.
Can I run it as text-to-video without an image?
Yes. Mute the image node and the model builds the entire clip from your prompt alone, with no start frame. Keep the image loaded for image-to-video, where the start frame sets the subject and framing.
What resolution and clip length can I get?
The model outputs 2K video at 24 frames per second, and you can set the duration to 5, 10, or 15 seconds. The scale setting controls how many pixels the clip has, and anything longer than 15 seconds needs a separate extend step rather than one generation.
Can I use MiniMax H3 commercially, and are there regional restrictions?
MiniMax H3 is open weights under the MiniMax H3 Community License. Non-commercial use is free, and commercial use is free for organisations under 20 million dollars in annual revenue with attribution, while larger organisations need written authorisation. The license also excludes the United States, the European Union, the United Kingdom, and South Korea from running the weights and their outputs locally, so review the current terms for your region before you rely on it.
How to run MiniMax H3 online?
You can run MiniMax H3 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload a start image, write a prompt, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A creator generates a clip and likes it. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Upload a start image and a prompt, then run it. Mute the image to go text-only.
Questions? Watch the free course or check the FAQ above.
Read more




