MiniMax H3 Open Weights · Image & Text to Video

Generate 2K video with stereo sound from a start image and a prompt using MiniMax H3 (Hailuo 3.0), the open-weights model. Mute the image to go text-only.

Audio
hailuo 3.0
image to video
minimax h3
text to video
Video
video with audio

2.7k

Gen time: ~15 min 43 secs

Nodes & Models

ResolutionSelector
KSamplerSelect
LoadImage
UNETLoader
BasicScheduler
VAEDecode
VAELoader
CLIPLoader
SamplerCustomAdvanced
MarkdownNote
RandomNoise
BasicGuider
SaveVideo
VAEDecodeAudio
CreateVideo
MiniMaxH3ImageToVideo
ComfyMathExpression
PrimitiveFloat

ABOUT THE WORKFLOW

Generate Video With Sound Upload a start image and describe what happens next. MiniMax H3 reads the picture and the words together, then writes the frames and the stereo audio in the same pass and returns a finished clip. Turn the image off and the model builds everything from your prompt.

Model

  • MiniMax H3 by MiniMax, the lab behind the Hailuo video line, also called Hailuo 3.0. A 33 billion parameter open-weight model that takes text, images, video, and audio as one input and generates picture and sound together.


HOW IT WORKS

Step 1. Upload your start image The frame the clip begins from. It sets the subject, the scene, and the framing. Works great with: portraits · product shots · character art · location photos

Step 2. Write your prompt Describe the motion, the camera, the light, and the sound you want. "A woman walks through a sunlit garden, camera tracks beside her, birds singing."

Step 3. Run from text alone (optional) No image? Mute the image loader with Ctrl+M and the model builds the whole clip from your prompt.

Step 4. Set your aspect ratio and clip length Pick a shape on the resolution selector and type a length in seconds. Both settings drive how long a run takes.

Step 5. Hit run and download You get one video file per run at 24 frames per second with stereo sound, saved under video/MiniMax_H3. Rename that path to sort your output. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects

First time? Leave every setting as-is. The defaults (16:9 · 5 seconds · random seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard clip (most people) — 16:9 · 5 seconds · random seed. The right starting point for almost everyone.

  • Faster runs at the same quality — Drop megapixels to about 1.0. The open weights are built around a 768 pixel short edge, so the 1.5 default already sits above their range and costs time without adding detail.

  • Quick test before a full run — Drop megapixels to 0.4 and keep the clip at 5 seconds. Cheapest way to check whether the motion and the sound read the way you wanted.

  • Need a longer clip — Raise the duration value. The model handles 5 to 15 seconds. Duration and megapixels together are what drive run time, so raise one at a time.

  • Want the clip to land on a specific frame — The group node has a last frame input sitting open. Connect a second image loader to it and the model builds the motion from your first still to your last.

  • Repeat a take you liked — Set a fixed seed number instead of leaving it on random. The same image, prompt, and seed give you the same clip back.

  • The motion or the sound is wrong — Rewrite the prompt before you touch a setting. Picture and audio come from the same prompt, so name the sound you want as well as the movement.

  • Clip length looks slightly off — Frame counts snap to a fixed ladder, so a 5 second request runs 124 frames, which is about 5.2 seconds at 24 fps. That is expected.

Prompt: Write the shot, not the subject. "The chef lifts the pan, flames rise, oil sizzles, camera pushes in slowly" gives you more than "a chef cooking." Name the sound directly, like "rain on a metal roof" or "a low crowd murmur," because the audio track comes from the same prompt as the picture.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Social Video Turn one still into a short clip with matching sound for Reels, TikTok, or Shorts, with no separate audio pass.

🗣️ Dialogue and Voice Write speech into the prompt and get it spoken in the clip, with support for eleven languages including English, Spanish, Japanese, and Arabic.

📢 Product and Ad Spots Animate a product shot into a 5 to 15 second spot with room tone and effects already on the track.

🎨 Previz and Mood Films Test how a frame moves and sounds before booking a shoot or committing to a full animation pass.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-lit start images with one main subject

  • Prompts that name the movement, the camera, and the sound

  • Shot lists broken into timed beats for longer clips

  • Clips of 5 to 10 seconds

⚠️ May produce softer results

  • Megapixel values pushed well past the 768 pixel short edge

  • Prompts that describe the image instead of what happens

  • Crowded scenes with several subjects moving at once

  • Dialogue written longer than the clip can hold


FAQ

What is MiniMax H3? MiniMax H3 is an open-weight video model from MiniMax, launched 31 July 2026 and widely called Hailuo 3.0. It is a 33 billion parameter dense transformer with a Qwen3-VL-32B text encoder that reads text, images, video, and audio as one input, then returns video at 24 frames per second with 32 kHz stereo audio written in the same pass. Clips run 5 to 15 seconds.

Is MiniMax H3 open source? The weights went public on Hugging Face on 3 August 2026, but the license is not a standard open-source license. The MiniMax H3 Community License covers commercial use free for organizations under 20 million USD in yearly revenue, requires products built on it to display "MiniMax H3" in the interface, and excludes the US, EU, UK, and South Korea from its applicable territory for local deployment. Read the license text before you build on it.

Does MiniMax H3 generate audio with the video? Yes. The model predicts audio and video latents together, then decodes them through separate video and audio stages, so dialogue, sound effects, and room tone land on the same timeline as the picture. There is no second model and no separate audio step. Describe the sound in your prompt and it goes onto the track.

Can MiniMax H3 do text to video as well as image to video? Yes. Text-to-video and image-to-video share the same checkpoint, so this workflow covers both. Mute the image loader and the model generates the whole clip from your prompt. Leave it on and your uploaded image becomes the opening frame.

What resolution does MiniMax H3 output? The open checkpoint writes 768p class video, built around a 768 pixel short edge. The 2K output MiniMax advertises comes from a separate finishing stage that stayed behind the API and was not part of the open-weight release, so treat 768p as the ceiling here and upscale afterwards if you need more.

What hardware do you need to run MiniMax H3 locally? The download alone is around 42 GB for the text-to-video and image-to-video path. Community reports put 16GB of VRAM at workable with quantized files and offloading, 24GB to 32GB at comfortable, and 64GB or more of system RAM as the practical target once weights start moving off the card. MiniMax has published no official minimum.

How to run MiniMax H3 online? You can run MiniMax H3 online through Floyo. No installation, no setup, no 42GB download. Open the workflow in your browser, upload your image or write a prompt, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a start image, describe the shot and the sound, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N
N/A
N/A
4 weeks ago
@Ahmed jisan:@亗SK๛SINDHI:ka ek seamless Reverse World cinematic sequence: premium Pixar-inspired 3D animation, 9:16 vertical. Pehle 3 seconds mein concept instantly clear ho: ek colorful city street mein ek normal-looking human character actually ek giant animated pencil ki tarah behave kar raha hai, jabki ek cute living pencil bilkul insaan ki tarah kapde pehne hue chal raha hai. Is duniya mein yeh sab completely normal hai—kisi ko surprise nahi hota. Human-pencil character sidewalk par ek pencil sharpener ke paas jaakar apne aap ko sharpen karwata hai aur casually bolta hai, “Bhai, thoda aur tez kar dena.” Saath khadi human-like pencil apni notebook kholkar us pencil-human ko dekhte hue bolti hai, “Aaj tumhari writing kaafi smart lag rahi hai.” Pencil-human seedha jawab deta hai, “Haan, aaj mood mein hoon.” Dono naturally walk karte hain, background mein aur bhi pencil-characters coffee pee rahe hain aur humans stationery ki tarah shelves par neatly arranged hain. Camera continuously unke saath move kare, dynamic tracking shot, funny visual escalation, expressive animated body language, believable physics, detailed city environment, cinematic lighting, polished textures, natural ambient sounds, pencil scratching SFX, light comedic music. No real human faces, no live action, no photorealistic people, no reaction faces, no text, subtitles, logos, watermark or UI. Characters aur environment ki design throughout consistent rahe, no random cuts, no location change, no continuity errors. Dialogues Hindi only.

Reply

k
kokekokko
1 month ago
H3

Reply