Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Minimax H3 Max for Text to Video

Generate video with native stereo sound from a text prompt using MiniMax H3 Max, fal's post-trained version of MiniMax H3. Write a prompt, pick a duration, and hit run.

api
Minimax H3 Max
T2V
Video

71

Gen time: ~43 secs

Nodes & Models

MiniMaxH3MaxTextToVideo_floyo
VideoToFrames
CreateVideo
SaveVideo

ABOUT THE WORKFLOW

Generate a Video from Text
Write a prompt describing a scene, action, camera move, and sound. The model generates a video clip with stereo audio in a single pass. Pick a resolution, duration, and aspect ratio, then download the result.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • MiniMax H3 Max by fal. A post-trained version of MiniMax's open-weight H3 video model, tuned for prompt adherence, aesthetics, and speed. Generates video with native stereo sound from text alone.


HOW IT WORKS

Step 1. Write a prompt
Describe the scene, action, camera movement, lighting, and sound. The model reads all of it together and generates both picture and audio from the same text.

Step 2. Pick your settings
Choose resolution (768P default), duration (5, 10, or 15 seconds), and aspect ratio (16:9 default). The defaults are a good starting point.

Step 3. Hit run and download
MiniMax H3 Max generates the clip with audio in one pass. Preview it in the workflow, then download the video file.
Ready for: Premiere · DaVinci Resolve · After Effects · any NLE

First time? Leave every setting as-is. The defaults (768P · 5 seconds · 16:9 · balanced prompt expansion) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard generation (most people) — 768P · 5 seconds · 16:9 · balanced prompt expansion · random seed. The right starting point for almost everyone.

  • Quick draft or test — 480P · 5 seconds · random seed. Lower resolution generates faster and costs less. Good for checking motion before committing to full quality.

  • Longer scene — 768P · 10 or 15 seconds · 16:9. Extend the duration when you need more time for the action to play out. Longer clips cost more per run.

  • Vertical or square video — Switch ratio to 9:16 or 1:1. Use 9:16 for Reels, TikTok, and Shorts. Use 1:1 for social posts.

  • The model is changing your prompt too much — Set prompt expansion to "direct" so the model follows your words more closely. "Balanced" gives the model room to interpret. "Direct" keeps it literal.

  • Reproduce a result you liked — Lock the seed to the number from your previous run. Same seed plus same prompt produces the same output.

  • The motion or sound is off — Rewrite the prompt before changing any setting. Describe the motion step by step and name the sounds you want. The prompt drives everything in this workflow.

Prompt: Describe four things: the subject, the action, the camera, and the sound. "A lone Joshua tree in the Mojave desert. Clouds accelerate overhead as the sky shifts from golden hour to starry night. Hyper-lapse style, smooth transitions" works because it covers scene, motion, and style. "Cool desert video" does not, because the model has nothing specific to follow.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Filmmakers & Directors
Generate pre-viz clips with camera moves, lighting, and ambient sound from a text brief. Test a shot before committing to a shoot or a full CG build.

🎵 Music & Content Creators
Turn a scene description into a video with matching audio for social content, lyric videos, or mood reels. No separate sound design pass needed.

🎮 Game Developers
Produce cinematic cutscene drafts or trailer material from text. The native audio means you get atmosphere and foley in the same generation.

💼 Marketing & Ads
Create short video concepts for pitches, social campaigns, or storyboard validation. Run several takes with different prompts and pick the strongest.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Detailed scene descriptions with named actions

  • Cinematic camera directions (dolly, pan, hyper-lapse)

  • Prompts that describe sound alongside visuals

  • Landscape, nature, and architectural scenes

⚠️ May produce softer results

  • Vague prompts with no specific motion or framing

  • Prompts requesting specific real people or celebrities

  • Complex multi-character dialogue scenes

  • Requests for output longer than 15 seconds


FAQ

What is MiniMax H3 Max and who made it?
MiniMax H3 Max is a video generation model built by fal. fal post-trained MiniMax's open-weight H3 base model for stronger prompt adherence, better aesthetics, and faster inference. It is fal's tuned edition of an open MiniMax model, not a MiniMax product. It generates video clips with native stereo audio from a text prompt in a single pass.

Does MiniMax H3 Max generate sound with the video?
Yes. The model produces stereo audio alongside the video frames in the same generation. You describe the sound you want in the prompt, the same way you describe the visuals. The output is a single video file with audio already synced. No separate sound design step needed.

What resolution and duration does MiniMax H3 Max support?
H3 Max outputs at 480P or 768P, in clips of 5, 10, or 15 seconds. The default is 768P at 5 seconds in 16:9. The base MiniMax H3 model reaches 2K, but H3 Max trades maximum resolution for faster generation. For most use cases, 768P is sharp enough for social content, pre-viz, and concept work.

How is MiniMax H3 Max different from the base MiniMax H3?
H3 Max is post-trained by fal for better prompt adherence, aesthetics, and speed. It generates a 5-second 768P clip in under 3 seconds, which is roughly 35x the throughput of the official MiniMax H3 endpoint. The trade-off: H3 Max caps at 768P (versus 2K on base H3) and supports text-to-video and image-to-video only, while base H3 also offers reference generation from mixed images, clips, and audio.

What does prompt expansion do?
Prompt expansion controls how much the model rewrites your prompt before generating. "Balanced" (the default) lets the model add cinematic details and fill in gaps. "Direct" keeps your words closer to what you wrote. Use "balanced" when your prompt is short or loose. Switch to "direct" when you need the model to follow specific framing or action exactly as described.

Can I use MiniMax H3 Max output commercially?
MiniMax H3 Max is built on MiniMax's open-weight H3 model under the MiniMax H3 Community License. That license permits commercial use for organizations under $20 million in yearly revenue, with attribution, but excludes the US, EU, UK, and South Korea from local weight deployment. Because this workflow runs through a hosted API (fal via Floyo), separate API terms apply rather than the open-weight license. Review fal's and MiniMax's current terms for your specific use case before shipping commercially.

How to run MiniMax H3 Max online?
You can run MiniMax H3 Max online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, write your prompt, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Write a prompt describing your scene, pick a duration, and hit run. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N