Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX 2.5 for Text to Video

Generate video with synchronized audio from a text prompt using LTX-2.5, the 22B open-weights model by Lightricks. Write a scene, hit run, get a clip with sound.

102

Gen time: ~1 min 31 secs

Nodes & Models

ResolutionSelector
ManualSigmas
MarkdownNote
KSamplerSelect
RandomNoise
LTXVConcatAVLatent
SamplerCustomAdvanced
SaveVideo
LTXVLatentUpsampler
ComfyMathExpression
EmptyLTXVLatentVideo
LTXVAudioVAEDecode
PrimitiveInt
CLIPTextEncode
LTXVConditioning
LTXVEmptyLatentAudio
LTXVSeparateAVLatent
CreateVideo
LatentUpscaleModelLoader
VAEDecodeTiled
PrimitiveStringMultiline
PreviewAny
ComfySwitchNode
PrimitiveBoolean
UNETLoader
VAELoader
CLIPLoader
LTXVDualCFGGuider

ABOUT THE WORKFLOW

Generate a Video from Text Write a prompt describing a scene, and LTX-2.5 generates a video clip with synchronized audio in a single pass. No separate audio step. The built-in prompt enhancer rewrites your description for better results, and a latent upscaler sharpens the frames before the final file is saved.

Model

  • LTX-2.5 (22B distilled) by Lightricks. An open-weights video model that generates picture and sound together from a text prompt. Strong at realistic motion, steady subjects, and native audio.


HOW IT WORKS

Step 1. Write your prompt Describe the scene you want: subject, action, camera movement, and any sounds. The more specific, the better. Works great with: landscapes · people · animals · ambient scenes

Step 2. Hit run The prompt enhancer rewrites your description into one the model reads better (you can turn this off). LTX-2.5 generates the video and audio together, sharpens the frames, and saves the clip. Ready for: Premiere Pro · DaVinci Resolve · After Effects · any video editor

First time? Leave every setting as-is. The defaults (16:9 widescreen, 5 seconds, 24 fps, prompt enhance on) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard video (most people) — 16:9 at 0.9 MP, 5 seconds, 24 fps, prompt enhance on, random seed. The right starting point for almost everyone.

  • Quick test before a longer run — Drop megapixels to 0.3 or 0.4. You get a smaller, faster preview to check motion and composition before committing to full resolution.

  • Higher resolution for production — Raise megapixels to 1.2 or higher. Sharper frames, longer generation time. Resolution and duration together multiply the compute cost.

  • Longer or shorter clip — Change duration. Default is 5 seconds. Shorter clips render faster. Longer clips hold more action but take more time.

  • Square or vertical format — Switch the aspect ratio in the Resolution Selector. Options include 1:1, 9:16, 4:3, and others. Dimensions snap to multiples of 32 automatically.

  • Reproduce a result you liked — Lock the seed to the number from that run. Same seed plus same prompt gives the same clip, so you can nudge one thing at a time.

  • The prompt enhancer is overwriting your intent — Turn prompt enhance off. The model reads your exact words instead of a rewritten version. Good for precise direction.

Prompt: Describe the scene, the action, the camera, and the sound. "A woman walks through a rainy night market, neon reflections on wet pavement, ambient city noise" is clearer than "cool city scene." Say what moves and how. Mention sounds you want in the clip.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Filmmakers and Video Editors Generate a scene with matching audio from a text description. Use it for storyboard previsualization, placeholder shots, or b-roll that fits a specific mood and soundscape.

🎵 Musicians and Audio-Visual Artists Create music video concepts or ambient visual loops with synchronized sound. One prompt produces both the picture and the audio, so the two stay in sync from the start.

📱 Social Media and Content Creators Produce short-form video with sound for reels, stories, or ads. Write the scene, run it, and get a clip ready for posting or editing.

🎮 Game Developers and World Builders Generate environment previews, cinematic intros, or ambient scene prototypes with matching audio. Test how a location looks and sounds before building it in an engine.

📚 Educators and Explainer Creators Create illustrated scenes with narration-ready audio for courses, training materials, or presentations. Describe the setting and the sounds, and get a clip to drop into your timeline.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Scenes with clear subjects and defined motion

  • Landscape and environment shots

  • Ambient audio (rain, city noise, wind, crowds)

  • Camera movements described in the prompt (pan, zoom, tracking)

⚠️ May produce softer results

  • Many small, fast-moving subjects in one frame

  • Resolutions above 1.0 MP on consumer hardware

  • Long durations combined with high resolution

  • Scenes requiring precise lip-synced dialogue


FAQ

What is LTX-2.5 and who made it? LTX-2.5 is a 22-billion parameter open-weights video model by Lightricks, released August 2026. It generates video and synchronized audio together from a text prompt in a single pass. It succeeds LTX-2.3 and uses a Gemma 4 text encoder for stronger prompt understanding.

Does LTX-2.5 generate audio, or do I need a separate step? Audio is generated in the same pass as the video. The model produces picture and sound together through bidirectional cross-attention, so there is no separate audio stage. The output is one video file with synchronized sound baked in.

What resolution and duration can LTX-2.5 produce? This workflow defaults to 16:9 at 0.9 megapixels (about 1280x736) and 5-second clips at 24 fps. You can raise megapixels up to 2.0 (about 1920x1088 at 16:9) and adjust duration and aspect ratio. Higher values take longer to generate.

Is LTX-2.5 free for commercial use? LTX-2.5 ships under the LTX-2.x Community License, not a standard open-source license. Organizations with under $10 million in annual revenue can use it commercially for free. Organizations above that threshold need a paid commercial license from Lightricks. Review the license terms before production deployment.

How does LTX-2.5 compare to other text-to-video models like Wan or CogVideo? LTX-2.5 stands out for native audio generation in the same pass as video, where most competitors generate silent video and require a separate audio model. Published inference speed is also fast: 6.8 seconds for a 10-second clip on high-end hardware. The tradeoff is that audio quality may not match dedicated audio synthesis models for complex dialogue or music.

What hardware does LTX-2.5 need to run locally? LTX-2.5 needs a minimum of 16 GB of VRAM for the distilled checkpoint. On consumer GPUs, expect longer generation times, especially above 1.0 megapixels. This workflow runs on cloud GPUs, so local hardware is not a concern.

How to run LTX-2.5 online? You can run LTX-2.5 online through Floyo. No installation, no setup, no model downloads to manage. Open the workflow in your browser, write your prompt, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A director runs a generation and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Write your first prompt and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N