Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX 2.3 and LTXDirector for Video Generation

Build a clip on a timeline with LTX 2.3, the open-weight 22B model that writes picture and synced audio in one pass. Drop an image in, write the shot, run.

130

Generates in about 4 mins 54 secs

Nodes & Models

KSamplerSelect
LatentUpscaleModelLoader
RandomNoise
UNETLoader
VAELoaderKJ
DualCLIPLoader
VAELoader
MarkdownNote
ModelPreviewOverrideKJ
ConditioningZeroOut
LTXVConditioning
LTXVConcatAVLatent
BasicScheduler
CFGGuider
SamplerCustomAdvanced
LTXVSeparateAVLatent
LTXVLatentUpsampler
LTXVAudioVAEDecode
VAEDecode
CreateVideo
SaveVideo
LTXDirector
LTXDirectorGuide
LTXDirectorCropGuides

ABOUT THE WORKFLOW

Direct a Video on a Timeline Drop a start image onto the timeline, write what happens across each stretch of time, and hit run. The model draws picture and sound at the same moment, so footsteps, waves, and voices land on the frames that show them. You get back one video file with the audio already inside it.

Model

  • LTX 2.3 by Lightricks. A 22 billion parameter audio-video model that writes video and synchronized sound in a single pass. This build is the distilled version, tuned to finish in 8 steps.


HOW IT WORKS

Step 1. Drop a start image on the timeline The first frame of your clip. It sets the subject, the scene, and the shape of the frame. Works great with: photos · renders · film stills · concept art

Step 2. Write the global prompt The look that holds across the whole clip. Subject, setting, light, film style. Example: "An elderly fisherman in a worn knit sweater on a rocky shoreline at golden hour, cinematic documentary style, shallow depth of field, 35mm film look."

Step 3. Write the segment prompt What happens during that stretch of time. Example: "The camera drifts right in a slow arc as he mends the net, waves roll over the rocks behind him." Split the timeline and give each stretch its own prompt to change the action partway through.

Step 4. Add audio or a motion clip (optional) The audio track takes a voice line or a music bed. The motion track takes a video whose movement you want copied. Leave both empty and the model writes its own sound.

Step 5. Hit run and download A first pass draws picture and sound together at half size. The result scales up and a short second pass adds detail. You get one video file with audio. Ready for: Premiere Pro · DaVinci Resolve · After Effects · CapCut

First time? Leave every setting as-is. The defaults (5 seconds · 24 fps · 8 steps then 4) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard clip (most people) — 5 seconds · 24 fps · 8 steps then 4 · seed fixed. The right starting point for almost everyone.

  • Want a longer scene — Drag the segment out on the timeline. Frame count is read from the timeline, so the length follows what you draw rather than a number in a box.

  • Need several beats in one shot — Split the timeline and write a prompt for each stretch. The whole run is generated as one continuous clip, so the beats flow into each other instead of cutting.

  • Want your own music or voice over — Drop a clip on the audio track. Audio inpainting stays on, so generated sound blends into what you bring in rather than stopping where your clip ends.

  • Want to copy motion from a video — Drop a reference clip on the motion track. The movement carries over while your image and prompt set the look.

  • The motion is wrong — Rewrite the segment prompt before you touch a number. Global prompt controls the look, segment prompt controls the action.

  • The result looks flat — Leave guidance at 1. The distilled build is tuned for that value, and raising it washes out contrast and motion.

  • A prompt box will not open on a segment — Give the image at least one second of length on the timeline first. Sizes also snap to multiples of 32, so odd dimensions get rounded.

Prompt: Write the global prompt for the look and the segment prompt for the action, and say what you want to hear as well as what you want to see. "The camera drifts right in a slow arc, waves break on wet rock, gulls call in the distance" gives the model far more to work with than "make it cinematic."


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Previz and Short Film Block out a scene with sound before you shoot it, and show a director how the beat plays rather than describing it.

📱 Social Video With Native Audio Turn a still into a vertical clip that already has ambience or a voice on it, ready to post without an edit pass.

🎙️ Dialogue and Voice-Led Scenes Drop a recorded voice line on the audio track and let the picture follow the performance.

📦 Product and Ad Spots Animate a product still into a short spot with room tone and motion, then hand the file straight to an editor.

🎞️ Motion Reference Work Feed a reference clip to the motion track and carry its camera move onto a new subject and setting.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Well-lit stills with one clear subject

  • Prompts that name a camera move and a sound

  • Short clips of 5 to 10 seconds

  • Vertical and portrait framing

  • Fine detail like hair, fabric, and water

⚠️ May produce softer results

  • Readable text inside the frame

  • Story logic that spans many cuts

  • Crowded scenes with several competing subjects

  • Vague prompts like "make it cinematic"

  • Guidance pushed above 1 on the distilled build


FAQ

What is LTX 2.3? LTX 2.3 is an open-weight audio-video model from Lightricks, released in March 2026. It is a 22 billion parameter diffusion transformer that generates video and matching audio in a single pass rather than adding sound afterward. The family supports up to native 4K at 50 frames per second and clips up to 20 seconds. This workflow runs the distilled build at 24 frames per second.

Does LTX 2.3 generate audio with the video? Yes. Sound and picture come out of the same pass, so dialogue, ambience, and effects line up with what is happening on screen without a separate dubbing step. That was the headline change when Lightricks open-sourced the LTX-2 family in January 2026, since open video models before it produced silent output.

What is the difference between the LTX 2.3 dev and distilled models? The dev checkpoint is the full trainable model at bf16 and carries the higher quality ceiling. The distilled checkpoint finishes in 8 steps at guidance 1, which cuts generation time and memory use for a small drop in fidelity. This workflow uses the distilled build, then recovers detail with a spatial upscaler and a short second sampling pass.

Can I use LTX 2.3 for commercial projects? LTX 2.3 ships under the LTX-2 Community License, not Apache 2.0. Lightricks makes it free for academic research and for commercial use by companies under $10M in annual recurring revenue. Organizations above that threshold need a commercial license from Lightricks. Read the license on the LTX-2 GitHub repository before you ship client work.

What GPU do you need to run LTX 2.3 locally? A 22 billion parameter model is a heavy local install. The distilled checkpoint runs on 16GB to 24GB cards with FP8 or GGUF quantization and sequential offloading, and the full bf16 dev pipeline wants 32GB or more. You also need the Gemma 3 12B text encoder, both video and audio VAEs, and the spatial upscaler on disk before the first run.

How is a timeline different from generating clips and stitching them together? Separate clips drift. Lighting shifts, the subject changes shape, and the cuts show. Here the timeline is read as one continuous generation, so each segment prompt changes what happens without breaking the shot. One image, one global look, several beats inside a single take.

How to run LTX 2.3 online? You can run LTX 2.3 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, drop an image on the timeline, write your prompts, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Drop an image on the timeline, write what happens, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N