LTX 2.3 Image to Video For AD Film
Animate any photo into a 1080p video with synchronized audio using LTX-Video 2.3, Lightricks' 22B open-source model. Upload an image, describe the motion, and hit run.
audio video
image to video
ltx video
video generation
0
27
Nodes & Models
FloyoStickyNote
ResizeImagesByLongerEdge
GetNode
LoadImage
PrimitiveBoolean
SaveVideo
PrimitiveInt
RandomNoise
LTXAVTextEncoderLoader
gemma_3_12B_it_fp4_mixed.safetensors
ltx-2.3/ltx-2.3-22b-dev.safetensors
LTXVAudioVAELoader
ltx-2.3/ltx-2.3-22b-dev.safetensors
LatentUpscaleModelLoader
ltx-2.3-spatial-upscaler-x2-1.0.safetensors
ManualSigmas
KSamplerSelect
CheckpointLoaderSimple
ltx-2.3/ltx-2.3-22b-dev.safetensors
LTXVConcatAVLatent
CFGGuider
SamplerCustomAdvanced
LoraLoaderModelOnly
ltx-2.3-22b-distilled-lora-384.safetensors
LTXVPreprocess
ComfyMathExpression
LTXVAudioVAEDecode
CLIPTextEncode
LTXVEmptyLatentAudio
LTXVSeparateAVLatent
CreateVideo
ImageResizeKJv2
VAEDecodeTiled
LTXVConditioning
EmptyLTXVLatentVideo
LTXVImgToVideoInplace
LTXVCropGuides
LTXVLatentUpsampler
SetNode
ABOUT THE WORKFLOW
Animate a Photo into 1080p Video with Audio
Upload a photo, describe the motion you want, and get a sharp 1080p video with synchronized audio in about 5 seconds of output. A two-pass pipeline generates at low resolution first, then upscales and refines in a second pass. The model produces audio that matches the scene automatically. That's it.
Model
LTX-Video 2.3 (22B) by Lightricks. A 22B parameter DiT-based audio-video foundation model (Apache 2.0) with a rebuilt VAE for sharper detail and a 4x larger text connector for precise prompt following. Paired with the official Distilled LoRA for fewer-step generation and a 2x spatial upscaler for high-resolution output.
HOW IT WORKS
Step 1. Upload your image
The photo you want to animate. The workflow runs in image-to-video mode by default.
Works great with: landscapes · portraits · product shots · environments
Step 2. Describe the motion
Write what moves and how. Be specific about camera movement, subject action, and scene changes. "The camera slowly pushes in as the woman turns her head and smiles, warm afternoon light" works better than "make it move."
Step 3. Write a negative prompt (optional)
List anything to keep out of the result, like "cartoon, video game, blurry." A negative prompt is not required but helps steer the output.
Step 4. Hit run and download
The model generates the video in two passes (low-res, then upscaled), adds synchronized audio, and returns the final 1080p result. Preview it in the workflow, then download.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any editor
First time? Leave every setting as-is. The defaults (1920×1080 · 121 frames · 24 fps) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard generation (most people) — 1920×1080 · 121 frames · 24 fps · fixed seeds. About 5 seconds of 1080p video with audio. The right starting point for almost everyone.
Want a shorter clip — Lower the frame length. At 24 fps, 73 frames is about 3 seconds. Faster generation, less credit usage.
Want a longer clip — Raise the frame length. Keep in mind that frame counts must follow the formula: divisible by 8, plus 1 (so 97, 105, 113, 121, 129). Longer clips take more time and credits.
Want portrait (vertical) video — Swap width and height to 1080×1920. The model supports native 9:16 portrait generation without cropping.
Want a different take — Change both Seed Pass 1 and Seed Pass 2 together. Each seed pair produces a different interpretation of the same prompt.
The subject freezes or drifts — Add more specific motion direction to the prompt. Describe what the subject does, when the camera moves, and how the lighting changes. Detailed prompts reduce the "slow pan across a still image" effect.
Want to skip the image and generate from text only — A mode toggle inside the workflow switches it to text-to-video. No image upload needed in that mode.
Prompt: Write like a director. Include camera movement, subject action, and scene atmosphere. "A street-level tracking shot moves through a rain-soaked Tokyo alley at night, neon signs reflected in puddles, steam rising from a food stall" gives the model clear direction. Avoid single-word instructions like "dramatic" or "cinematic" without context.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Filmmakers and Content Creators
Animate a still frame into a cinematic shot with matched audio. Turn a location photo into a scene with ambient sound, camera drift, and subject motion. Preview a shot before committing to a full shoot.
🛍️ Product and E-commerce Video
Bring product photos to life with subtle motion: a slow orbit, a fabric ripple, ambient lighting changes. The synchronized audio adds environmental sound without a separate step.
🎨 Concept Art and Pre-visualization
Animate concept art or storyboard frames to test how a scene reads in motion. Faster and cheaper than a full CG render or animatic.
📱 Social Media and Short-Form Content
Generate vertical 9:16 clips natively for TikTok, Reels, and Shorts. The model composes for portrait framing instead of cropping from landscape.
🎧 Audio-Visual Content
Generate video with synchronized ambient sound in one pass. No separate audio dubbing step needed. The model produces environmental sounds, music cues, and speech-like audio that matches the visual scene.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Detailed prompts with camera direction, subject action, and atmosphere
Landscape, portrait, and square aspect ratios
Environmental scenes with natural motion (wind, water, crowds)
Photos with clear subjects and consistent lighting
⚠️ May produce softer results
Short or vague prompts like "make it cinematic"
Extreme close-ups with rapid complex motion
Very long clips (generation time and consistency degrade with length)
Source images that are blurry, heavily compressed, or low-resolution
FAQ
What is LTX-Video 2.3?
LTX-Video 2.3 is a 22B parameter open-source audio-video foundation model by Lightricks, released under the Apache 2.0 license. It uses an asymmetric dual-stream Diffusion Transformer with a 14B video stream and a 5B audio stream connected by bidirectional cross-attention. It generates synchronized video and audio in a single pass.
What does two-pass upscaling mean?
This workflow generates the video at low resolution first (768×512), then runs a second pass with the spatial upscaler to refine and sharpen the output to 1080p. The two-pass approach produces sharper detail in faces, textures, hair, and edges than generating at full resolution in a single pass.
Does LTX 2.3 generate audio automatically?
Yes. The model generates synchronized audio alongside the video in a single pass. Environmental sounds, ambient noise, and scene-appropriate audio are produced automatically based on your prompt and input image. You do not need a separate audio generation step.
How is LTX 2.3 different from other open-source video models like Wan 2.2?
LTX 2.3 is the only open-source model that generates synchronized audio and video in a single architecture. It also supports native 4K output, native portrait (9:16) generation, and comes with official spatial and temporal upscalers. Wan 2.2 generates silent video and focuses on strong motion quality at lower resolutions. The right choice depends on whether you need audio and resolution (LTX 2.3) or prioritize motion quality at standard resolution (Wan 2.2).
What aspect ratios and resolutions does LTX 2.3 support?
The model supports landscape (16:9), portrait (9:16), square (1:1), and other standard ratios. It generates natively up to 1080p, with upscaler support to 4K. Width and height must be divisible by 32, and frame counts must follow the formula: divisible by 8, plus 1.
Is LTX-Video 2.3 free to use commercially?
Yes. LTX-Video 2.3 is released under the Apache 2.0 license, which allows commercial use, modification, redistribution, and fine-tuning. You can use the outputs in client work, published content, and commercial products.
How to run LTX 2.3 image to video online?
You can run LTX 2.3 image to video online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your image, describe the motion, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A filmmaker generates a shot and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Upload a photo, describe the motion, and run it. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more

_1784546300569.webp?width=1400&height=620&quality=80&resize=cover)
_1784546302540.webp?width=1400&height=620&quality=80&resize=cover)

_1784546300569.webp?width=104&height=104&quality=80&resize=cover)
_1784546302540.webp?width=104&height=104&quality=80&resize=cover)
_1784206805705.webp?width=400&height=300&quality=80&resize=cover)


_1783028563354.gif?width=400&height=300&quality=80&resize=cover)


