LTX's open-weight video model — spun out of Lightricks — that turns an image into a 10-second clip in seconds, and runs on a GPU you already own.
Last verified · 2026-08-12 · by Moe Ameen
LTX-2.5 is an open-weight video and world model released on August 11, 2026 by LTX, the open-model company spun out of Lightricks. Its headline claim is speed with ownership: on two NVIDIA GB200 chips it generates a 10-second, 720p image-to-video clip in about 6.8 seconds — faster than the clip itself runs — and through LTX's managed API it returns a 1080p clip in roughly 23.7 seconds. It does both text-to-video and image-to-video, and it keeps the synchronized native audio LTX introduced in the 2.3 generation, so a render can arrive with matching sound rather than as a silent file.
What sets LTX-2.5 apart from the closed frontier is that the weights are open. The model is free to use for organizations under $10 million in annual recurring revenue, with larger companies negotiating a license, and it ships day-one on Hugging Face, inside ComfyUI, and through fal.ai and the LTX API. LTX tuned it to run on hardware creators already have: it needs roughly 16GB of VRAM at minimum and is optimized for local NVIDIA RTX GPUs and the NVIDIA DGX Spark desktop, with a distilled variant for faster local inference. A frontier-class video model that runs on one desk instead of a rented cluster is the real story here.
On capability, LTX-2.5 adds native multishot — one generation produces several connected shots that hold character, environment, lighting, and voice across cuts, instead of you stitching separate clips — plus a new diffusion video decoder for cleaner high-motion footage, a custom Gemma-based text encoder with a prompt enhancer for stronger prompt-following, automatic clip duration, native 4K HDR support with a RAW finishing workflow, and a beta precise-editing mode. The honest framing: LTX-2.5 is a model, not a content studio. It generates a clip (with sound), and that is where it stops — there is no caption burner, no per-platform reframing, no brand or persona system, and no scheduler or publishing. Treat specific speeds, resolutions, and license terms as the launch snapshot and confirm them on ltx.io before you build against them.
The thing that makes LTX-2.5 special — open weights on your own GPU — is also what leaves you holding a folder of raw clips at the end of the night. Run it in [ComfyUI](/ai-tools/comfyui) on an RTX card and you can batch a week of b-roll and image-to-video shots locally, for free, while you sleep. What you wake up to is exactly the problem: a stack of 10-second MP4s with no captions, no aspect ratios, no hook on the opening frame, and no way to post them. Native multishot keeps a character consistent inside one clip, but nothing keeps Monday's render looking like Thursday's, or makes either read as your brand. [Kompozy](/) is the no-code layer that turns that local render farm into a published channel. Drop the clips in and it burns branded, on-style captions for the sound-off feed, reframes each one to 9:16, 1:1, and 16:9 per destination, and stacks a hook overlay through [HyperFrames](/glossary/hyperframes) so the muted autoplay opening lands — with the [Persona Brief](/glossary/persona-brief) enforcing one voice across every caption regardless of what prompt produced the clip.
Then it does the two things a raw model can't. It generates the formats LTX-2.5 will never make — [carousels](/glossary/output-buckets) that break a clip into beats, quote cards, face-locked persona photos, blog articles, and email newsletters — and its own [Persona Shorts](/glossary/persona-shorts) and [Persona Frames](/glossary/persona-frames) avatar video, so you aren't limited to what the model rendered. And it publishes: scheduling and fanning the finished set across the eight social platforms plus blog and email from one queue with a per-post review gate and [Autopilot](/glossary/autopilot). LTX-2.5 is the fastest way to make the clip cheaply and locally; Kompozy is the no-code way to make it a brand and ship it everywhere.
LTX-2.5 is an open-weight AI video and world model released August 11, 2026 by LTX, the open-model company spun out of Lightricks. It generates a 10-second, 720p image-to-video clip in about 6.8 seconds on NVIDIA GB200 hardware, supports text-to-video and synchronized native audio, and runs locally on RTX GPUs.
The weights are open and free to use for organizations under $10 million in annual recurring revenue; larger companies negotiate a license. It is available on Hugging Face, inside ComfyUI, and through fal.ai and the LTX API. Confirm current license terms on ltx.io.
LTX tuned it for local inference on NVIDIA RTX GPUs and the DGX Spark desktop, with a minimum of roughly 16GB of VRAM and a distilled variant for faster generation. On two GB200 chips it renders a 10-second clip in about 6.8 seconds; local speed depends on your GPU.
Yes. It keeps the synchronized native audio introduced in LTX-2.3, so a clip can arrive with matching dialogue, effects, and ambient sound in the same generation rather than added afterward. It also adds native multishot for consistency across cuts.
Generate the clips in ComfyUI or via the LTX API, then bring them into Kompozy — no code needed. Kompozy adds branded captions, reframes per platform, and unifies brand styling, then fans the idea into a carousel, quote card, blog, and captions in your voice, scheduled and published across the eight social platforms plus blog and email.