replicate/cog-ltx-video-0.9.7-distilled: The Repo That Makes Video AI Deployable

A Cog wrapper for LTX Video that hides cold starts, GPU memory pressure, and awkward geometry rules behind a clean API.

8 min read • View on GitHub • More from replicate

A black-ink editorial scene of a conveyor belt carrying a film reel toward a padded shipping crate while clamps, gauges, and measuring tools keep every piece aligned. It explains that the repo's real value is packaging, not the model headline itself.
The wrapper acts like packaging, not the payload.
Key Takeaways

The fastest way to miss this repo is to treat it like a model release. It is really a delivery system for LTX Video, which means the interesting work lives in the wrapper: how weights arrive, how inputs are resized, how memory is trimmed, and how the API hides the ugly parts. That is why replicate/cog-ltx-video-0.9.7-distilled reads less like a demo and more like a service manual.

Why the wrapper exists

Replicate is not trying to invent a new video model here. It is packaging Lightricks' LTX Video into Cog, the thin deployment layer that makes a model feel like an API instead of a research repo. The point is compatibility, repeatability, and a sane path from prompt to output.

A high-quality video generation model that creates videos from text prompts, images, or existing videos using the LTX Video 0.9.7-distilled architecture.

lucataco, Top Contributor · Repo README

That description is accurate, but it undersells the real job. The repo is built to survive production constraints that the model itself does not care about. `predict.py` carries the inference path, `cog.yaml` defines the runtime, `requirements.txt` pins the stack, and `.github/workflows/push.yaml` automates the release loop. That is the shape of a deployment artifact, not a notebook.

How the machinery works

The heart of the repo is the split inside `predict.py`. One pipeline handles base generation, another handles latent upsampling, and the wrapper decides when each stage runs. Before either stage begins, the code makes the input geometry fit the model's expectations, which matters because video systems tend to fail on the edges, not the center.

The wrapper sits between user input and the model's two-stage output path.

That split is the real engineering story. Base generation gets the motion and timing right. Latent upsampling restores spatial detail without asking the whole pipeline to carry full resolution from the start. The wrapper also rounds dimensions to acceptable sizes, which is the kind of math users never want to think about and model code can never quite ignore.

A close-up of a metal ruler, a film frame, and a smaller nested frame being expanded into a larger grid. It explains why the wrapper pads and rounds dimensions before inference so the model can process the video cleanly.
Geometry rules are not a footnote. They are the difference between a smooth run and a broken one.

The dull parts are the product

The code also spends time on the things nobody markets. `pget` pulls large weights fast enough to make cold starts less painful. `enable_tiling()`, `enable_attention_slicing()`, and `enable_vae_slicing()` cut memory pressure so the model can fit on practical GPUs. Even the `go_fast` path is a product choice, not a science fair trick. It says the repo is optimizing for usable latency, not just benchmark glory.

The first DiT-based video generation model capable of real-time performance, `ltx-video-0.9.7` produces high-quality 30 FPS videos at 1216×704 resolution.

aimodels-fyi, Author · DEV Community

That is the distinction. The model headline gets attention. The wrapper decides whether the model can be called repeatedly without a human babysitting the machine. In that sense, the most important code here is the code that makes the rest of the code disappear.

How it compares

PathWhat you getWhat you still manage
Raw Lightricks model codeDirect access to the model and its latest featuresEnvironment setup, weight delivery, VRAM limits, input sizing, and API wrapping
replicate/cog-ltx-video-0.9.7-distilledA ready-to-call service with a predictable deployment shapePrompt design, cost, and the product decisions around resolution and speed
Generic managed endpointLess infrastructure work at launchLess control over model-specific tricks like latent upsampling and geometry handling

If you self-host the original stack, you inherit the full burden of environment drift, weight delivery, and sizing mistakes. If you use a generic managed endpoint, you gain convenience but lose some model-specific control. This repo sits in the narrow middle: specialized enough to respect the model's quirks, standardized enough to ship on Replicate without a custom deployment playbook.