This model card focuses on the model associated with the LTX-Video model, codebase available here. LTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-time. It produces 30 FPS videos at a 1216×704 resolution faster than they can be watched. Trained on a large-scale dataset of diverse videos, the model generates high-resolution videos with realistic and varied content. You can use the model for purposes under the license: - 2B version 0.9: license - 2B version 0.9.1 license - 2B version 0.9.5 license - 2B version 0.9.6-dev license - 2B version 0.9.6-distilled license - 13B version 0.9.7-dev license - 13B version 0.9.7-dev-fp8 license…
This model card focuses on the LTX-2 model, as presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. The codebase is available here.
Runs On
What it takes to serve LTX-2 (18.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 37.8 GB | 45.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 18.9 GB | 22.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 9.4 GB | 11.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
SAVRN's Notes on LTX-2
Video with its own soundtrack is the job here. LTX-2 generates picture and audio together in one 18.9 billion parameter model, and the publisher's set runs to 69 files and 314 GB once the distilled and upscaler variants are counted. At 16-bit the weights are 37.8 GB and need about 45.3 GB in use, so one MI300X with 192 GB, at $1.85 an hour on-demand, carries it with room left over. At 8-bit the need drops to 22.7 GB, and at 4-bit to 11.3 GB.
The license is filed as other, with no summary in our record, so read the publisher's terms before a commercial run; they cover the full, distilled and upscaler models and any derivatives. No host on the SAVRN Index sells it by the token, so $1.85 an hour is the only price reference. The paper is arXiv:2601.03233. Released January 3, 2026, last updated August 4, 2026.
Model Card
This model card focuses on the LTX-2 model, as presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. The codebase is available here. LTX-2 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2 is accessible right away via the following links: You can use the models - full, distilled, upscalers and any derivatives of the models - for purposes under the license. We recommend you use the built-in LTXVideo nodes that can be found in the ComfyUI Manager. For manual…
Excerpt from the card by LTX.io, licensed other.
Identity and Version
- Repository
- Lightricks/LTX-2
- Publisher
- LTX.io
- Task
- Image to video
- Modality
- Other
- Library
- diffusers
- Parameters
- 18.9B parameters
- Languages
- en, de, es, fr, ja, ko, zh, it
- Revision
- dfcc2108383fe1aaa0584bdf55d368a4bdadd90c
- First published
- 2026-01-03
- Last updated
- 2026-08-04
Files and Weights
69 files, 314.3 GB in total. The weights are 44 files totalling 314.3 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| audio_vae/diffusion_pytorch_model.safetensors | Weights | 106.5 MB | b36ce4066065 |
| connectors/diffusion_pytorch_model.safetensors | Weights | 2.9 GB | c7c0ad36c2d0 |
| latent_upsampler/diffusion_pytorch_model.safetensors | Weights | 995.7 MB | c56276acbffb |
| ltx-2-19b-dev-fp4.safetensors | Weights | 20.0 GB | 08a28245d841 |
| ltx-2-19b-dev-fp8.safetensors | Weights | 27.1 GB | 8a67e709b6d1 |
| ltx-2-19b-dev.safetensors | Weights | 43.3 GB | 4a51e70aad66 |
| ltx-2-19b-distilled-fp8.safetensors | Weights | 27.1 GB | 8ae14327130c |
| ltx-2-19b-distilled-lora-384.safetensors | Weights | 7.7 GB | 2718f8958200 |
| ltx-2-19b-distilled.safetensors | Weights | 43.3 GB | c4006d689061 |
| ltx-2-spatial-upscaler-x2-1.0.safetensors | Weights | 995.8 MB | 3160fabf8edf |
| ltx-2-temporal-upscaler-x2-1.0.safetensors | Weights | 262.0 MB | 9a35c2eb92f6 |
| text_encoder/diffusion_pytorch_model-00001-of-00012.safetensors | Weights | 1.7 GB | 06ffef2cbc99 |
| text_encoder/diffusion_pytorch_model-00002-of-00012.safetensors | Weights | 5.0 GB | 308270af3b7c |
| text_encoder/diffusion_pytorch_model-00003-of-00012.safetensors | Weights | 4.8 GB | 523d1b6d3ba4 |
| text_encoder/diffusion_pytorch_model-00004-of-00012.safetensors | Weights | 5.0 GB | cf426cf00fe6 |
| text_encoder/diffusion_pytorch_model-00005-of-00012.safetensors | Weights | 4.9 GB | 01c8cec1fc6d |
| text_encoder/diffusion_pytorch_model-00006-of-00012.safetensors | Weights | 5.0 GB | f1a5ec996bdd |
| text_encoder/diffusion_pytorch_model-00007-of-00012.safetensors | Weights | 4.9 GB | 5594441c5b83 |
| text_encoder/diffusion_pytorch_model-00008-of-00012.safetensors | Weights | 5.0 GB | b333d3bb4764 |
| text_encoder/diffusion_pytorch_model-00009-of-00012.safetensors | Weights | 4.9 GB | 34db39ec863e |
| text_encoder/diffusion_pytorch_model-00010-of-00012.safetensors | Weights | 5.0 GB | 0e72953188ec |
| text_encoder/diffusion_pytorch_model-00011-of-00012.safetensors | Weights | 5.0 GB | 29993bd9711e |
| text_encoder/diffusion_pytorch_model-00012-of-00012.safetensors | Weights | 589.9 MB | 19a8f0f23c87 |
| text_encoder/model-00001-of-00011.safetensors | Weights | 1.7 GB | cbc6e8132e49 |
| text_encoder/model-00002-of-00011.safetensors | Weights | 5.0 GB | b95e7ab472b8 |
| text_encoder/model-00003-of-00011.safetensors | Weights | 4.8 GB | 3731e7c18280 |
| text_encoder/model-00004-of-00011.safetensors | Weights | 5.0 GB | e9d1ce8b472f |
| text_encoder/model-00005-of-00011.safetensors | Weights | 4.9 GB | cb478659a67b |
| text_encoder/model-00006-of-00011.safetensors | Weights | 5.0 GB | a190581d8719 |
| text_encoder/model-00007-of-00011.safetensors | Weights | 4.9 GB | c347de789ff3 |
| text_encoder/model-00008-of-00011.safetensors | Weights | 5.0 GB | 2ec7525b89b0 |
| text_encoder/model-00009-of-00011.safetensors | Weights | 4.9 GB | 2b0117ecf1d8 |
| text_encoder/model-00010-of-00011.safetensors | Weights | 5.0 GB | d9f665a74358 |
| text_encoder/model-00011-of-00011.safetensors | Weights | 2.7 GB | 999bf4706d4f |
| transformer/diffusion_pytorch_model-00001-of-00008.safetensors | Weights | 5.0 GB | c4cebec5231d |
| transformer/diffusion_pytorch_model-00002-of-00008.safetensors | Weights | 5.0 GB | fbbcbd973b51 |
| transformer/diffusion_pytorch_model-00003-of-00008.safetensors | Weights | 5.0 GB | 2767c94f485e |
| transformer/diffusion_pytorch_model-00004-of-00008.safetensors | Weights | 5.0 GB | 54598964c133 |
| transformer/diffusion_pytorch_model-00005-of-00008.safetensors | Weights | 5.0 GB | 2eb2b5aa625a |
| transformer/diffusion_pytorch_model-00006-of-00008.safetensors | Weights | 4.9 GB | 068c360cde8b |
| transformer/diffusion_pytorch_model-00007-of-00008.safetensors | Weights | 5.0 GB | a55908e7fde1 |
| transformer/diffusion_pytorch_model-00008-of-00008.safetensors | Weights | 2.9 GB | 8a7e2d0941dc |
| vae/diffusion_pytorch_model.safetensors | Weights | 2.4 GB | 107cc359e3c4 |
| vocoder/diffusion_pytorch_model.safetensors | Weights | 111.2 MB | 15855fc59233 |
| audio_vae/config.json | Configuration | 505 B | — |
| connectors/config.json | Configuration | 649 B | — |
| latent_upsampler/config.json | Configuration | 266 B | — |
| model_index.json | Configuration | 616 B | — |
| scheduler/scheduler_config.json | Configuration | 487 B | — |
| text_encoder/config.json | Configuration | 3.0 KB | — |
| text_encoder/diffusion_pytorch_model.safetensors.index.json | Configuration | 150.0 KB | — |
| text_encoder/generation_config.json | Configuration | 168 B | — |
| text_encoder/model.safetensors.index.json | Configuration | 108.6 KB | — |
| tokenizer/added_tokens.json | Configuration | 35 B | — |
| tokenizer/preprocessor_config.json | Configuration | 570 B | — |
| tokenizer/processor_config.json | Configuration | 70 B | — |
| tokenizer/special_tokens_map.json | Configuration | 662 B | — |
| transformer/config.json | Configuration | 1.1 KB | — |
| transformer/diffusion_pytorch_model.safetensors.index.json | Configuration | 377.7 KB | — |
| vae/config.json | Configuration | 1.3 KB | — |
| vocoder/config.json | Configuration | 544 B | — |
| LICENSE | Documentation | 21.5 KB | — |
| README.md | Documentation | 9.6 KB | — |
| ltx-2-running-local.mp4 | Other | 15.7 MB | 8c0bde52079d |
| tokenizer/chat_template.jinja | Other | 1.5 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer/tokenizer.json | Tokenizer | 33.4 MB | 4667f2089529 |
| tokenizer/tokenizer.model | Tokenizer | 4.7 MB | 1299c11d7cf6 |
| tokenizer/tokenizer_config.json | Tokenizer | 1.2 MB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 314.3 GB
Released by LTX.io through its official repository on Hugging Face. Read the license.
Built From
- Described by arXiv:2601.03233
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 314.3 GB |
| 16-bit | 37.8 GB |
| 8-bit | 18.9 GB |
| 4-bit | 9.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About LTX-2
How much GPU memory does LTX-2 need?
About 45.3 GB at 16-bit and 11.3 GB at 4-bit: the weights (18.9B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run LTX-2 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What license is LTX-2 released under?
other, as its publisher declares it. Read the license text before commercial use.
Similar Models
Repackaged model files for ComfyUI. - https://huggingface.co/nvidia/ChronoEdit-14B-Diffusers - https://huggingface.co/Wan-AI/Wan2.2-Animate-14B - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control - https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control - https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP - https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B - https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B - https://huggingface.co/Wan-AI/Wan2.2-S2V-14B - https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B - https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B…
LTX-2.5 is an open world model with open weights, built for local execution and fine-tuning. Its established use is generating synchronized, high-fidelity video and audio from text, image, and video inputs; applicability to emerging domains such as robotics and physical AI is developing. Full control and customization — self-host on your own infrastructure. No per-generation billing, no per-seat lock-in, no forced API dependency. Revenue is measured across the whole entity, including subsidiaries and affiliates under common control. The full, binding terms live in LICENSE. - Native multishot generation — generate connected scenes in a single pass: multiple shots that hold character…
Please check our repository or the LightX2V MiniMax-H3 examples to reproduce the results. Please check the model specifications for more details. Try the MiniMax-H3 Turbo LoRA directly in LightX2V Studio: The Studio currently uses the FL2V 8-step v1.0 768p LoRA, which provides improved video and audio generation quality with 8-step inference. Integrate MiniMax-H3 Turbo into your application through the LightX2V API
This model card focuses on the LTX-2.3 model, which is a significant update to the LTX-2 model with improved audio and visual quality as well as enhanced prompt adherence. LTX-2 was presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. If you want to dive in right to the code - it is available here. LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2.3 is accessible right away via the API Playground. You can use the models - full, distilled, upscalers and any…
This repository (Abiray/MiniMax-H3-GGUF) provides GGUF quantized versions and necessary component files for the MiniMax H3 model. MiniMax H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. If you are looking for a smaller model with the same great quality that fits better on consumer-tier GPUs, please check out the MiniMax-H3-Pruned-GGUF repository. The pruned architecture is compressed down to 8.9 GB – 21.6 GB, bringing MiniMax H3 execution directly to consumer hardware. This…