SAVRN
Search Contact SAVRN

Open-weight model · Text to video

FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree

by FastVideo FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree

The recommended FastH3 Preview v1 checkpoint from FastVideo. It generates synchronized video and audio from text with four transformer forwards. This step-1300 model was trained with data-free DMD2 and VSA-H3 at 90% sparsity.

Parameters35B
Context
Weights147.8 GB
Licenseother
AccessOpen weights
Monthly Downloads279.5k

Runs On

What it takes to serve FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree (35B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 70.1 GB 84.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 35.0 GB 42.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 17.5 GB 21.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

The recommended FastH3 Preview v1 checkpoint from FastVideo. It generates synchronized video and audio from text with four transformer forwards. This step-1300 model was trained with data-free DMD2 and VSA-H3 at 90% sparsity. Install uv, then use the CUDA 13 / Blackwell path below. It selects FastVideo's published CUDA kernel wheel instead of compiling the kernel locally. See the for other platforms. The tested defaults use four B200 GPUs and the trained four-forward schedule. On other multi-GPU CUDA systems, follow the installation guide and add --no-replicated-dit --vsa-kernel triton --no-fa4. The GPU count must divide H3's 56 attention heads. This preview supports text-to-audio-video…

Excerpt from the card by FastVideo, licensed other.

Identity and Version

Repository
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
Publisher
FastVideo
Task
Text to video
Modality
Video
Library
diffusers
Parameters
35B parameters
Languages
few-step
Revision
5ea076f35b84da4c3c82217112fa733d8eea2ae1
First published
2026-08-27
Last updated
2026-09-04

Files and Weights

67 files, 147.9 GB in total. The weights are 32 files totalling 147.8 GB in safetensors.

Weights32 files · 147.8 GB
Configuration20 files · 276.3 KB
Tokenizer12 files · 34.5 MB
Documentation2 files · 21.0 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
audio_vae/diffusion_pytorch_model.safetensorsWeights605.4 MB 52c59e67ba8d
text_encoder/model-00001-of-00014.safetensorsWeights4.9 GB 6b9dfbc930e5
text_encoder/model-00002-of-00014.safetensorsWeights4.9 GB d8bb44b4ff30
text_encoder/model-00003-of-00014.safetensorsWeights4.9 GB 54f22e8b3168
text_encoder/model-00004-of-00014.safetensorsWeights4.9 GB ad09c74d3c13
text_encoder/model-00005-of-00014.safetensorsWeights4.9 GB fc993c8a0e2a
text_encoder/model-00006-of-00014.safetensorsWeights4.9 GB 82f05620d1f7
text_encoder/model-00007-of-00014.safetensorsWeights4.9 GB fb91da8cb01f
text_encoder/model-00008-of-00014.safetensorsWeights4.9 GB 431ca56535c8
text_encoder/model-00009-of-00014.safetensorsWeights4.9 GB 3825e3f4302f
text_encoder/model-00010-of-00014.safetensorsWeights4.9 GB aded5a4d1d5e
text_encoder/model-00011-of-00014.safetensorsWeights4.9 GB 3820ffe8d8d6
text_encoder/model-00012-of-00014.safetensorsWeights4.9 GB 05ad2d08ce71
text_encoder/model-00013-of-00014.safetensorsWeights4.9 GB b64f22898712
text_encoder/model-00014-of-00014.safetensorsWeights3.3 GB e45b6c9998c7
transformer/diffusion_pytorch_model-00001-of-00014.safetensorsWeights5.3 GB d4563741673a
transformer/diffusion_pytorch_model-00002-of-00014.safetensorsWeights5.3 GB 8920795048b2
transformer/diffusion_pytorch_model-00003-of-00014.safetensorsWeights5.3 GB 15b4102c4bf1
transformer/diffusion_pytorch_model-00004-of-00014.safetensorsWeights4.9 GB 793ce754a418
transformer/diffusion_pytorch_model-00005-of-00014.safetensorsWeights5.3 GB b91266bb970e
transformer/diffusion_pytorch_model-00006-of-00014.safetensorsWeights5.2 GB 219cf3f31b51
transformer/diffusion_pytorch_model-00007-of-00014.safetensorsWeights5.3 GB 583020d48bf9
transformer/diffusion_pytorch_model-00008-of-00014.safetensorsWeights5.3 GB 7edf1065d51a
transformer/diffusion_pytorch_model-00009-of-00014.safetensorsWeights4.9 GB dc9a291ef15c
transformer/diffusion_pytorch_model-00010-of-00014.safetensorsWeights5.3 GB 10c1b8173929
transformer/diffusion_pytorch_model-00011-of-00014.safetensorsWeights5.2 GB ba7a12ac11e6
transformer/diffusion_pytorch_model-00012-of-00014.safetensorsWeights5.3 GB cea2d272d196
transformer/diffusion_pytorch_model-00013-of-00014.safetensorsWeights5.3 GB efbbe1952958
transformer/diffusion_pytorch_model-00014-of-00014.safetensorsWeights2.1 GB c4c204080a1f
vae/diffusion_pytorch_model-00001-of-00003.safetensorsWeights5.1 GB 72f4c6be84ac
vae/diffusion_pytorch_model-00002-of-00003.safetensorsWeights5.0 GB 2e05e8bc23fa
vae/diffusion_pytorch_model-00003-of-00003.safetensorsWeights398.5 MB c05d6ac4b1a3
audio_scheduler/scheduler_config.jsonConfiguration96 B
audio_vae/config.jsonConfiguration2.3 KB
checkpoint_content.jsonConfiguration3.8 KB
checkpoint_metadata.jsonConfiguration6.2 KB
fastvideo_inference.jsonConfiguration859 B
modular_model_index.jsonConfiguration3.2 KB
processor/chat_template.jsonConfiguration5.5 KB
processor/preprocessor_config.jsonConfiguration390 B
processor/video_preprocessor_config.jsonConfiguration385 B
provenance.jsonConfiguration986 B
scheduler/scheduler_config.jsonConfiguration97 B
text_encoder/chat_template.jsonConfiguration5.5 KB
text_encoder/config.jsonConfiguration1.5 KB
text_encoder/model.safetensors.index.jsonConfiguration97.8 KB
text_encoder/preprocessor_config.jsonConfiguration390 B
text_encoder/video_preprocessor_config.jsonConfiguration385 B
transformer/config.jsonConfiguration546 B
transformer/diffusion_pytorch_model.safetensors.index.jsonConfiguration70.1 KB
vae/config.jsonConfiguration2.0 KB
vae/diffusion_pytorch_model.safetensors.index.jsonConfiguration74.2 KB
LICENSEDocumentation17.6 KB
README.mdDocumentation3.4 KB
.gitattributesRepository1.5 KB
processor/merges.txtTokenizer1.7 MB
processor/tokenizer.jsonTokenizer7.0 MB
processor/tokenizer_config.jsonTokenizer11.0 KB
processor/vocab.jsonTokenizer2.8 MB
text_encoder/merges.txtTokenizer1.7 MB
text_encoder/tokenizer.jsonTokenizer7.0 MB
text_encoder/tokenizer_config.jsonTokenizer11.0 KB
text_encoder/vocab.jsonTokenizer2.8 MB
tokenizer/merges.txtTokenizer1.7 MB
tokenizer/tokenizer.jsonTokenizer7.0 MB
tokenizer/tokenizer_config.jsonTokenizer11.0 KB
tokenizer/vocab.jsonTokenizer2.8 MB

License and Download

License
other
Access
Open weights, no gate
Download size
147.8 GB
Download from FastVideo

Released by FastVideo through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published147.8 GB
16-bit70.1 GB
8-bit35.0 GB
4-bit17.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree

How much GPU memory does FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree need?

About 84.1 GB at 16-bit and 21 GB at 4-bit: the weights (35B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Text to video

LTX-2.5-Diffusers

LTX.io

LTX-2.5 is an open world model with open weights, built for local execution and fine-tuning. Its established use is generating synchronized, high-fidelity video and audio from text, image, and video inputs; applicability to emerging domains such as robotics and physical AI is developing. Full control and customization — self-host on your own infrastructure. No per-generation billing, no per-seat lock-in, no forced API dependency. Revenue is measured across the whole entity, including subsidiaries and affiliates under common control. The full, binding terms live in LICENSE. Encoding always uses vae/, and LTX2Pipeline decodes with vae/ too. The diffusion decoder is a diffusion model in its…

Access requested at publisher other 19B parameters diffusers

We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations: This repository contains our T2V-A14B model, which supports generating 5s videos at both 480P and 720P resolutions. Built with a Mixture-of-Experts (MoE) architecture, it delivers outstanding video generation quality. On our new benchmark Wan-Bench 2.0, the model surpasses leading commercial models across most key evaluation dimensions. Your browser does not support the video tag. If your research or project builds upon Wan2.1 or Wan2.2, we welcome you to share it with us so we can highlight it for the broader community. - Wan2.2…

Open weights apache-2.0 14.3B parameters diffusers

Model · Text to video

Wan2.2-T2V-A14B-Diffusers

Wan-AI

We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations: This repository contains our T2V-A14B model, which supports generating 5s videos at both 480P and 720P resolutions. Built with a Mixture-of-Experts (MoE) architecture, it delivers outstanding video generation quality. On our new benchmark Wan-Bench 2.0, the model surpasses leading commercial models across most key evaluation dimensions. Your browser does not support the video tag. If your research or project builds upon Wan2.1 or Wan2.2, we welcome you to share it with us so we can highlight it for the broader community. - Wan2.2…

Open weights apache-2.0 14.3B parameters diffusers

Model · Text to video

Wan2.1-T2V-14B-Diffusers

Wan-AI

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: This repository features our T2V-14B model, which establishes a new SOTA performance benchmark among both open-source and closed-source models. It demonstrates exceptional capabilities in generating high-quality visuals with significant motion dynamics. It is also the only video model capable of producing both Chinese and English text and supports video generation at both 480P and 720P resolutions. Your browser does not support the video tag. - Wan2.1 Text-to-Video - [x] Multi-GPU Inference code of the 14B and 1.3B…

Open weights apache-2.0 14.3B parameters diffusers

Model · Text to video

Wan2.1-T2V-14B

Wan-AI

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: This repository features our T2V-14B model, which establishes a new SOTA performance benchmark among both open-source and closed-source models. It demonstrates exceptional capabilities in generating high-quality visuals with significant motion dynamics. It is also the only video model capable of producing both Chinese and English text and supports video generation at both 480P and 720P resolutions. Your browser does not support the video tag. - Wan2.1 Text-to-Video - [x] Multi-GPU Inference code of the 14B and 1.3B…

Open weights apache-2.0 14.3B parameters diffusers

Model · Text to video

Wan2.2-TI2V-5B-Diffusers

Wan-AI

We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations: This repository contains our TI2V-5B model, built with the advanced Wan2.2-VAE that achieves a compression ratio of 16×16×4. This model supports both text-to-video and image-to-video generation at 720P resolution with 24fps and can runs on single consumer-grade GPU such as the 4090. It is one of the fastest 720P@24fps models available, meeting the needs of both industrial applications and academic research. Your browser does not support the video tag. If your research or project builds upon Wan2.1 or Wan2.2, we welcome you to share it…

Open weights apache-2.0 5B parameters diffusers