SAVRN
Search Contact SAVRN

Open-weight model · Text to video

minimax-h3-nvfp4-convrot

by rockerBOO rockerBOO/minimax-h3-nvfp4-convrot

Quantized diffusion transformers for MiniMax H3, a 33B omni-modal video+audio generator, built from the ComfyUI repack at Comfy-Org/MiniMax-H3. Output is 768p / 24 fps / 4–15 s with synchronized 32 kHz stereo audio.

Parameters
Context
Weights231.7 GB
Licenseother
AccessOpen weights
Monthly Downloads17.8k

Model Card

Quantized diffusion transformers for MiniMax H3, a 33B omni-modal video+audio generator, built from the ComfyUI repack at Comfy-Org/MiniMax-H3. Output is 768p / 24 fps / 4–15 s with synchronized 32 kHz stereo audio. (2K output requires the separate H3-Regenerate-2K module, which is not part of this or Comfy-Org's release.) These are ComfyUI single-file checkpoints, not diffusers models. Filenames follow minimaxh3.safetensors. - fl2va — first/last-frame mode. Zero images = text-to-video, one or two = frame-conditioned. - ref2va — omni-reference mode (up to 9 images / 3 video clips / 3 audio clips). Both get identical treatment; pick the one matching your workflow. The three INT4 variants…

Excerpt from the card by rockerBOO, licensed other.

Identity and Version

Repository
rockerBOO/minimax-h3-nvfp4-convrot
Publisher
rockerBOO
Task
Text to video
Modality
Video
Library
minimax-h3
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
d06455b58c5f84bb7831536a672245d6b2b0d014
First published
2026-08-03
Last updated
2026-08-21

Files and Weights

15 files, 231.7 GB in total. The weights are 10 files totalling 231.7 GB in safetensors.

Weights10 files · 231.7 GB
Configuration2 files · 1.1 KB
Documentation2 files · 31.4 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
minimax_h3_fl2va_int4_convrot_simple.safetensorsWeights25.3 GB 068d462742d2
minimax_h3_fl2va_nvfp4.safetensorsWeights34.4 GB dd5c44466f0d
minimax_h3_fl2va_pruned_int4_convrot_simple.safetensorsWeights16.8 GB 41bf98f26018
minimax_h3_fl2va_pruned_mixed_int4_int8_convrot_simple.safetensorsWeights20.3 GB 89c0c93aacf3
minimax_h3_fl2va_pruned_nvfp4.safetensorsWeights20.1 GB 8111a496dae7
minimax_h3_fl2va_pruned_nvfp4_convrot_int8.safetensorsWeights20.1 GB 41e9f92df81a
minimax_h3_fl2va_pruned_nvfp4_fp8.safetensorsWeights20.1 GB a9356b900b8f
minimax_h3_ref2va_nvfp4.safetensorsWeights34.4 GB e25881cb906f
minimax_h3_ref2va_pruned_nvfp4.safetensorsWeights20.1 GB 672c1be80533
minimax_h3_ref2va_pruned_nvfp4_convrot_int8.safetensorsWeights20.1 GB 9593074fc37e
minimax_h3_layer_config.jsonConfiguration531 B
minimax_h3_layer_config_int4_convrot.jsonConfiguration545 B
LICENSEDocumentation17.6 KB
README.mdDocumentation13.8 KB
.gitattributesRepository1.5 KB

License and Download

License
other
Access
Open weights, no gate
Download size
231.7 GB
Download from rockerBOO

Released by rockerBOO through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published231.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About minimax-h3-nvfp4-convrot

What license is minimax-h3-nvfp4-convrot released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Text to video

Wan2.2-T2V-A14B-GGUF

QuantStack

This GGUF file is a direct conversion of Wan-AI/Wan2.2-T2V-A14B Since this is a quantized model, all original licensing terms and usage restrictions remain in effect. Usage The model can be used with the ComfyUI custom node ComfyUI-GGUF by city96 Place model files in ComfyUI/models/unet see the GitHub readme for further installation instructions.

Open weights apache-2.0 gguf

Model · Text to video

MiniMax-H3-Turbo-Lora

Larryvrh

A LoRA for MiniMax-H3 that renders joint video + synchronized stereo audio in as few as 4 sampling steps instead of the usual ~20 — a ~5× sampling speedup — and keeps getting better as you add steps. For most work, use minimaxh3turbov4step600ema.safetensors. It's the markedly better micro-detail (faces, fingers, fine texture), and the over-sharpening / plastic look of the earlier v1 (~850) line is fully resolved. v4 introduced a static-frame enhancement — a big win for static and small-motion content. The one trade-off shows up only at 4 steps with large, fast motion, where v4 can produce motion-smear / trailing ghosting (we're actively fixing this). Two things address it: - Use 6–8 steps.…

Open weights apache-2.0 minimax-h3

Model · Text to video

MiniMax-H3-Turbo-Lora-ComfyUI

DRBAPH

This repository contains MiniMax-H3 Turbo LoRAs converted and optimized for ComfyUI: These LoRAs accelerate MiniMax-H3 video and synchronized-audio generation by reducing the required number of sampling steps. Newly added LoRA, located in the experimental/ folder: Manual recommended sigmas: 3-step 1.0, 0.961165, 0.853333, 0.0 4-step 1.0, 0.970874, 0.907249, 0.640000, 0.0 Three LoRAs extracted from VDN-H3 8 step: The main 8-step LoRA works on both FL2VA and Ref2VA. If you are running a pruned base, choose the pruned version that corresponds to your base — the pruned versions need their own matching pruned base. Three dynamically resized BF16 LoRAs are now included. Their source weights were…

Open weights apache-2.0 minimax-h3

Model · Text to video

Sulphur-2-base

Sulphur

Sulphur 2 An uncensored video generation model based on LTX 2.3 supporting both t2v and i2v natively, as well as all of the other ltx 2.3 formats. Follow us on X Join our Discord Support the next version of the project, even just a few dollars would go a long way: Kofi To get started with the model, I recommend downloading either of the dev versions, (fp8mixed or bf16) and downloading the distill lora provided. By the way, I'm aware the workflows contain sulphurfinal right now, just use the lora or use the full models, don't use both at the same time. This model contains a prompt enhancer. The easiest way to get started with the prompt enhancer is by using it on lmstudio. The way to…

Open weights diffusers

Model · Text to video

Wan2.1-VACE-1.3B-GGUF

Sam

Wan2.1 is an open-source suite of video foundation models, compatible with consumer-grade GPUs, that excels in various video generation tasks like text-to-video, image-to-video, and video editing, even supporting visual text generation. Download models using huggingface-cli: You can also download directly from this page. This model is a derivative work of the original model licensed under the Apache 2.0 License, and is therefore distributed under the terms of the same license. Thanks to Patrick Gillespie for creating the ASCII text art tool used in this project https://patorjk.com/software/taag/ Wan-AI for the Wan model https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B…

Open weights apache-2.0 diffusers

Model · Text to video

Wan2.1-T2V-1.3B-GGUF

Sam

Wan2.1 is an open-source suite of video foundation models, compatible with consumer-grade GPUs, that excels in various video generation tasks like text-to-video, image-to-video, and video editing, even supporting visual text generation. Download models using huggingface-cli: You can also download directly from this page. This model is a derivative work of the original model licensed under the Apache 2.0 License, and is therefore distributed under the terms of the same license. Thanks to Patrick Gillespie for creating the ASCII text art tool used in this project https://patorjk.com/software/taag/ Wan-AI for the Wan model https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B…

Open weights apache-2.0 diffusers