This GGUF file is a direct conversion of Wan-AI/Wan2.2-T2V-A14B Since this is a quantized model, all original licensing terms and usage restrictions remain in effect. Usage The model can be used with the ComfyUI custom node ComfyUI-GGUF by city96 Place model files in ComfyUI/models/unet see the GitHub readme for further installation instructions.
Open-weight model · Text to video
minimax-h3-nvfp4-convrot
by rockerBOO rockerBOO/minimax-h3-nvfp4-convrot
Quantized diffusion transformers for MiniMax H3, a 33B omni-modal video+audio generator, built from the ComfyUI repack at Comfy-Org/MiniMax-H3. Output is 768p / 24 fps / 4–15 s with synchronized 32 kHz stereo audio.
Model Card
Quantized diffusion transformers for MiniMax H3, a 33B omni-modal video+audio generator, built from the ComfyUI repack at Comfy-Org/MiniMax-H3. Output is 768p / 24 fps / 4–15 s with synchronized 32 kHz stereo audio. (2K output requires the separate H3-Regenerate-2K module, which is not part of this or Comfy-Org's release.) These are ComfyUI single-file checkpoints, not diffusers models. Filenames follow minimaxh3.safetensors. - fl2va — first/last-frame mode. Zero images = text-to-video, one or two = frame-conditioned. - ref2va — omni-reference mode (up to 9 images / 3 video clips / 3 audio clips). Both get identical treatment; pick the one matching your workflow. The three INT4 variants…
Excerpt from the card by rockerBOO, licensed other.
Identity and Version
- Repository
- rockerBOO/minimax-h3-nvfp4-convrot
- Publisher
- rockerBOO
- Task
- Text to video
- Modality
- Video
- Library
- minimax-h3
- Parameters
- Not stated by the source
- Languages
- Not stated by the source
- Revision
- d06455b58c5f84bb7831536a672245d6b2b0d014
- First published
- 2026-08-03
- Last updated
- 2026-08-21
Files and Weights
15 files, 231.7 GB in total. The weights are 10 files totalling 231.7 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| minimax_h3_fl2va_int4_convrot_simple.safetensors | Weights | 25.3 GB | 068d462742d2 |
| minimax_h3_fl2va_nvfp4.safetensors | Weights | 34.4 GB | dd5c44466f0d |
| minimax_h3_fl2va_pruned_int4_convrot_simple.safetensors | Weights | 16.8 GB | 41bf98f26018 |
| minimax_h3_fl2va_pruned_mixed_int4_int8_convrot_simple.safetensors | Weights | 20.3 GB | 89c0c93aacf3 |
| minimax_h3_fl2va_pruned_nvfp4.safetensors | Weights | 20.1 GB | 8111a496dae7 |
| minimax_h3_fl2va_pruned_nvfp4_convrot_int8.safetensors | Weights | 20.1 GB | 41e9f92df81a |
| minimax_h3_fl2va_pruned_nvfp4_fp8.safetensors | Weights | 20.1 GB | a9356b900b8f |
| minimax_h3_ref2va_nvfp4.safetensors | Weights | 34.4 GB | e25881cb906f |
| minimax_h3_ref2va_pruned_nvfp4.safetensors | Weights | 20.1 GB | 672c1be80533 |
| minimax_h3_ref2va_pruned_nvfp4_convrot_int8.safetensors | Weights | 20.1 GB | 9593074fc37e |
| minimax_h3_layer_config.json | Configuration | 531 B | — |
| minimax_h3_layer_config_int4_convrot.json | Configuration | 545 B | — |
| LICENSE | Documentation | 17.6 KB | — |
| README.md | Documentation | 13.8 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 231.7 GB
Released by rockerBOO through its official repository on Hugging Face. Read the license.
Built From
- Derived from MiniMaxAI/MiniMax-H3
- Described by arXiv:2512.03673
- Quantized from MiniMaxAI/MiniMax-H3
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 231.7 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About minimax-h3-nvfp4-convrot
What license is minimax-h3-nvfp4-convrot released under?
other, as its publisher declares it. Read the license text before commercial use.
Similar Models
A LoRA for MiniMax-H3 that renders joint video + synchronized stereo audio in as few as 4 sampling steps instead of the usual ~20 — a ~5× sampling speedup — and keeps getting better as you add steps. For most work, use minimaxh3turbov4step600ema.safetensors. It's the markedly better micro-detail (faces, fingers, fine texture), and the over-sharpening / plastic look of the earlier v1 (~850) line is fully resolved. v4 introduced a static-frame enhancement — a big win for static and small-motion content. The one trade-off shows up only at 4 steps with large, fast motion, where v4 can produce motion-smear / trailing ghosting (we're actively fixing this). Two things address it: - Use 6–8 steps.…
This repository contains MiniMax-H3 Turbo LoRAs converted and optimized for ComfyUI: These LoRAs accelerate MiniMax-H3 video and synchronized-audio generation by reducing the required number of sampling steps. Newly added LoRA, located in the experimental/ folder: Manual recommended sigmas: 3-step 1.0, 0.961165, 0.853333, 0.0 4-step 1.0, 0.970874, 0.907249, 0.640000, 0.0 Three LoRAs extracted from VDN-H3 8 step: The main 8-step LoRA works on both FL2VA and Ref2VA. If you are running a pruned base, choose the pruned version that corresponds to your base — the pruned versions need their own matching pruned base. Three dynamically resized BF16 LoRAs are now included. Their source weights were…
Sulphur 2 An uncensored video generation model based on LTX 2.3 supporting both t2v and i2v natively, as well as all of the other ltx 2.3 formats. Follow us on X Join our Discord Support the next version of the project, even just a few dollars would go a long way: Kofi To get started with the model, I recommend downloading either of the dev versions, (fp8mixed or bf16) and downloading the distill lora provided. By the way, I'm aware the workflows contain sulphurfinal right now, just use the lora or use the full models, don't use both at the same time. This model contains a prompt enhancer. The easiest way to get started with the prompt enhancer is by using it on lmstudio. The way to…
Wan2.1 is an open-source suite of video foundation models, compatible with consumer-grade GPUs, that excels in various video generation tasks like text-to-video, image-to-video, and video editing, even supporting visual text generation. Download models using huggingface-cli: You can also download directly from this page. This model is a derivative work of the original model licensed under the Apache 2.0 License, and is therefore distributed under the terms of the same license. Thanks to Patrick Gillespie for creating the ASCII text art tool used in this project https://patorjk.com/software/taag/ Wan-AI for the Wan model https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B…
Wan2.1 is an open-source suite of video foundation models, compatible with consumer-grade GPUs, that excels in various video generation tasks like text-to-video, image-to-video, and video editing, even supporting visual text generation. Download models using huggingface-cli: You can also download directly from this page. This model is a derivative work of the original model licensed under the Apache 2.0 License, and is therefore distributed under the terms of the same license. Thanks to Patrick Gillespie for creating the ASCII text art tool used in this project https://patorjk.com/software/taag/ Wan-AI for the Wan model https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B…