Model · Image text to video
Jay
This repository is a community-compiled collection of quantized and pruned weights for MiniMax H3 (Hailuo 3.0), optimized for local inference environments like ComfyUI. By unifying various quantization formats (INT4, INT8, Mixed, and NVFP4) into a single structured repository, this hub makes it easier for users with consumer GPUs (16GB - 24GB VRAM) to experiment with MiniMax H3's powerful omni-modal text/image/audio-to-video generation capabilities. If you are new to local generation and aren't sure what to download, use this guide based on your graphics card. Perfect for RTX 4070 Ti Super, RTX 4080, etc. Perfect for RTX 3090, RTX 4090, etc. Exclusively for RTX 5090, PRO 6000, and other…
Open weights
other
diffusers
Model · Image text to video
Jay
This repository (Abiray/MiniMax-H3-Pruned-GGUF) provides pruned and quantized GGUF weights for the MiniMax H3 omni-modal generative model. MiniMax H3 is designed for unified multimodal context processing, capable of generating synchronized high-definition video and 32 kHz stereo audio from text, image, audio, and video inputs. 1. Download your desired.gguf variant from the table above. 2. Place the downloaded.gguf file into the ComfyUI/models/unet/ directory. 3. In your ComfyUI workflow, load the model using the UnetLoaderGGUF node. MiniMax H3 is released under the MiniMax H3 Community License Agreement. Please refer to the MiniMaxAI/MiniMax-H3 repository and the repository's LICENSE file…
Open weights
other
Model · Image text to video
Fal
A LoRA adapter for MiniMax H3 specialized in realistic people: faces that hold up in close-up, natural skin texture, believable expressions and gestures, film-style lighting and documentary camera movement. Same prompt, same seed — base model on the left, this adapter on the right: 19 pairs, same prompt, same seed, adapter on vs off — the only variable is the LoRA. The trigger word is present on both sides, so it is not doing the work. Each pair plays the base model first, then freezes and dims while the adapted version plays beside it. Close-up talking faces, arguments, several people speaking at once, weathered skin, children, ritual and travel scenes. Nothing cherry-picked from a larger…
Open weights
other
minimax-h3
Unmodified 4-step and 8-step DaSiWa MiniMax H3 Hybrid SafeTensors checkpoints mirrored for BRP Canvas downloads. The hybrid checkpoint supports text-to-video, reference-to-video, and first/last-frame-to-video through the same ComfyUI workflow. This is not an official MiniMax or DaSiWa distribution. Review the original model page and the included upstream MiniMax license before use. BRP Canvas defaults to shift video 9 and shift audio 4 for both distilled checkpoints. - 4-step: 56c52c7890c105308d28fe9c25c25fdb80e6cd6a54e2604d8af71732ba4ed74f - 8-step: e0441d26414f6e0c28f43d580e6cc56fad424da0fa4d261b698ca73188aa6332
Open weights
other
minimax-h3
Model · Image text to video
MiniMax
Offical skills to improve prompt writing: skills on github Use MiniMax\-H3 directly via API\. Use MiniMax\-H3 directly via App\. MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output…
Open weights
other
33.1B parameters
minimax-h3