Instructions further below. GGUF for MiniMax-H3, compatible on most platforms including stablediffusion.cpp and Unsloth. You can run MiniMax-H3 via Unsloth: https://github.com/unslothai/unsloth/ GGUF quantizations of MiniMaxAI/MiniMax-H3 MiniMax H3 is an omni-modal generative system that produces video with native stereo audio, up to 15 seconds at 24 FPS with 32 kHz stereo audio. Both halves of the runtime are in this repo: the denoisers and the Qwen3-VL text encoder they need. H3 ships two denoisers, and which one you load decides what the model can be given: - fl2vapruned, the H3-Base first-and-last-frame variant. Text, plus zero, one or two frames. - ref2vapruned, the reference variant.…
Open weights
other
gguf
Model · Image text to video
Jay
This repository (Abiray/MiniMax-H3-Pruned-GGUF) provides pruned and quantized GGUF weights for the MiniMax H3 omni-modal generative model. MiniMax H3 is designed for unified multimodal context processing, capable of generating synchronized high-definition video and 32 kHz stereo audio from text, image, audio, and video inputs. 1. Download your desired.gguf variant from the table above. 2. Place the downloaded.gguf file into the ComfyUI/models/unet/ directory. 3. In your ComfyUI workflow, load the model using the UnetLoaderGGUF node. MiniMax H3 is released under the MiniMax H3 Community License Agreement. Please refer to the MiniMaxAI/MiniMax-H3 repository and the repository's LICENSE file…
Open weights
other
Model · Image text to video
Fal
A LoRA adapter for MiniMax H3 specialized in realistic people: faces that hold up in close-up, natural skin texture, believable expressions and gestures, film-style lighting and documentary camera movement. Same prompt, same seed — base model on the left, this adapter on the right: 19 pairs, same prompt, same seed, adapter on vs off — the only variable is the LoRA. The trigger word is present on both sides, so it is not doing the work. Each pair plays the base model first, then freezes and dims while the adapted version plays beside it. Close-up talking faces, arguments, several people speaking at once, weathered skin, children, ritual and travel scenes. Nothing cherry-picked from a larger…
Open weights
other
minimax-h3
Unmodified 4-step and 8-step DaSiWa MiniMax H3 Hybrid SafeTensors checkpoints mirrored for BRP Canvas downloads. The hybrid checkpoint supports text-to-video, reference-to-video, and first/last-frame-to-video through the same ComfyUI workflow. This is not an official MiniMax or DaSiWa distribution. Review the original model page and the included upstream MiniMax license before use. BRP Canvas defaults to shift video 9 and shift audio 4 for both distilled checkpoints. - 4-step: 56c52c7890c105308d28fe9c25c25fdb80e6cd6a54e2604d8af71732ba4ed74f - 8-step: e0441d26414f6e0c28f43d580e6cc56fad424da0fa4d261b698ca73188aa6332
Open weights
other
minimax-h3
Model · Image text to video
MiniMax
Offical skills to improve prompt writing: skills on github Use MiniMax\-H3 directly via API\. Use MiniMax\-H3 directly via App\. MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output…
Open weights
other
33.1B parameters
minimax-h3