SAVRN
Search Contact SAVRN

Open-weight model · Image to video

Minimax-h3_Singularity

by AIGC Singularity WarmBloodAban/Minimax-h3_Singularity

Minimax-h3Singularity is a comprehensive fine-tuned fusion model specialized in enhancing the capabilities of MiniMax-H3.

Parameters
Context
Weights66.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads181.8k

Model Card

By AIGC Singularity, published under apache-2.0, revision 7aa89f2bcb53.

Minimax-h3Singularity is a comprehensive fine-tuned fusion model specialized in enhancing the capabilities of MiniMax-H3. Designed as a versatile multimodal video generation model, it natively supports Text-to-Video (T2V), Image-to-Video (I2V), Reference-to-Video (Ref2V), and Video-to-Video (V2V) workflows within ComfyUI. Built upon a strategic fusion of key checkpoints (including ref, fl, b25-49, etc.), this model underwent deep high-step fine-tuning. To preserve the original model's foundational strengths and broad generalization while solving artifacts introduced by high-step training, we spent 3 full days on precise model pruning and weight optimization. The result is a clean, sharp…

Read AIGC Singularity's full model card

Model Overview

Minimax-h3_Singularity is a comprehensive fine-tuned fusion model specialized in enhancing the capabilities of MiniMax-H3. Designed as a versatile multimodal video generation model, it natively supports Text-to-Video (T2V), Image-to-Video (I2V), Reference-to-Video (Ref2V), and Video-to-Video (V2V) workflows within ComfyUI.

Built upon a strategic fusion of key checkpoints (including ref, fl, b25-49, etc.), this model underwent deep high-step fine-tuning. To preserve the original model's foundational strengths and broad generalization while solving artifacts introduced by high-step training, we spent 3 full days on precise model pruning and weight optimization. The result is a clean, sharp, and highly dynamic video generation model.


Key Improvements & Features

  • HDR Image Quality & Blur Reduction: Fine-tuned on high-dynamic-range (HDR) video datasets to significantly enhance visual clarity and eliminate motion blur during high-speed action.
  • Distant Face Restoration: Drastically reduces facial distortion, blurriness, and collapsing in medium-to-long shots.
  • Clean & De-Oiled Aesthetic: Removes heavy, unnatural skin shine and glossy textures, rendering natural lighting and photorealistic materials.
  • Enhanced Dynamic Motion: Boosts motion fluidity and physical impact, excels in complex action sequences such as sword fighting and martial arts/melee combat.
  • VFX & Fantasy Effects: Specifically optimized for fantasy spellcasting, particle aura, and magical combat visual effects.
  • Expressive Facial Dynamics: Captures subtle facial expressions and emotional nuances more vividly.
  • Cinematography & Camera Control: Strengthens responsiveness to camera movements (pan, tilt, zoom, tracking shots) for cinematic storytelling.
  • Full Base Capability Retention: 100% preserves MiniMax-H3's original prompt adherence, style adaptability, and base multimodal generation strength.

Showcase


Usage Guide

Multimodal Pipeline Support

This model is fully compatible with ComfyUI and supports: * Text-to-Video (T2V) * Image-to-Video (I2V) * Reference-to-Video (Ref2V) * Video-to-Video (V2V)

Recommended Acceleration LoRA

For high-speed generation with minimal quality loss, we strongly recommend pairing with: * minimax_h3_ref2v_turbo_4step_v0.1 (Enables 4-step fast inference)

Online Interactive Demo

Test the model directly in your browser without local GPU setup: Try it on RunningHub Workflows


Acknowledgements

Special thanks to the MiniMax open-source team for creating and releasing the powerful MiniMax-H3multimodal video model, providing a solid foundation for the open-source community!


Community & Commercial Inquiries

Feel free to connect for tutorials, community discussions, workflow sharing, or commercial collaborations:

  • YouTube Channel: AIGC-Singularity
  • Bilibili Channel: AIGC-Singularity Space
  • QQ Group 1: 1058747239 (Request to join)
  • QQ Group 2: 1072010342 (Request to join)
  • Business Inquiries (WeChat): aigctyd
  • Email: [email protected]

Identity and Version

Repository
WarmBloodAban/Minimax-h3_Singularity
Publisher
AIGC Singularity
Task
Image to video
Modality
Other
Library
minimax-h3
Parameters
Not stated by the source
Languages
en, zh
Revision
7aa89f2bcb53c426428acdeba17d53c513393609
First published
2026-09-05
Last updated
2026-09-17

Files and Weights

14 files, 66.8 GB in total. The weights are 3 files totalling 66.7 GB in safetensors.

Weights3 files · 66.7 GB
Documentation2 files · 24.7 KB
Other7 files · 64.1 MB
Repository2 files · 2.1 KB
Every file
FileTypeSizeSHA-256
Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8.safetensorsWeights21.0 GB 412a7b126595
Minimax-h3_Singularity_ref2va_v1.3_int8.safetensorsWeights34.0 GB 551915097b87
minimax_h3_ref2va_pruned_w4a8_mixed.safetensorsWeights11.8 GB de2c6c29c4ee
MiniMax_H3_Singularity_Prompt_Writing_Specification_Enhanced_EN.mdDocumentation19.8 KB
README.mdDocumentation4.9 KB
video/1.mp4Other10.9 MB 48e1856ff43c
video/2.mp4Other8.5 MB 005c07600e9a
video/3.mp4Other12.3 MB 4569d372d414
video/AIGCTYD2_00001_p84-audio_ganvc_1788634594 (1).mp4Other7.2 MB 8fa0174b2f05
video/AIGCTYD2_00001_p85-audio_jrski_1788643962 (1).mp4Other12.2 MB 5da1f0ac234d
video/AIGCTYD2_00003_p82-audio_aqukl_1788646377 (29).mp4Other8.3 MB ff231c3b0bc4
video/AIGCTYD2_00010_p87-audio_alipa_1788652671.mp4Other4.6 MB fe63d4094633
.gitattributesRepository2.1 KB
video/.gitkeepRepository

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
66.7 GB
Download from AIGC Singularity

Released by AIGC Singularity through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published66.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Minimax-h3_Singularity

Can I use Minimax-h3_Singularity commercially?

Yes. Minimax-h3_Singularity is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image to video

Wan_2.2_ComfyUI_Repackaged

Comfy Org

Repackaged model files for ComfyUI. - https://huggingface.co/nvidia/ChronoEdit-14B-Diffusers - https://huggingface.co/Wan-AI/Wan2.2-Animate-14B - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control - https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control - https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP - https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B - https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B - https://huggingface.co/Wan-AI/Wan2.2-S2V-14B - https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B - https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B…

Open weights apache-2.0 diffusion-single-file

Model · Image to video

LTX-2.5

LTX.io

LTX-2.5 is an open world model with open weights, built for local execution and fine-tuning. Its established use is generating synchronized, high-fidelity video and audio from text, image, and video inputs; applicability to emerging domains such as robotics and physical AI is developing. Full control and customization — self-host on your own infrastructure. No per-generation billing, no per-seat lock-in, no forced API dependency. Revenue is measured across the whole entity, including subsidiaries and affiliates under common control. The full, binding terms live in LICENSE. - Native multishot generation — generate connected scenes in a single pass: multiple shots that hold character…

Access requested at publisher other diffusion-single-file

Model · Image to video

Minimax-h3-Turbo

Lightx2v

Please check our repository or the LightX2V MiniMax-H3 examples to reproduce the results. Please check the model specifications for more details. Try the MiniMax-H3 Turbo LoRA directly in LightX2V Studio: The Studio currently uses the FL2V 8-step v1.0 768p LoRA, which provides improved video and audio generation quality with 8-step inference. Integrate MiniMax-H3 Turbo into your application through the LightX2V API

Open weights apache-2.0 diffusers

Model · Image to video

LTX-2.3

LTX.io

This model card focuses on the LTX-2.3 model, which is a significant update to the LTX-2 model with improved audio and visual quality as well as enhanced prompt adherence. LTX-2 was presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. If you want to dive in right to the code - it is available here. LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2.3 is accessible right away via the API Playground. You can use the models - full, distilled, upscalers and any…

Open weights other diffusers

Model · Image to video

MiniMax-H3-GGUF

Jay

This repository (Abiray/MiniMax-H3-GGUF) provides GGUF quantized versions and necessary component files for the MiniMax H3 model. MiniMax H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. If you are looking for a smaller model with the same great quality that fits better on consumer-tier GPUs, please check out the MiniMax-H3-Pruned-GGUF repository. The pruned architecture is compressed down to 8.9 GB – 21.6 GB, bringing MiniMax H3 execution directly to consumer hardware. This…

Open weights other

Model · Image to video

LTX-2.3-fp8

LTX.io

This is the FP8 versions of the LTX-2.3 model. All information below is derived from the base model. This model card focuses on the LTX-2.3 model, which is a significant update to the LTX-2 model with improved audio and visual quality as well as enhanced prompt adherence. LTX-2 was presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. If you want to dive in right to the code - it is available here. LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2.3 is…

Open weights other diffusers