SAVRN
Search Contact SAVRN

SAVRN Model Hub · Models by Task

Image to video Models

18 models in the SAVRN Model Hub for image to video, from publishers including LTX.io, Joey, Daniel, AIGC Singularity.

18 models.

Model · Image to video

Wan_2.2_ComfyUI_Repackaged

Comfy Org

Repackaged model files for ComfyUI. - https://huggingface.co/nvidia/ChronoEdit-14B-Diffusers - https://huggingface.co/Wan-AI/Wan2.2-Animate-14B - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control - https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control - https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP - https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP - https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B - https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B - https://huggingface.co/Wan-AI/Wan2.2-S2V-14B - https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B - https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B…

Open weights apache-2.0 diffusion-single-file

Model · Image to video

LTX-Video

LTX.io

This model card focuses on the model associated with the LTX-Video model, codebase available here. LTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-time. It produces 30 FPS videos at a 1216×704 resolution faster than they can be watched. Trained on a large-scale dataset of diverse videos, the model generates high-resolution videos with realistic and varied content. You can use the model for purposes under the license: - 2B version 0.9: license - 2B version 0.9.1 license - 2B version 0.9.5 license - 2B version 0.9.6-dev license - 2B version 0.9.6-distilled license - 13B version 0.9.7-dev license - 13B version 0.9.7-dev-fp8 license…

Open weights other 1.9B parameters diffusers

Model · Image to video

LTX-2

LTX.io

This model card focuses on the LTX-2 model, as presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. The codebase is available here. LTX-2 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2 is accessible right away via the following links: You can use the models - full, distilled, upscalers and any derivatives of the models - for purposes under the license. We recommend you use the built-in LTXVideo nodes that can be found in the ComfyUI Manager. For manual…

Open weights other 18.9B parameters diffusers

Model · Image to video

LTX-2.5

LTX.io

LTX-2.5 is an open world model with open weights, built for local execution and fine-tuning. Its established use is generating synchronized, high-fidelity video and audio from text, image, and video inputs; applicability to emerging domains such as robotics and physical AI is developing. Full control and customization — self-host on your own infrastructure. No per-generation billing, no per-seat lock-in, no forced API dependency. Revenue is measured across the whole entity, including subsidiaries and affiliates under common control. The full, binding terms live in LICENSE. - Native multishot generation — generate connected scenes in a single pass: multiple shots that hold character…

Access requested at publisher other diffusion-single-file

Model · Image to video

Minimax-h3-Turbo

Lightx2v

Please check our repository or the LightX2V MiniMax-H3 examples to reproduce the results. Please check the model specifications for more details. Try the MiniMax-H3 Turbo LoRA directly in LightX2V Studio: The Studio currently uses the FL2V 8-step v1.0 768p LoRA, which provides improved video and audio generation quality with 8-step inference. Integrate MiniMax-H3 Turbo into your application through the LightX2V API

Open weights apache-2.0 diffusers

Model · Image to video

LTX-2.3

LTX.io

This model card focuses on the LTX-2.3 model, which is a significant update to the LTX-2 model with improved audio and visual quality as well as enhanced prompt adherence. LTX-2 was presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. If you want to dive in right to the code - it is available here. LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2.3 is accessible right away via the API Playground. You can use the models - full, distilled, upscalers and any…

Open weights other diffusers

Model · Image to video

MiniMax-H3-GGUF

Jay

This repository (Abiray/MiniMax-H3-GGUF) provides GGUF quantized versions and necessary component files for the MiniMax H3 model. MiniMax H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. If you are looking for a smaller model with the same great quality that fits better on consumer-tier GPUs, please check out the MiniMax-H3-Pruned-GGUF repository. The pruned architecture is compressed down to 8.9 GB – 21.6 GB, bringing MiniMax H3 execution directly to consumer hardware. This…

Open weights other

Model · Image to video

LTX-2.3-fp8

LTX.io

This is the FP8 versions of the LTX-2.3 model. All information below is derived from the base model. This model card focuses on the LTX-2.3 model, which is a significant update to the LTX-2 model with improved audio and visual quality as well as enhanced prompt adherence. LTX-2 was presented in the paper LTX-2: Efficient Joint Audio-Visual Foundation Model. If you want to dive in right to the code - it is available here. LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution. LTX-2.3 is…

Open weights other diffusers

Model · Image to video

LTX-2.3-GGUF

Unsloth AI

This is a GGUF quantized version of LTX-2.3. unsloth/LTX-2.3-GGUF uses Unsloth Dynamic 2.0 methodology for SOTA performance. - Important layers are upcasted to higher precision. - Uses tooling from ComfyUI-GGUF by city96. There are two sets of GGUF's published. One for the dev model and one for the distilled. The distilled model is optimized for few step generation, think 4-8 steps. dev on the other hand needs more steps at least 20, but you get better outputs. The distilled variant is useful as a drafting model or a refining model. In fact the workflow published below, uses the distilled lora on top of the dev model to refine the intial output. Download the mp4 in the repo and open it with…

Open weights other ggml

Model · Image to video

MiniMax-H3-encoder-GGUF

Joey

GGUF quantizations of the Qwen3-VL-32B vision-language text encoder used by MiniMax-H3 in ComfyUI. The H3 DiT quants are here: joeygambino/MiniMax-H3-GGUF. You need one file from each repo to run H3 — the DiT alone will not generate anything. Load these encoders with H3 Clip Loader (Any) from not the stock CLIPLoaderGGUF node. The H3 text encoder is a truncated Qwen3-VL-32B - 50 layers, no final norm, no lmhead - and its vision tower ships separately as the -mmproj-F16.gguf sidecar. Stock ComfyUI-GGUF only merges an mmproj when the encoder's architecture is qwen2vl; Qwen3-VL reports qwen3vl, so the sidecar is never merged at all, and the resulting missing vision tensors surface as a…

Open weights gguf

Model · Image to video

Minimax-h3_Singularity

AIGC Singularity

Minimax-h3Singularity is a comprehensive fine-tuned fusion model specialized in enhancing the capabilities of MiniMax-H3. Designed as a versatile multimodal video generation model, it natively supports Text-to-Video (T2V), Image-to-Video (I2V), Reference-to-Video (Ref2V), and Video-to-Video (V2V) workflows within ComfyUI. Built upon a strategic fusion of key checkpoints (including ref, fl, b25-49, etc.), this model underwent deep high-step fine-tuning. To preserve the original model's foundational strengths and broad generalization while solving artifacts introduced by high-step training, we spent 3 full days on precise model pruning and weight optimization. The result is a clean, sharp…

Open weights apache-2.0 minimax-h3

Model · Image to video

LTX-2.5

Comfyicu

LTX-2.5 is an open world model with open weights, built for local execution and fine-tuning. Its established use is generating synchronized, high-fidelity video and audio from text, image, and video inputs; applicability to emerging domains such as robotics and physical AI is developing. Full control and customization — self-host on your own infrastructure. No per-generation billing, no per-seat lock-in, no forced API dependency. Revenue is measured across the whole entity, including subsidiaries and affiliates under common control. The full, binding terms live in LICENSE. - Native multishot generation — generate connected scenes in a single pass: multiple shots that hold character…

Open weights other diffusion-single-file

Model · Image to video

MiniMax-H3-GGUF

Leejet

The license of the quantized files follows the license of the original model: These files are converted using https://github.com/leejet/stable-diffusion.cpp This model can be used with stable-diffusion.cpp. For setup instructions and usage details, please refer to: To use this model in ComfyUI, first install the following custom node: An example ComfyUI workflow is available here: src="https://huggingface.co/leejet/MiniMax-H3-GGUF/resolve/main/assets/example.mp4" controls muted

Open weights

Model · Image to video

MiniMax-H3-GGUF

Bálint Molnár-Kaló

This repository (molbal/MiniMax-H3-GGUF) provides GGUF quantized versions and necessary component files for the MiniMax H3 model. MiniMax H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. This repository includes quantized versions of both the FL2VA (First-and-last-frame mode) and Ref2VA (Omni-reference mode) base models. The FL2VA builds are pruned to FP8 first and then quantized to GGUF. minimaxh3fl2vaprunedfp8Q40.gguf (11.4 GB) minimaxh3fl2vaprunedfp8Q80.gguf (20.2 GB)…

Open weights other

Model · Image to video

Lightricks-LTX-2

Daniel

LTX-2.5 is an open world model with open weights, built for local execution and fine-tuning. Its established use is generating synchronized, high-fidelity video and audio from text, image, and video inputs; applicability to emerging domains such as robotics and physical AI is developing. Full control and customization — self-host on your own infrastructure. No per-generation billing, no per-seat lock-in, no forced API dependency. Revenue is measured across the whole entity, including subsidiaries and affiliates under common control. The full, binding terms live in LICENSE. - Native multishot generation — generate connected scenes in a single pass: multiple shots that hold character…

Open weights other diffusion-single-file

Model · Image to video

MiniMax-H3-curve-GGUF

Joey

GGUF quantizations of MiniMax-H3's 33B video+audio DiTs, built from MiniMax's pruned checkpoints. Same model, same quant tiers as the original-form repo, about 40% smaller - and a Q80 that fits a 24 GB card. fl2va = text/first-last-frame to video+audio (T2V and I2V). ref2va = reference-conditioned generation (identity from up to 9 images, 3 videos with soundtracks, 3 voice clips). Runs in ComfyUI with ComfyUI-GGUF plus a one-line architecture patch (node pack: ComfyUI-H3-Multishot; workflows: On anything older these will not load at all, because the shape of the modulation weights changed and older builds do not know how to read them. That is the only catch - everything else is a drop-in…

Open weights other minimax-h3

Model · Image to video

Lightricks-LTX-2-DISTILLED-10-Eros

Daniel

LTX2.5 temp v3 has 10eros audio while v2 is sulphur audio, hopefully the 85 can fix it if you have audio problems,LTX2.5DEV10Eros15r512.safetensors seems most stable LTX2.3DISTILLEDBAKEDLTXSULPHURSTYLEIS10Erosr256.safetensors this seems to be the best for stacking lora i'm just guessing the strength and not 100% sure which reason is best 10 Eros v1.4 Changelog: Built off 1.3 and bringing back explicit prompting and motion hopefully without any kind of anatomy redraw or negative tendency. Still requires intense prompt refinement. This version is set up to be trained on to fix it into a real base, it doesn't depict anatomy well but it also isn't confused by it which is priority for the first…

Open weights other diffusers

Model · Image to video

LTX-2.3-GGUF

QuantStack

This GGUF file is a direct conversion of Lightricks/LTX-2.3 is a quantized model, all original licensing terms and usage restrictions remain in effect.

Open weights other gguf

Who Publishes These Models

Questions

Which Image to video models are most downloaded?

By monthly downloads reported by the Hugging Face Hub: Wan_2.2_ComfyUI_Repackaged (5.8M); LTX-Video (762k); LTX-2 (324.4k).

Other tasks

See all