Creating these models takes significant time, work and compute. If you find them useful consider supporting me: Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs. GGUF quantizations of llmfan46/gemma-4-Ortenzya-The-Creative-Wordsmith-31B-it-uncensored-heretic. attn.oproj Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections. The aim of this finetune was to improve this…
Open weights
apache-2.0
transformers
Model · Text to video
Cαlcμ
drag gguf to >./ComfyUI/models/diffusionmodels - drag t5xxl-um to >./ComfyUI/models/textencoders - drag vae to >./ComfyUI/models/vae - for i2v model, drag clip-vision-h to >./ComfyUI/models/clipvision - run the.bat file in the main directory (assume you are using gguf pack below) - if you opt to use fp8 scaled umt5xxl encoder (if applies to any fp8 scale t5 actually), please use cpu offload (switch from default to cpu under device in gguf clip loader; won't affect speed); btw, it works fine for both gguf umt5xxl and gguf vae - drag any demo video (below) to > your browser for workflow - pig is a lazy architecture for gguf node; it applies to all model, encoder and vae gguf file(s); if you…
Open weights
apache-2.0
This is an uncensored version of google/gemma-4-12B-it-qat-q40-unquantized created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. Note: For this model, both the thinking mode and the non-thinking mode have been completely abliterated. Only layers 23-36 have been abliterated. Please use the latest version of ollama You can use huihuiai/gemma-4-abliterated:12b-qat directly, Please use the latest version of ggml-org/llama.cpp - Risk of Sensitive or Controversial Outputs: This model’s safety filtering has been significantly reduced, potentially…
Open weights
apache-2.0
transformers
Model · Audio classification
Ivan
MLX-compatible weights for WeSpeaker ResNet34-LM, converted from the pyannote speaker embedding model with BatchNorm fused into Conv2d. WeSpeaker ResNet34-LM is a speaker embedding model (~6.6M params) that produces 256-dimensional L2-normalized speaker embeddings from audio. Trained on VoxCeleb for speaker verification and diarization. BatchNorm is fused into Conv2d at conversion time — no BN layers in the MLX model. Part of speech-swift. Converts the original pyannote/wespeaker-voxceleb-resnet34-LM checkpoint using a custom unpickler (no pyannote.audio dependency required). Key transformations: - Fuse BatchNorm into Conv2d: wfused = w × γ/√(σ²+ε), bfused = β − μ×γ/√(σ²+ε) - Transpose…
Open weights
mit
7M parameters
mlx
BEiT model pre-trained in a self-supervised fashion on ImageNet-21k (14 million images, 21,841 classes) at resolution 224x224, and fine-tuned on ADE20k (an important benchmark for semantic segmentation of images) at resolution 640x640. It was introduced in the paper BEIT: BERT Pre-Training of Image Transformers by Hangbo Bao, Li Dong and Furu Wei and first released in this repository. Disclaimer: The team releasing BEiT did not write a model card for this model so this model card has been written by the Hugging Face team. The BEiT model is a Vision Transformer (ViT), which is a transformer encoder model (BERT-like). In contrast to the original ViT model, BEiT is pretrained on a large…
Open weights
apache-2.0
transformers
GGUF quantizations of Lightricks' LTX-2.5, converted for ComfyUI with ComfyUI-GGUF. LTX-2.5 generates video and synchronized audio in a single pass — a dual-stream DiT with a 4096-wide video path, a 2048-wide audio path, and cross-modal attention joining them. The bf16 transformer is 39 GB. These quants bring it to 8-22 GB. All original licensing terms and usage restrictions carry over from the base model. If you convert LTX-2.5 to GGUF yourself, it will not load. These files have the fix baked in; this section explains what it is, because it isn't obvious and it cost a night to find. ComfyUI derives most model dimensions from tensor shapes. For LTX-2.5 it can't: the connector widths, the…
Open weights
other
gguf
FlowState is the first time-scale adjustable Time Series Foundation Model (TSFM), open-sourced by IBM Research. Combining a State Space Model (SSM) Encoder with a Functional Basis Decoder allows FlowState to transition into a timescale invariant coefficient space and make a continuous forecast from this space. This allows FlowState to seamlessly adjust to all possible sampling rates. Therefore, training in one time-scale helps for inference at all scales, allowing for drastically improved utilization of training data across time-scales. This innovation leads to a significant improvement in performance, making FlowState the new state-of-the art in zero-shot time series forecasting.…
Open weights
apache-2.0
9M parameters
The model is a fine-tuned version of jonatasgrosman/wav2vec2-large-xlsr-53-english for a Speech Emotion Recognition (SER) task. The dataset used to fine-tune the original pre-trained model is the RAVDESS dataset. This dataset provides 1440 samples of recordings from actors performing on 8 different emotions in English, which are: It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 0.0001 - trainbatchsize: 4 - evalbatchsize: 4 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 8 - lrschedulertype: linear - numepochs: 3 - mixedprecisiontraining: Native AMP Any doubt, contact me on Twitter. - Transformers 4.8.2…
Open weights
apache-2.0
316M parameters
transformers
Model · Text to video
Hpmg
空间思维与物理逻辑 LoRA,基于 MiniMax-H3(Comfy-Org/MiniMax-H3)训练,让模型学会纯物体的空间关系与物理运动(碰撞、堆叠、掉落、遮挡等)。 目前还是训练和测试阶段,一些素材片段以及打标问题导致LORA还不是特别稳定。下一阶段准备修复后重新训练。目前LORA也是可以使用,强度建议0.3 ~ 0.5。纯属是在原模型基础上稍微增强一点物理反馈。 最新是重新训练到了wushuspatialphysicsclean3000pruned.safetensors 版本。 实测用了这个LORA,视频整体提升真实感物理的逻辑,比如物体碰撞的真实反馈。也可以用于一些打斗场景,人物真实碰撞的效果。 用了空间物理LORA的武打片段(强度 0.3) 没有用空间物理LORA的武打片段(同样提示词) 1. 下载.safetensors,放入 ComfyUI models/loras/ 2. LoraLoader 加载,strength 建议 0.8~1.0 3. 用空间/物理语言 prompt 描述物体运动 - several colored objects 多个彩色物体(CLEVRER 风格) - rigid objects 刚体 / elastic objects 弹性物体 - metal objects 金属物体 - a ball / balls 球 - billiard balls 台球 - objects 通用物体 - on a table 桌面上(PhyCo 台球场景) - in a synthetic scene 合成场景(CLEVRER 风格)…
Open weights
apache-2.0
minimax-h3
Model · Time series forecasting
Datadog
Toto (Time Series Optimized Transformer for Observability) is a family of time series foundation models for multivariate forecasting developed by Datadog. Toto 2.0 is the current generation, featuring u-μP-scaled transformers ranging from 4m to 2.5B parameters, all trained from a single recipe. Forecast quality improves reliably with parameter count across the family. The family sets a new state of the art on three forecasting benchmarks: BOOM, our observability benchmark; GIFT-Eval, the standard general-purpose benchmark; and the recent contamination-resistant TIME benchmark. Inference code is available on GitHub. For more examples, see the Quick Start notebook and GluonTS integration…
Open weights
apache-2.0
313M parameters
pytorch
TinyTimeMixers (TTMs) are compact pre-trained models for Multivariate Time-Series Forecasting, open-sourced by IBM Research. With less than 1 Million parameters, TTM (accepted in NeurIPS 24) introduces the notion of the first-ever “tiny” pre-trained models for Time-Series Forecasting. TTM outperforms several popular benchmarks demanding billions of parameters in zero-shot and few-shot forecasting. TTMs are lightweight forecasters, pre-trained on publicly available time series data with various augmentations. TTM provides state-of-the-art zero-shot forecasts and can easily be fine-tuned for multi-variate forecasts with just 5% of the training data to be competitive. Refer to our paper for…
Open weights
apache-2.0
805,280 parameters
granite-tsfm
https://huggingface.co/microsoft/Florence-2-base-ft with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Example: Perform image captioning with onnx-community/Florence-2-base-ft. We also released an online demo, which you can try yourself: https://huggingface.co/spaces/Xenova/florence2-webgpu Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).
Open weights
mit
1,024 tokens
transformers.js
Unified Layout Module for PaddleOCR-VL 1.5/1.6 & GLM-OCR. This is the PP-Doclayoutv3 model weights for the PaddlePaddle framework. Get safetensors weights at PP-DocLayoutV3safetensors PP-DocLayoutV3 is specifically engineered to handle non-planar document images. It can directly predict multi-point bounding boxes for layout elements—as opposed to standard two-point boxes—and determine logical reading orders for skewed and curved surfaces within a single forward pass, significantly reducing cascading errors. This model is an essential component of PaddleOCR-VL-1.5, providing crucial layout analysis for the high-precision parsing of various real-world documents in PaddleOCR-VL. This work has…
Open weights
apache-2.0
PaddleOCR
This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m for voicemail detection. It is trained on a dataset of call recordings to distinguish between voicemail greetings and live human responses. This model builds on wav2vec2-xls-r-300m, a self-supervised speech model trained on large-scale multilingual data. We fine-tuned it on the first two seconds of a call. - Automated voicemail detection in AI-powered call assistants. - Filtering voicemail responses in customer service and sales call automation. - Only trianed on the English language. - Assumes the voicemail track is isolated and contains no audio from the caller. - Designed for the first two seconds of audio when calling a…
Open weights
apache-2.0
316M parameters
transformers
π₀.₅ is a Vision-Language-Action (VLA) model with open-world generalization from Physical Intelligence, co-trained on robot demonstrations and large-scale multimodal data to execute long-horizon tasks in unseen real-world environments. Checkpoint trained and evaluated on LIBERO tasks Note: This model currently supports only the flow-matching action head for π₀.₅ training and inference. Other components from the original work (e.g., subtask prediction, action tokenization, or RL) were not released upstream and are not included here, though the LeRobot team is actively working to support them. Original paper: π0.5: A Vision-Language-Action Model with Open-World Generalization For full…
Open weights
gemma
3.6B parameters
lerobot
LTX2.5 temp v3 has 10eros audio while v2 is sulphur audio, hopefully the 85 can fix it if you have audio problems,LTX2.5DEV10Eros15r512.safetensors seems most stable LTX2.3DISTILLEDBAKEDLTXSULPHURSTYLEIS10Erosr256.safetensors this seems to be the best for stacking lora i'm just guessing the strength and not 100% sure which reason is best 10 Eros v1.4 Changelog: Built off 1.3 and bringing back explicit prompting and motion hopefully without any kind of anatomy redraw or negative tendency. Still requires intense prompt refinement. This version is set up to be trained on to fix it into a real base, it doesn't depict anatomy well but it also isn't confused by it which is priority for the first…
Open weights
other
diffusers
Janus-Pro is a novel autoregressive framework that unifies multimodal understanding and generation. It addresses the limitations of previous approaches by decoupling visual encoding into separate pathways, while still utilizing a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder’s roles in understanding and generation, but also enhances the framework’s flexibility. Janus-Pro surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus-Pro make it a strong candidate for next-generation unified multimodal models.…
Open weights
mit
2.1B parameters
16,384 tokens
transformers
Using llama.cpp release b10630 for quantization. Don't know which to choose? Grab Q4KM (17.04GB) - usually a good mix of size and performance. Download instructions available here First, make sure you have the Hugging Face CLI installed: These quants run with llama.cpp - installable in one line via llama.app: llama-server includes a built-in chat web UI, served at http://localhost:8080 by default. These quants were made with llama.cpp release b10630 - if this model's architecture is newly supported, you'll need that release or newer to run them. They also work in: model supports image and audio input. Alongside the quants, this repo includes the multimodal projector files…
Open weights
TabPFN-3 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/TabPFN. More details can be found in the Model Report. Fitting a classifier and predicting looks like this: For more examples (e.g. how to train a regressor), see the github repo: https://github.com/PriorLabs/tabPFN! TabPFN-3 ships with default classification and regression checkpoints, plus a few experimental specialized variants. We recommend starting with the defaults — the variants can be useful in ensembling or HPO setups, or tried manually in the regime they were trained for. Their name…
Open weights
other
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in four distinct sizes: E2B, E4B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from high-end phones…
Open weights
apache-2.0
5.2B parameters
131,072 tokens
transformers
Deformable DEtection TRansformer (DETR) model trained end-to-end on COCO 2017 object detection (118k annotated images). It was introduced in the paper Deformable DETR: Deformable Transformers for End-to-End Object Detection by Zhu et al. and first released in this repository. Disclaimer: The team releasing Deformable DETR did not write a model card for this model so this model card has been written by the Hugging Face team. The DETR model is an encoder-decoder transformer with a convolutional backbone. Two heads are added on top of the decoder outputs in order to perform object detection: a linear layer for the class labels and a MLP (multi-layer perceptron) for the bounding boxes. The…
Open weights
apache-2.0
40M parameters
1,024 tokens
transformers
MobileBERT is a thin version of BERTLARGE, while equipped with bottleneck structures and a carefully designed balance between self-attentions and feed-forward networks. This model was fine-tuned from the HuggingFace checkpoint google/mobilebert-uncased on SQuAD2.0. CPU: Intel(R) Core(TM) i7-6800K CPU @ 3.40GHz Memory: 32 GiB GPUs: 2 GeForce GTX 1070, each with 8GiB memory GPU driver: 418.87.01, CUDA: 10.1 It took about 3.5 hours to finish. Note that the above results didn't involve any hyperparameter search.
Open weights
mit
25M parameters
512 tokens
transformers
Donut model fine-tuned on DocVQA. It was introduced in the paper OCR-free Document Understanding Transformer by Geewok et al. and first released in this repository. Disclaimer: The team releasing Donut did not write a model card for this model so this model card has been written by the Hugging Face team. Donut consists of a vision encoder (Swin Transformer) and a text decoder (BART). Given an image, the encoder first encodes the image into a tensor of embeddings (of shape batchsize, seqlen, hiddensize), after which the decoder autoregressively generates text, conditioned on the encoding of the encoder. This model is fine-tuned on DocVQA, a document visual question answering dataset. We…
Open weights
mit
transformers
Unleashing the full potential of the previous sugoi 14B model, Sugoi 14B Ultra delivers near-double translation accuracy compared to its quantized predecessor—achieving a BLEU score of 21.38 vs 13.67. Its prompt-following skills rival those of Qwen 2.5 Base, especially when handling the bracket-heavy text commonly found in RPG Maker projects. - Key Improvements Nearly 2× BLEU score boost over previous quantized version (21.38 vs 13.67). Stronger prompt adherence, especially with RPGM-style bracketed text. - Ideal Use Cases Japanese → English translation—especially for game dialogue or RPG text. Interactive environments—works well with chat UIs like LM Studio. Must include a system prompt…
Open weights
apache-2.0