An HTR model for historical Swedish developed by the Swedish National Archives in collaboration with the Stockholm City Archives, the Finnish National Archives and Jämtlands Fornskriftsällskap. The model is trained on Swedish handwriting from the period 1600-1900. The model is trained on Swedish running-text handwriting dating from the start of the 17th century to the end of the 19th century. Like most current HTR models it operates on a text-line level, so its intended use is within an HTR pipeline that segments the text into text lines, which are transcribed by the model. The model can be used without fine-tuning on all handwriting but performs best on the type of handwriting it was…
Open weights
apache-2.0
385M parameters
htrflow
Model · Video classification
Google
ViViT model as introduced in the paper ViViT: A Video Vision Transformer by Arnab et al. and first released in this repository. Disclaimer: The team releasing ViViT did not write a model card for this model so this model card has been written by the Hugging Face team. ViViT is an extension of the Vision Transformer (ViT) to video. We refer to the paper for details. The model is mostly meant to intended to be fine-tuned on a downstream task, like video classification. See the model hub to look for fine-tuned versions on a task that interests you. For code examples, we refer to the documentation.
Open weights
mit
transformers
Model · Time series forecasting
Amazon
Update Feb 14, 2025: Chronos-Bolt & original Chronos models are now available on Amazon SageMaker JumpStart! Check out the tutorial notebook to learn how to deploy Chronos endpoints for production use in a few lines of code. Update Nov 27, 2024: We have released Chronos-Bolt models that are more accurate (5% lower error), up to 250 times faster and 20 times more memory-efficient than the original Chronos models of the same size. Check out the new models here. Chronos is a family of pretrained time series forecasting models based on language model architectures. A time series is transformed into a sequence of tokens via scaling and quantization, and a language model is trained on these…
Open weights
apache-2.0
709M parameters
chronos-forecasting
Model · Text to video
Jay
This repository contains GGUF format model files for SulphurAI's Sulphur-2-base. The following quantization tiers are provided to accommodate different hardware capabilities and VRAM constraints.
Open weights
gguf
NVIDIA Isaac GR00T N1.7 is an open foundation model for generalized humanoid robot reasoning and skills. This cross-embodiment model takes multimodal input, including language and images, to perform manipulation tasks in diverse environments. Developers and researchers can post-train GR00T N1.7 with real or synthetic data for their specific humanoid robot or task. Isaac GR00T N1.7 is the medium-sized version of our model built using pre-trained vision and language encoders, and uses a flow matching action transformer to model a chunk of actions conditioned on vision, language and proprioception. A detailed description of the Isaac GR00T N1.X architecture is provided in the GROOT N1 White…
Open weights
3.1B parameters
Update Feb 14, 2025: Chronos-Bolt & original Chronos models are now available on Amazon SageMaker JumpStart! Check out the tutorial notebook to learn how to deploy Chronos endpoints for production use in a few lines of code. Update Nov 27, 2024: We have released Chronos-Bolt models that are more accurate (5% lower error), up to 250 times faster and 20 times more memory-efficient than the original Chronos models of the same size. Check out the new models here. Chronos is a family of pretrained time series forecasting models based on language model architectures. A time series is transformed into a sequence of tokens via scaling and quantization, and a language model is trained on these…
Open weights
apache-2.0
20M parameters
transformers
Fish Audio S2 Pro is a leading text-to-speech (TTS) model with fine-grained inline control of prosody and emotion. Trained on over 10M+ hours of audio data across 80+ languages, the system combines reinforcement learning alignment with a dual-autoregressive architecture. The release includes model weights, fine-tuning code, and an SGLang-based streaming inference engine. S2 Pro builds on a decoder-only transformer combined with an RVQ-based audio codec (10 codebooks, ~21 Hz frame rate) using a Dual-Autoregressive (Dual-AR) architecture: - Slow AR (4B parameters): Operates along the time axis and predicts the primary semantic codebook. - Fast AR (400M parameters): Generates the remaining 9…
Open weights
other
4.6B parameters
Original model is here. This model created by maxfeifei8.
Open weights
other
2.6B parameters
diffusers
This GGUF file is a direct conversion of Wan-AI/Wan2.2-TI2V-5B Since this is a quantized model, all original licensing terms and usage restrictions remain in effect. Usage The model can be used with the ComfyUI custom node ComfyUI-GGUF by city96 Place model files in ComfyUI/models/unet see the GitHub readme for further installation instructions.
Open weights
apache-2.0
gguf
source languages: fr,frBE,frCA,frFR,wa,frp,oc,ca,rm,lld,fur,lij,lmo,es,esAR,esCL,esCO,esCR,esDO,esEC,esES,esGT,esHN,esMX,esNI,esPA,esPE,esPR,esSV,esUY,esVE,pt,ptbr,ptBR,ptPT,gl,lad,an,mwl,it,itIT,co,nap,scn,vec,sc,ro,la; target languages: en; OPUS readme: fr+frBE+frCA+frFR+wa+frp+oc+ca+rm+lld+fur+lij+lmo+es+esAR+esCL+esCO+esCR+esDO+esEC+esES+esGT+esHN+esMX+esNI+esPA+esPE+esPR+esSV+esUY+esVE+pt+ptbr+ptBR+ptPT+gl+lad+an+mwl+it+itIT+co+nap+scn+vec+sc+ro+la-en; dataset: opus; model: transformer; pre-processing: normalization + SentencePiece.
Open weights
apache-2.0
512 tokens
transformers
Nori is a tabular foundation model for regression via in-context learning (ICL). Given a few labeled rows as context, it predicts on new query rows in a single forward pass, with no task-specific training or fine-tuning. The model is trained entirely on synthetic data. Mean and median R² of the base model across 96 regression tasks from three public benchmark suites (single H200, up to 50K context rows per dataset): Large-N / long-context tables (common in TabArena) are the current focus of the large-table training stages. These numbers are reproducible end-to-end with one command — see Reproducing these numbers. Paste this into Claude Code, Cursor, or any AI coding assistant and it will…
Open weights
apache-2.0
synthefy-nori
Model · Audio classification
MuQ
This is the official repository for the paper "MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization". For more detailed information, we strongly recommend referring to https://github.com/tencent-ailab/MuQ and the paper). In this repo, the following models are released: - MuQ(see this link): A large music foundation model pre-trained via Self-Supervised Learning (SSL), achieving SOTA in various MIR tasks. - MuQ-MuLan(see this link): A music-text joint embedding model trained via contrastive learning, supporting both English and Chinese texts. To begin with, please use pip to install the official muq lib, and ensure that your python>=3.8: Using MuQ-MuLan to…
Open weights
cc-by-nc-4.0
Latent Consistency Model (LCM) LoRA was proposed in LCM-LoRA: A universal Stable-Diffusion Acceleration Module by Simian Luo, Yiqin Tan, Suraj Patil, Daniel Gu et al. It is a distilled consistency adapter for runwayml/stable-diffusion-v1-5 that allows to reduce the number of inference steps to only between 2 - 8 steps. LCM-LoRA is supported in Hugging Face Diffusers library from version v0.23.0 onwards. To run the model, first install the latest version of the Diffusers library as well as peft, accelerate and transformers. audio dataset from the Hugging Face Hub: Note: For detailed usage examples we recommend you to check out our official LCM-LoRA docs The adapter can be loaded with SDv1-5…
Open weights
openrail++
diffusers
In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: This repository features our T2V-14B model, which establishes a new SOTA performance benchmark among both open-source and closed-source models. It demonstrates exceptional capabilities in generating high-quality visuals with significant motion dynamics. It is also the only video model capable of producing both Chinese and English text and supports video generation at both 480P and 720P resolutions. Your browser does not support the video tag. - Wan2.1 Text-to-Video - [x] Multi-GPU Inference code of the 14B and 1.3B…
Open weights
apache-2.0
14.3B parameters
diffusers
Original model is here. This model created by Ikena.
Open weights
other
2.6B parameters
diffusers
Model · Image and text to text
NVIDIA
LocateAnything is a vision-language model for fast and high-quality visual grounding, enabling precise object localization, dense detection, and point-based localization across diverse domains in both Enterprise Intelligence and Physical AI. The model adopts a generalist design, supporting tasks such as referring expression grounding, multi-object detection, GUI element grounding, and text localization, with strong performance in complex and cluttered scenes. Its core innovation, Parallel Box Decoding (PBD), predicts complete bounding box coordinates in a single parallel step rather than autoregressive token-by-token decoding, improving efficiency while preserving geometric consistency.…
Open weights
other
3.8B parameters
32,768 tokens
transformers
You can try our models here! We're excited to introduce the FastWan2.2 series—a new line of models finetuned with our novel Sparse-distill strategy. This approach jointly integrates DMD and VSA in a single training process, combining the benefits of both distillation to shorten diffusion steps and sparse attention to reduce attention computations, enabling even faster video generation. FastWan2.2-TI2V-5B-Full-Diffusers is built upon Wan-AI/Wan2.2-TI2V-5B-Diffusers. It supports efficient 3-step inference and produces high-quality videos at 121×704×1280 resolution. For training, we used simulated forward for the generator model, making the process data-free. The current…
Open weights
apache-2.0
5B parameters
diffusers
Model · Speech recognition
Handy
GGUF conversions of nvidia/canary-1b-v2 for use with transcribe.cpp. Ported from upstream commit pinned 2026-05-08. Validated against the NeMo reference at transcribe.cpp commit Offline multilingual speech-to-text and translation across 25 European languages. A 978M-parameter multitask AED with a 32-layer FastConformer encoder and an 8-layer Transformer decoder. Supports automatic speech recognition for any of the 25 supported languages, plus translation between supported language pairs (per the upstream model card). Takes a 16 kHz mono WAV and produces a transcript. Not a streaming model; word and segment timestamps from the upstream model are not exposed in the v1 port. WER on the full…
Open weights
cc-by-4.0
transcribe.cpp
This is the model card of NLLB-200's 1.3B variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. - Paper or other resource for more information NLLB Team et al, No Language Left Behind: Scaling Human-Centered Machine Translation, Arxiv, 2022 - Where to send questions or comments about the model: https://github.com/facebookresearch/fairseq/issues • Model performance measures: NLLB-200 model was…
Open weights
cc-by-nc-4.0
1,024 tokens
transformers
This is a quantization of Tongyi-MAI/Z-Image-Turbo to FP8 E5M2 and FP8 E4M3FN. This model strictly follows the original licensing terms and usage restrictions. Please refer to the original model card for details.
Open weights
apache-2.0
diffusers
OneFormer model trained on the COCO dataset (large-sized version, Swin backbone). It was introduced in the paper OneFormer: One Transformer to Rule Universal Image Segmentation by Jain et al. and first released in this repository. OneFormer is the first multi-task universal image segmentation framework. It needs to be trained only once with a single universal architecture, a single model, and on a single dataset, to outperform existing specialized models across semantic, instance, and panoptic segmentation tasks. OneFormer uses a task token to condition the model on the task in focus, making the architecture task-guided for training, and task-dynamic for inference, all with a single model.…
Open weights
mit
transformers
We apply Parallel Decoding Distillation (PDD) 1 to MiniMax-H3, enabling efficient video generation in only a few inference steps. For more details, please refer to our GitHub repo. Set modelpath and pddlorapath to the MiniMax-H3 model and the matching acceleration LoRA checkpoint in predictt2v.py for FL2VA or predictref2v.py for Ref2VA, then run the corresponding script. Each example uses applypddlora to load the checkpoint and derive the required number of inference steps from its configuration.
Open weights
other
videox_fun
Original model is here. This model created by janxd.
Open weights
other
2.6B parameters
diffusers
SDXL-Lightning is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps. For more information, please refer to our research paper: SDXL-Lightning: Progressive Adversarial Diffusion Distillation. We open-source the model as part of the research. Our models are distilled from stabilityai/stable-diffusion-xl-base-1.0. This repository contains checkpoints for 1-step, 2-step, 4-step, and 8-step distilled models. The generation quality of our 2-step, 4-step, and 8-step model is amazing. Our 1-step model is more experimental. We provide both full UNet and LoRA checkpoints. The full UNet models have the best quality while the LoRA models can be…
Open weights
openrail++
diffusers