SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

SmolVLM2-500M-Video-Instruct

by Hugging Face Smol Models Research HuggingFaceTB/SmolVLM2-500M-Video-Instruct

SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content.

Parameters507M
Context8,192
Weights7.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.3M

Runs On

What it takes to serve SmolVLM2-500M-Video-Instruct (507M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.5 GB 0.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.3 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on SmolVLM2-500M-Video-Instruct

Ask this model what happens in a clip and it answers from 507M parameters. At 16-bit the weights load in 1.0 GB and need 1.2 GB; 8-bit needs 0.6 GB, 4-bit 0.3 GB, and the publisher puts video inference at 1.8 GB of GPU RAM. The cheapest metered host in our data is one MI300X with 192 GB at $1.85 per hour, so the hardware question is how many copies you pack per card. Context is 8,192 tokens.

Apache 2.0 permits commercial use, modification and redistribution provided the license, copyright notices and any NOTICE file travel with it. Two checks before committing: the repository is 38 files and 7.9 GB because it ships safetensors and ONNX, well above the 1.0 GB you load, and it is quantized from SmolVLM-500M-Instruct and trained on seven listed sets including the_cauldron, Docmatix and LLaVA-Video-178K, which your data policy should clear.

Model Card

By Hugging Face Smol Models Research, published under apache-2.0, revision 7b375e1b73b1.

SmolVLM2-500M-Video

SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its compact size, requiring only 1.8GB of GPU RAM for video inference, it delivers robust performance on complex multimodal tasks. This efficiency makes it particularly well-suited for on-device applications where computational resources may be limited.

Model Summary

  • Developed by:Hugging Face
  • Model type: Multi-modal model (image/multi-image/video/text)
  • Language(s) (NLP): English
  • License: Apache 2.0
  • Architecture: Based on Idefics3 (see technical summary)

Resources

Uses

Read the full model card (764 words)

Configuration

Architecture
SmolVLMForConditionalGeneration
Context length (tokens)
8,192
Layers
32
Hidden size
960
Feed-forward size
2,560
Attention heads
15
Key/value heads
5
Head dimension
64
Vocabulary size
49,280
RoPE base
100,000
Stored precision
float32
Model type
smolvlm

Identity and Version

Repository
HuggingFaceTB/SmolVLM2-500M-Video-Instruct
Publisher
Hugging Face Smol Models Research
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
507M parameters
Languages
en
Revision
7b375e1b73b11138ff12fe22c8f2822d8fe03467
First published
2025-02-11
Last updated
2025-04-08

Files and Weights

38 files, 7.9 GB in total. The weights are 25 files totalling 7.9 GB in onnx, safetensors.

Weights25 files · 7.9 GB
Configuration7 files · 10.6 KB
Tokenizer4 files · 4.8 MB
Documentation1 file · 10.9 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights2.0 GB b9bfd456c947
onnx/decoder_model_merged.onnxWeights1.5 GB 2c8743f02060
onnx/decoder_model_merged_bnb4.onnxWeights206.5 MB c20111c92ece
onnx/decoder_model_merged_fp16.onnxWeights725.5 MB ace259f4ff3a
onnx/decoder_model_merged_int8.onnxWeights365.0 MB 7d0a356562c7
onnx/decoder_model_merged_q4.onnxWeights229.1 MB 69561703c238
onnx/decoder_model_merged_q4f16.onnxWeights205.3 MB 6265f0aa571a
onnx/decoder_model_merged_quantized.onnxWeights365.0 MB a5aaa12d62d4
onnx/decoder_model_merged_uint8.onnxWeights365.0 MB a5aaa12d62d4
onnx/embed_tokens.onnxWeights189.2 MB 4696831e7181
onnx/embed_tokens_bnb4.onnxWeights189.2 MB dc09c2b32643
onnx/embed_tokens_fp16.onnxWeights94.6 MB bc9e790f2e90
onnx/embed_tokens_int8.onnxWeights47.3 MB 0e6d275190cd
onnx/embed_tokens_q4.onnxWeights189.2 MB dc09c2b32643
onnx/embed_tokens_q4f16.onnxWeights94.6 MB a330cc444671
onnx/embed_tokens_quantized.onnxWeights47.3 MB 0e6d275190cd
onnx/embed_tokens_uint8.onnxWeights47.3 MB 0e6d275190cd
onnx/vision_encoder.onnxWeights393.2 MB d4797cdad0ba
onnx/vision_encoder_bnb4.onnxWeights60.7 MB 749b03c76adc
onnx/vision_encoder_fp16.onnxWeights196.7 MB 9faa8f3b4a92
onnx/vision_encoder_int8.onnxWeights99.0 MB d931eb204e21
onnx/vision_encoder_q4.onnxWeights66.7 MB 2fb71423919f
onnx/vision_encoder_q4f16.onnxWeights57.7 MB f7e811e4cef2
onnx/vision_encoder_quantized.onnxWeights99.0 MB 8e8e2950e683
onnx/vision_encoder_uint8.onnxWeights99.0 MB 8e8e2950e683
added_tokens.jsonConfiguration4.7 KB
chat_template.jsonConfiguration430 B
config.jsonConfiguration3.8 KB
generation_config.jsonConfiguration136 B
preprocessor_config.jsonConfiguration599 B
processor_config.jsonConfiguration67 B
special_tokens_map.jsonConfiguration868 B
README.mdDocumentation10.9 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer466.4 KB
tokenizer.jsonTokenizer3.5 MB
tokenizer_config.jsonTokenizer28.6 KB
vocab.jsonTokenizer800.7 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
7.9 GB
Download from Hugging Face Smol Models Research

Released by Hugging Face Smol Models Research through its official repository on Hugging Face. Read the license.

Built From

  • Derived from HuggingFaceTB/SmolVLM-500M-Instruct
  • Described by arXiv:2504.05299
  • Quantized from HuggingFaceTB/SmolVLM-500M-Instruct
  • Trained on (disclosed) Enxin/MovieChat-1K_train
  • Trained on (disclosed) HuggingFaceFV/finevideo
  • Trained on (disclosed) HuggingFaceM4/Docmatix
  • Trained on (disclosed) HuggingFaceM4/the_cauldron
  • Trained on (disclosed) MAmmoTH-VL/MAmmoTH-VL-Instruct-12M
  • Trained on (disclosed) Mutonix/Vript
  • Trained on (disclosed) ShareGPT4Video/ShareGPT4Video
  • Trained on (disclosed) TIGER-Lab/VISTA-400K
  • Trained on (disclosed) lmms-lab/LLaVA-OneVision-Data
  • Trained on (disclosed) lmms-lab/LLaVA-Video-178K
  • Trained on (disclosed) lmms-lab/M4-Instruct-Data
  • Trained on (disclosed) orrzohar/Video-STaR

Memory Requirements

PrecisionWeights in memory
As published7.9 GB
16-bit1.0 GB
8-bit0.5 GB
4-bit0.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare SmolVLM2-500M-Video-Instruct

Questions About SmolVLM2-500M-Video-Instruct

How much GPU memory does SmolVLM2-500M-Video-Instruct need?

About 1.2 GB at 16-bit and 0.3 GB at 4-bit: the weights (507M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run SmolVLM2-500M-Video-Instruct on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use SmolVLM2-500M-Video-Instruct commercially?

Yes. SmolVLM2-500M-Video-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is SmolVLM2-500M-Video-Instruct's context length?

8,192 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

surya-ocr-2

Datalab

Surya is a 650M param OCR model with these features: - Accuracy - scores 83.3% on olmOCR-bench (top under 3B params) - Multilingual - scores 87.2% on an internal benchmark set of 91 languages (more here) - Layout analysis (table, image, header, etc.) with reading order - Table recognition (rows + columns) It works on a range of documents (see usage and benchmarks). Our managed platform runs both Surya, and variants of our highest accuracy model, Chandra. Get started with $5 in free credits — sign up (takes under 30 seconds) or try our free public playground. Surya is named for the Hindu sun god, who has universal vision. The Surya code is licensed under Apache 2.0. The model weights use a…

Open weights openrail 686M parameters 262,144 tokens transformers

Model · Image and text to text

Florence-2-base

Microsoft

This Hub repository contains a HuggingFace's transformers implementation of Florence-2 model from Microsoft. Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks. Florence-2 can interpret simple text prompts to perform tasks like captioning, object detection, and segmentation. It leverages our FLD-5B dataset, containing 5.4 billion annotations across 126 million images, to master multi-task learning. The model's sequence-to-sequence architecture enables it to excel in both zero-shot and fine-tuned settings, proving to be a competitive vision foundation model. Use the code below to get started with the…

Open weights mit 232M parameters 1,024 tokens transformers

Model · Image and text to text

Qwen3.5-0.8B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Scores of Qwen3.5 models are reported…

Open weights apache-2.0 873M parameters 262,144 tokens transformers

Model · Image and text to text

PaddleOCR-VL-1.6

PaddlePaddle

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to those regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score…

Open weights apache-2.0 959M parameters 131,072 tokens PaddleOCR

Model · Image and text to text

PaddleOCR-VL-1.5

PaddlePaddle

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing PaddleOCR-VL-1.5 is an advanced next-generation model of PaddleOCR-VL, achieving a new state-of-the-art accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions—including scanning artifacts, skew, warping, screen photography, and illumination—we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model’s capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM…

Open weights apache-2.0 959M parameters 131,072 tokens PaddleOCR

Model · Image and text to text

Omni-Edu-27B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.8-27B on the on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 16 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 128 - totalevalbatchsize: 128 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 3M parameters 262,144 tokens transformers