SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

surya-ocr-2

by Datalab datalab-to/surya-ocr-2

Surya is a 650M param OCR model with these features: - Accuracy - scores 83.3% on olmOCR-bench (top under 3B params) - Multilingual - scores 87.2% on an internal benchmark set of 91 languages (more here) - Layout analysis (table, image, header, etc.) with…

Parameters686M
Context262,144
Weights1.4 GB
Licenseopenrail
AccessOpen weights
Monthly Downloads1.3M

Runs On

What it takes to serve surya-ocr-2 (686M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.4 GB 1.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.7 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on surya-ocr-2

A context window of 262,144 tokens on a 686M parameter OCR model tells you what Datalab built it for: whole documents, with layout, reading order and table rows and columns alongside the text. The 16-bit weights are 1.4 GB and need 1.6 GB; 8-bit needs 0.8 GB, 4-bit 0.4 GB. The cheapest setup in our data is one MI300X with 192 GB at $1.85 per hour, so one card runs many instances side by side and the buying question is pages per hour per dollar.

Open RAIL permits commercial use subject to use-based restrictions you must pass on to anyone who receives the model or a derivative, so read that list before it goes into a product. Then check the model card's own olmOCR-bench numbers, 83.3 overall but 41.8 on old scans, against your scanned backfile. The publisher calls it 650M parameters while the configuration counts 686M.

Model Card

By Datalab, published under openrail, revision 3b3d4cdf88d6.

Datalab

State of the Art models for Document Intelligence

Surya

Surya is a 650M param OCR model with these features:

  • Accuracy - scores 83.3% on olmOCR-bench (top under 3B params)
  • Speed - throughput of 5 pages/s on an RTX 5090
  • Multilingual - scores 87.2% on an internal benchmark set of 91 languages (more here)
  • Layout analysis (table, image, header, etc.) with reading order
  • Table recognition (rows + columns)

It works on a range of documents (see usage and benchmarks).

Try Datalab's Managed Platform

Our managed platform runs both Surya, and variants of our highest accuracy model, Chandra.

Get started with $5 in free creditssign up (takes under 30 seconds) or try our free public playground.

Model Information

Detection OCR
Layout Table Recognition

Surya is named for the Hindu sun god, who has universal vision.

Examples

Read the full model card (2,491 words)

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
24
Hidden size
1,024
Feed-forward size
3,584
Attention heads
8
Key/value heads
2
Head dimension
256
Vocabulary size
65,425
Model type
qwen3_5

Identity and Version

Repository
datalab-to/surya-ocr-2
Publisher
Datalab
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
686M parameters
Languages
ocr, pdf
Revision
3b3d4cdf88d6928b0acdc75181b13206ea67c4a3
First published
2026-05-14
Last updated
2026-05-27

Files and Weights

42 files, 1.4 GB in total. The weights are 1 file totalling 1.4 GB in safetensors.

Weights1 file · 1.4 GB
Configuration6 files · 6.8 KB
Tokenizer2 files · 1.7 MB
Documentation2 files · 37.3 KB
Other30 files · 25.5 MB
Repository1 file · 3.0 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.4 GB 5755f82a997d
.eval_results/olmocrbench.yamlConfiguration1.7 KB
config.jsonConfiguration2.6 KB
generation_config.jsonConfiguration131 B
preprocessor_config.jsonConfiguration482 B
processor_config.jsonConfiguration1.3 KB
video_preprocessor_config.jsonConfiguration615 B
LICENSEDocumentation14.7 KB
README.mdDocumentation22.5 KB
assets/corporate.pngOther168.2 KB 03e5004c5ee8
assets/corporate_layout.pngOther167.4 KB c472d16817b8
assets/corporate_reading.pngOther168.2 KB fa9471e99827
assets/corporate_tablerec.pngOther163.1 KB c1861b87e15f
assets/corporate_text.pngOther166.2 KB 893c5924fe66
assets/excerpt.pngOther338.8 KB 9d1913fc79fb
assets/excerpt_layout.pngOther345.4 KB 924b81dba774
assets/excerpt_text.pngOther542.6 KB baddec3a08b6
assets/form.pngOther513.3 KB 60fb6bfde4d7
assets/form_layout.pngOther508.0 KB 631c09dc21bd
assets/form_reading.pngOther518.7 KB e88bac7fb15d
assets/form_tablerec.pngOther511.5 KB d799c78d495c
assets/form_text.pngOther329.5 KB 7c40b7233926
assets/handwritten.pngOther178.5 KB 3c7623a26db0
assets/handwritten_layout.pngOther184.7 KB ad0e4ae387b8
assets/handwritten_reading.pngOther185.4 KB 5cc6662224ca
assets/handwritten_tablerec.pngOther171.2 KB 5e3f7820dc76
assets/handwritten_text.pngOther298.5 KB 9686cfab491a
assets/newspaper.pngOther5.6 MB 1a07a43797b7
assets/newspaper_layout.pngOther5.6 MB 139c8fd41152
assets/newspaper_reading.pngOther5.6 MB c18c2eb0c39d
assets/newspaper_text.pngOther1.9 MB 364de91c602c
assets/olmocr_size_chart.pngOther82.5 KB 34addebba231
assets/scanned_tablerec.pngOther345.5 KB 86091b59d376
assets/textbook.pngOther193.1 KB 0070c7f61aae
assets/textbook_layout.pngOther194.9 KB 2711fcd2183f
assets/textbook_reading.pngOther200.8 KB 81d14c7a6d77
assets/textbook_text.pngOther225.7 KB 7950d981fdf9
chat_template.jinjaOther2.9 KB
datalab-logo.pngOther6.2 KB
.gitattributesRepository3.0 KB
tokenizer.jsonTokenizer1.7 MB
tokenizer_config.jsonTokenizer571 B

License and Download

License
openrail
Access
Open weights, no gate
Download size
1.4 GB
Download from Datalab

Released by Datalab through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
allenai/olmOCR-bench Task arxiv_mathMetric arxiv_mathComparison conditions not established 88.3 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task baselineMetric baselineComparison conditions not established 99.7 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task headers_footersMetric headers_footersComparison conditions not established 92.5 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task long_tiny_textMetric long_tiny_textComparison conditions not established 93.7 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task multi_columnMetric multi_columnComparison conditions not established 82.4 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task old_scansMetric old_scansComparison conditions not established 41.8 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task old_scans_mathMetric old_scans_mathComparison conditions not established 81.4 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task overallMetric overallComparison conditions not established 83.3 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
allenai/olmOCR-bench Task table_testsMetric table_testsComparison conditions not established 86.6 Surya OCR 2 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-27
llamaindex/ParseBench Task chartMetric chartSetup Pipeline name: surya2_sdkComparison conditions not established 21.95 ParseBench
Reported by a third party
Evaluated revision not stated 2026-06-01
llamaindex/ParseBench Task layoutMetric layoutSetup Pipeline name: surya2_sdkComparison conditions not established 71.59 ParseBench
Reported by a third party
Evaluated revision not stated 2026-06-01
llamaindex/ParseBench Task meanMetric meanSetup Pipeline name: surya2_sdkComparison conditions not established 64.83 ParseBench
Reported by a third party
Evaluated revision not stated 2026-06-01
llamaindex/ParseBench Task tableMetric tableSetup Pipeline name: surya2_sdkComparison conditions not established 82.68 ParseBench
Reported by a third party
Evaluated revision not stated 2026-06-01
llamaindex/ParseBench Task text_contentMetric text_contentSetup Pipeline name: surya2_sdkComparison conditions not established 86.57 ParseBench
Reported by a third party
Evaluated revision not stated 2026-06-01
llamaindex/ParseBench Task text_formattingMetric text_formattingSetup Pipeline name: surya2_sdkComparison conditions not established 61.36 ParseBench
Reported by a third party
Evaluated revision not stated 2026-06-01

Memory Requirements

PrecisionWeights in memory
As published1.4 GB
16-bit1.4 GB
8-bit0.7 GB
4-bit0.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare surya-ocr-2

Questions About surya-ocr-2

How much GPU memory does surya-ocr-2 need?

About 1.6 GB at 16-bit and 0.4 GB at 4-bit: the weights (686M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run surya-ocr-2 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use surya-ocr-2 commercially?

Yes, with conditions. surya-ocr-2 is released under Open RAIL License. Open RAIL licenses permit use, including commercial use, subject to the use-based restrictions listed in the license, which must be passed on to anyone who receives the model or a derivative.

What is surya-ocr-2's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its compact size, requiring only 1.8GB of GPU RAM for video inference, it delivers robust performance on complex multimodal tasks. This efficiency makes it particularly well-suited for on-device applications where computational resources may be limited. SmolVLM2 can be used for inference on multimodal (video / image / text) tasks where the input consists of text queries along with video or one or more images.…

Open weights apache-2.0 507M parameters 8,192 tokens transformers

Model · Image and text to text

Qwen3.5-0.8B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Scores of Qwen3.5 models are reported…

Open weights apache-2.0 873M parameters 262,144 tokens transformers

Model · Image and text to text

PaddleOCR-VL-1.6

PaddlePaddle

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to those regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score…

Open weights apache-2.0 959M parameters 131,072 tokens PaddleOCR

Model · Image and text to text

PaddleOCR-VL-1.5

PaddlePaddle

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing PaddleOCR-VL-1.5 is an advanced next-generation model of PaddleOCR-VL, achieving a new state-of-the-art accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions—including scanning artifacts, skew, warping, screen photography, and illumination—we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model’s capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM…

Open weights apache-2.0 959M parameters 131,072 tokens PaddleOCR

Model · Image and text to text

Florence-2-base

Microsoft

This Hub repository contains a HuggingFace's transformers implementation of Florence-2 model from Microsoft. Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks. Florence-2 can interpret simple text prompts to perform tasks like captioning, object detection, and segmentation. It leverages our FLD-5B dataset, containing 5.4 billion annotations across 126 million images, to master multi-task learning. The model's sequence-to-sequence architecture enables it to excel in both zero-shot and fine-tuned settings, proving to be a competitive vision foundation model. Use the code below to get started with the…

Open weights mit 232M parameters 1,024 tokens transformers

Model · Image and text to text

Omni-Edu-27B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.8-27B on the on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 16 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 128 - totalevalbatchsize: 128 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 3M parameters 262,144 tokens transformers