SAVRN
Search Contact SAVRN

Open-weight model · Feature extraction

Ovis-VL-Embedding-9B

by ATH-MaaS ATH-MaaS/Ovis-VL-Embedding-9B

Ovis-VL-Embedding-9B is an open-weight model for feature extraction from ATH-MaaS, released under Apache License 2.0. It has 262,144-token context. Its published files total 16.8 GB.

Ovis-VL-Embedding-9B is a high-capacity vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs.

Parameters—
Context262,144
Weights16.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By ATH-MaaS, published under apache-2.0, revision 724f4a25d5ed.

Ovis-VL-Embedding-9B is a high-capacity vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling high-accuracy cross-modal retrieval with a single encoder. The model is initialized from Qwen3.5-9B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added. Ovis-VL-Embedding-9B is designed for high-quality multimodal search…

Read ATH-MaaS's full model card

Ovis-VL-Embedding-9B is a high-capacity vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling high-accuracy cross-modal retrieval with a single encoder.

The model is initialized from Qwen3.5-9B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added.

Ovis-VL-Embedding-9B is designed for high-quality multimodal search, multimodal RAG, image and visual-document retrieval, video search and temporal localization, recommendation, and nearest-neighbor matching.

Architecture

Figure 2 from the technical report. Ovis-VL-Embedding-9B encodes text, images, visual documents, and sampled video frames as one interleaved sequence. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.

The Qwen3.5-9B backbone contains 32 language layers with hidden size 4096. Its hybrid stack repeats three Gated DeltaNet layers followed by one gated full-attention layer, combining efficient long-context processing with periodic global token interaction. Multimodal positional encoding preserves temporal and two-dimensional spatial coordinates for visual tokens.

Training

Training follows three stages:

  1. Multimodal contrastive pretraining. Large-scale, multi-task data establish broad alignment across text, image, document, and video inputs. The objective combines difficulty-aware focal contrastive learning with similarity-distribution distillation.
  2. Full-parameter homogeneous finetuning. High-quality downstream data refine fine-grained discrimination. Each micro-batch is drawn from one dataset so that gathered candidates form task-consistent negatives.
  3. Annealing Embedding Distillation. Teacher-correct examples are retained, unresolved student examples are emphasized, and confidence-adaptive forward-KL supervision transfers complementary expert capabilities.

Retrieval interface

Ovis-VL-Embedding-9B is a bi-encoder, not a cross-encoder:

  1. Pair the query with the task instruction and format it with the native processor and chat template.
  2. Encode queries and candidates independently.
  3. Extract the final-layer hidden state at the last non-padding token.
  4. L2-normalize the 4096-dimensional query and candidate embeddings.
  5. Rank candidates by cosine similarity, equivalently the dot product of the normalized vectors.

No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, and interleaved multimodal items are all treated as candidates in the same embedding space.

Performance

MMEB-v2

MMEB-v2 evaluates vision-language embeddings over 78 datasets spanning image, video, and visual-document tasks. Ovis-VL-Embedding-9B achieves 81.13 overall, outperforming the strongest compared baseline by 1.04 points.

Group Ovis-VL-Embedding-9B Best compared baseline Result
Image 83.96 81.86 +2.10
Video 72.90 75.95 -3.05
Visual document 83.06 82.38 +0.68
All 78 datasets 81.13 80.09 +1.04

The model ranks first on all four image sub-tasks, video classification, video moment retrieval, the visual-document aggregate, and ViDoRe-V1. Scaling from 2B to 9B improves the overall score by 3.67 points, with the largest gains on video question answering (+7.64), video moment retrieval (+7.23), and video retrieval (+5.74).

Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 78 MMEB-v2 datasets.

Embedding dimensions

The native output width is 4096, inherited directly from the Qwen3.5-9B backbone because no embedding projection head is added. Queries and candidates must use the same preprocessing, pooling rule, dimensionality, and L2 normalization.

Intended use

The model is intended for embedding extraction and retrieval over supported unimodal or interleaved multimodal content, including:

  • high-accuracy semantic and cross-modal search;
  • text-to-image, image-to-text, and image-to-image retrieval;
  • multimodal RAG indexing and retrieval;
  • visual-document and page retrieval;
  • text-to-video, video retrieval, and temporal localization;
  • recommendation and nearest-neighbor matching.

Limitations

  • This checkpoint does not natively support audio input. Use Ovis-Embedding-Omni-3B for audio and general omni-modal retrieval.
  • This checkpoint produces retrieval embeddings; it is not intended as a text or image generation model.
  • Retrieval quality depends on task-appropriate query instructions and the native preprocessing and chat template.
  • Performance varies by task. The reported model trails the strongest specialist baseline on the aggregate video score, video question answering, general video retrieval, ViDoRe-V2, VisRAG, and VisDoc-OOD.
  • The 4096-dimensional output and 9B-scale backbone require more memory, storage, and inference compute than the 2B variant.
  • Benchmark scores may not directly predict performance on a new domain. Evaluate with representative queries, candidates, and retrieval metrics before deployment.

Citation

If you find our embedding models useful, please consider citing our technical report:

@article{ovisembedding2026,
  title   = {Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings},
  author  = {{Ovis-Embedding Team}},
  journal = {arXiv preprint arXiv:2609.25165},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.25165}
}

License

This model is released under the Apache 2.0 license.

Resources

This repository contains the weights for Ovis-VL-Embedding-9B.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
4,096
Feed-forward size
12,288
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
ATH-MaaS/Ovis-VL-Embedding-9B
Publisher
ATH-MaaS
Task
Feature extraction
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
724f4a25d5ede0744eda3a11d2a1ec806a8f5cb0
First published
2026-09-21
Last updated
2026-09-24

Files and Weights

22 files, 16.8 GB in total. The weights are 4 files totalling 16.8 GB in safetensors.

Weights4 files · 16.8 GB
Configuration6 files · 76.5 KB
Tokenizer2 files · 20.0 MB
Documentation2 files · 19.3 KB
Other7 files · 2.4 MB
Repository1 file · 1.9 KB
Every file
FileTypeSizeSHA-256
model/model-00001-of-00004.safetensorsWeights4.9 GB ca2af5c966bd
model/model-00002-of-00004.safetensorsWeights5.0 GB 972358972d4d
model/model-00003-of-00004.safetensorsWeights5.0 GB 238c43ce3e38
model/model-00004-of-00004.safetensorsWeights1.9 GB b3409997921b
config.jsonConfiguration2.9 KB —
model/args.jsonConfiguration86 B —
model/config.jsonConfiguration2.9 KB —
model/model.safetensors.index.jsonConfiguration69.2 KB —
model/preprocessor_config.jsonConfiguration390 B —
model/processor_config.jsonConfiguration1.2 KB —
LICENSEDocumentation11.5 KB —
README.mdDocumentation7.8 KB —
figures/ovis_blog_table2.pdfOther482.1 KB 1cb316f64ae7
figures/ovis_blog_table2.pngOther262.2 KB 10f6a6b9e2fc
figures/ovis_embedding_data_centric.pngOther729.9 KB 0ad28dd7a765
figures/ovis_embedding_model_architecture.pngOther401.5 KB da256408d299
figures/ovis_embedding_train_inference.pngOther414.1 KB 49353acc8881
figures/ovis_logo.pngOther85.5 KB —
model/chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.9 KB —
model/tokenizer.jsonTokenizer20.0 MB 06b9509352d2
model/tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
16.8 GB
Download from ATH-MaaS

Released by ATH-MaaS through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2609.25165

Memory Requirements

PrecisionWeights in memory
As published16.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Ovis-VL-Embedding-9B

Can I use Ovis-VL-Embedding-9B commercially?

Yes. Ovis-VL-Embedding-9B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Ovis-VL-Embedding-9B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Feature extraction

all-MiniLM-L6-v2

Joshua

https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: You can then use the model to compute embeddings like this: You can convert this Tensor to a nested JavaScript array using.tolist(): Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights apache-2.0 512 tokens transformers.js

Model · Feature extraction

bge-base-en-v1.5

Joshua

https://huggingface.co/BAAI/bge-base-en-v1.5 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: You can then use the model to compute embeddings, as follows: You can also use the model for retrieval. For example: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights mit 512 tokens transformers.js

Model · Feature extraction

clap-htsat-unfused

LAION eV

The abstract of the paper states that: You can use this model for zero shot audio classification or extracting audio and/or textual features. You can also get the audio and text embeddings using ClapModel If you are using this model for your work, please consider citing the original paper

Open weights apache-2.0 514 tokens transformers

For more details please refer to our Github: FlagEmbedding. If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3. FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: - 1/30/2024: Release BGE-M3, a new member to BGE model series! M3 stands for Multi-linguality (100+ languages), Multi-granularities (input length up to 8192), Multi-Functionality (unification of dense, lexical, multi-vec/colbert retrieval). It is the first embedding model which supports all three retrieval methods, achieving new SOTA on multi-lingual (MIRACL) and cross-lingual (MKQA) benchmarks. Technical…

Open weights mit 512 tokens sentence-transformers

Model · Feature extraction

wavlm-large

Microsoft

The large model pretrained on 16kHz sampled speech audio. When using the model, make sure that your speech input is also sampled at 16kHz. Note: This model does not have a tokenizer as it was pretrained on audio alone. In order to use this model speech recognition, a tokenizer should be created and the model should be fine-tuned on labeled text data. Check out this blog for more in-detail explanation of how to fine-tune the model. - 60,000 hours of Libri-Light - 10,000 hours of GigaSpeech - 24,000 hours of VoxPopuli Authors: Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin…

Open weights transformers