SAVRN
Search Contact SAVRN

Open-weight model · Feature extraction

Ovis-Omni-Embedding-3B

by ATH-MaaS ATH-MaaS/Ovis-Omni-Embedding-3B

Ovis-Omni-Embedding-3B is an open-weight model for feature extraction from ATH-MaaS, released under Apache License 2.0. Its published files total 11.1 GB.

Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model for text, images, visual documents, video, audio, and interleaved multimodal inputs.

Parameters—
Context—
Weights11.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By ATH-MaaS, published under apache-2.0, revision 08547b8479ed.

Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model for text, images, visual documents, video, audio, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling any-to-any retrieval with a single encoder. The model is initialized from Qwen2.5-Omni-3B. Rather than attaching separate modality-specific embedding towers, it retains the native text tokenizer, vision encoder, audio encoder, and shared Thinker backbone. The speech-generation Talker and language-modeling head are removed, and the final-layer hidden state at the last non-padding token is used directly as the retrieval embedding. Ovis-Omni-Embedding-3B is…

Read ATH-MaaS's full model card

Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model for text, images, visual documents, video, audio, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling any-to-any retrieval with a single encoder.

The model is initialized from Qwen2.5-Omni-3B. Rather than attaching separate modality-specific embedding towers, it retains the native text tokenizer, vision encoder, audio encoder, and shared Thinker backbone. The speech-generation Talker and language-modeling head are removed, and the final-layer hidden state at the last non-padding token is used directly as the retrieval embedding.

Ovis-Omni-Embedding-3B is designed for multimodal search, retrieval-augmented generation, recommendation, visual-document retrieval, video and audio search, and agentic retrieval over tools, interfaces, and memory.

Architecture

Figure 2 from the technical report. Ovis-Omni-Embedding-3B encodes text, visual, and audio inputs as an interleaved token sequence processed by the Qwen2.5-Omni Thinker with TMRoPE. The final-layer hidden state at the last non-padding token becomes the retrieval embedding, with no modality-specific projection heads.

An input is formatted with a retrieval instruction through the native processor and chat template. Text, visual, and acoustic tokens are then processed as one interleaved sequence by the shared causal Transformer. Qwen2.5-Omni's time-aligned multimodal rotary position embedding preserves temporal alignment between audio and video.

Training

Figure 4 from the technical report. The training corpus spans text, images, video, audio, visual documents, and interleaved inputs. Homogeneous-source sampling forms each micro-batch from one dataset and deduplicates pooled candidates to create informative in-batch negatives without positive collisions.

Training follows three stages:

  1. Omni-modal contrastive pretraining. Globally mixed candidates and cross-device in-batch negatives establish broad alignment. The objective combines difficulty-aware focal contrastive learning with similarity-distribution distillation from complementary modality experts.
  2. Full-parameter homogeneous finetuning. High-quality downstream data refine fine-grained discrimination. Samples within each micro-batch come from one dataset, and candidate deduplication prevents positive collisions and false in-batch negatives.
  3. Annealing Embedding Distillation. Teacher-correct examples are retained, unresolved student examples are upsampled, and confidence-adaptive forward-KL supervision transfers complementary expert capabilities without adding inference-time towers.

Figure 5 from the technical report. Embedding Distillation transfers similarity distributions from complementary experts, while the inference-time low-rank module combines a shared PCA basis with lightweight residual adapters for compact embeddings.

Retrieval interface

Ovis-Omni-Embedding-3B is a bi-encoder, not a cross-encoder:

  1. Pair the query with the task instruction and format it with the model's native processor and chat template.
  2. Encode queries and candidates independently.
  3. Extract the final-layer hidden state at the last non-padding token.
  4. L2-normalize the query and candidate embeddings.
  5. Rank candidates by cosine similarity, equivalently the dot product of the normalized vectors.

No answer generation or query-candidate cross-attention is used during retrieval. Classification labels, passages, images, documents, videos, audio clips, and multimodal items are all treated as candidates in the same embedding space.

Performance

MMEB-v3

MMEB-v3 is an omni-modal benchmark comprising 190 datasets across image, video, visual-document, text, audio, and agent retrieval. Ovis-Omni-Embedding-3B achieves 58.46 overall, outperforming the strongest compared baseline by 5.19 points, and ranks first on the aggregate score of every modality group.

Group Ovis-Omni-Embedding-3B Best compared baseline Margin
Image 77.55 73.83 +3.72
Video 64.99 59.37 +5.62
Visual document 78.26 75.37 +2.89
Text 47.15 43.62 +3.53
Audio 50.08 43.17 +6.91
Agent 45.52 39.42 +6.10
All 190 datasets 58.46 53.27 +5.19

Across the 31 aggregate and sub-task entries in the complete comparison, Ovis-Omni-Embedding-3B ranks first on 22 and second on 8. MultiConIR is the only entry on which it falls outside the top two.

Scores are percentages and higher is better. Red marks the best result in each row, underlining marks the second best, and Ovis scores are bold. The overall score is the unweighted average over all 190 MMEB-v3 datasets. MMEB-v3 primarily uses Hit@1 for image, video, audio, and agent tasks and nDCG@5 for text and visual-document retrieval.

Additional benchmark results

Benchmark Ovis-Embedding-Omni-3B Best compared baseline Evaluation scope
MAEB (beta) 57.29 LCO-Embedding-Omni-7B: 53.54 Mean over 30 audio embedding tasks
MVEB (beta) 61.77 LCO-Embedding-Omni-7B: 57.58 Mean over 23 video and audio-video embedding tasks
RTEB 67.35 Qwen3-Embedding-4B: 67.27 15-task English public retrieval split

These benchmark families use their own official aggregation procedures, so their scores should not be averaged together. MAEB and MVEB results are local evaluations inserted into the corresponding leaderboard snapshots, as described in the technical report.

Embedding dimensions

The native output width is 2048 because no embedding projection head is added to the backbone. For deployments with tighter storage or latency budgets, the post-hoc elastic-dimension module supports 1024, 512, 256, and 128 dimensions. It combines a shared, modality-balanced PCA rotation with a zero-initialized residual linear adapter and folds both operations into one projection matrix at inference time.

Always use the same dimensionality and transformation for queries and candidates, and L2-normalize after projection.

Intended use

The model is intended for embedding extraction and retrieval over supported unimodal or interleaved multimodal content, including:

  • semantic and cross-modal search;
  • multimodal RAG indexing and retrieval;
  • image and visual-document retrieval;
  • video and audio retrieval;
  • recommendation and nearest-neighbor matching;
  • retrieval of tools, GUI states, and memory for agents.

Limitations

  • This checkpoint produces retrieval embeddings; it is not intended as a speech or text generation model.
  • Retrieval quality depends on using task-appropriate query instructions and the model's native preprocessing and chat template.
  • Performance varies by task. In the reported MMEB-v3 comparison, MultiConIR is the principal weakness, while several established visual-document suites and memory retrieval remain below the best specialist result.
  • Benchmark scores may not directly predict performance on a new domain. Evaluate with representative queries, candidates, and retrieval metrics before deployment.
  • For high-stakes applications, embeddings should be combined with domain-specific evaluation, access controls, and, where appropriate, a second-stage reranker.

Citation

If you find our embedding models useful, please consider citing our technical report:

@article{ovisembedding2026,
  title   = {Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings},
  author  = {{Ovis-Embedding Team}},
  journal = {arXiv preprint arXiv:2609.25165},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.25165}
}

License

This model is released under the Apache 2.0 license.

Resources

This repository contains the weights for Ovis-Omni-Embedding-3B.

Configuration

Architecture
Qwen2_5OmniForConditionalGeneration
Hidden size
2,048
Model type
qwen2_5_omni

Identity and Version

Repository
ATH-MaaS/Ovis-Omni-Embedding-3B
Publisher
ATH-MaaS
Task
Feature extraction
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
08547b8479edc10edc9878597438ebd076f74dff
First published
2026-09-20
Last updated
2026-09-24

Files and Weights

23 files, 11.1 GB in total. The weights are 3 files totalling 11.1 GB in safetensors.

Weights3 files · 11.1 GB
Configuration8 files · 304.9 KB
Tokenizer2 files · 11.4 MB
Documentation2 files · 21.4 KB
Other7 files · 2.6 MB
Repository1 file · 1.9 KB
Every file
FileTypeSizeSHA-256
model/model-00001-of-00003.safetensorsWeights5.0 GB 3859069435a0
model/model-00002-of-00003.safetensorsWeights5.0 GB 2ab4c3c8ab84
model/model-00003-of-00003.safetensorsWeights1.1 GB 6eebd641ef74
config.jsonConfiguration13.1 KB —
model/args.jsonConfiguration96 B —
model/config.jsonConfiguration13.1 KB —
model/generation_config.jsonConfiguration154 B —
model/model.safetensors.index.jsonConfiguration241.8 KB —
model/preprocessor_config.jsonConfiguration667 B —
model/processor_config.jsonConfiguration2.8 KB —
model/zero_to_fp32.pyConfiguration33.3 KB —
README.mdDocumentation9.9 KB —
model/LICENSEDocumentation11.5 KB —
figures/ovis_blog_table1.pdfOther484.4 KB ae5707e6bc8e
figures/ovis_blog_table1.pngOther436.4 KB e6cdde834fdb
figures/ovis_embedding_data_centric.pngOther729.9 KB 0ad28dd7a765
figures/ovis_embedding_model_architecture.pngOther401.5 KB da256408d299
figures/ovis_embedding_train_inference.pngOther414.1 KB 49353acc8881
figures/ovis_logo.pngOther85.5 KB —
model/chat_template.jinjaOther1.3 KB —
.gitattributesRepository1.9 KB —
model/tokenizer.jsonTokenizer11.4 MB 1ab7a851e5c6
model/tokenizer_config.jsonTokenizer939 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
11.1 GB
Download from ATH-MaaS

Released by ATH-MaaS through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2609.25165

Memory Requirements

PrecisionWeights in memory
As published11.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Ovis-Omni-Embedding-3B

Can I use Ovis-Omni-Embedding-3B commercially?

Yes. Ovis-Omni-Embedding-3B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Feature extraction

all-MiniLM-L6-v2

Joshua

https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: You can then use the model to compute embeddings like this: You can convert this Tensor to a nested JavaScript array using.tolist(): Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights apache-2.0 512 tokens transformers.js

Model · Feature extraction

bge-base-en-v1.5

Joshua

https://huggingface.co/BAAI/bge-base-en-v1.5 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: You can then use the model to compute embeddings, as follows: You can also use the model for retrieval. For example: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights mit 512 tokens transformers.js

Model · Feature extraction

clap-htsat-unfused

LAION eV

The abstract of the paper states that: You can use this model for zero shot audio classification or extracting audio and/or textual features. You can also get the audio and text embeddings using ClapModel If you are using this model for your work, please consider citing the original paper

Open weights apache-2.0 514 tokens transformers

For more details please refer to our Github: FlagEmbedding. If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3. FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: - 1/30/2024: Release BGE-M3, a new member to BGE model series! M3 stands for Multi-linguality (100+ languages), Multi-granularities (input length up to 8192), Multi-Functionality (unification of dense, lexical, multi-vec/colbert retrieval). It is the first embedding model which supports all three retrieval methods, achieving new SOTA on multi-lingual (MIRACL) and cross-lingual (MKQA) benchmarks. Technical…

Open weights mit 512 tokens sentence-transformers

Model · Feature extraction

wavlm-large

Microsoft

The large model pretrained on 16kHz sampled speech audio. When using the model, make sure that your speech input is also sampled at 16kHz. Note: This model does not have a tokenizer as it was pretrained on audio alone. In order to use this model speech recognition, a tokenizer should be created and the model should be fine-tuned on labeled text data. Check out this blog for more in-detail explanation of how to fine-tune the model. - 60,000 hours of Libri-Light - 10,000 hours of GigaSpeech - 24,000 hours of VoxPopuli Authors: Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin…

Open weights transformers