SAVRN
Search Contact SAVRN

Organization

ATH-MaaS

ATH-MaaS

Models in Library3
Datasets in Library0
Models on Hugging Face62
Followers1.1k

Models

Model · Feature extraction

Ovis-VL-Embedding-9B

ATH-MaaS

Ovis-VL-Embedding-9B is a high-capacity vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling high-accuracy cross-modal retrieval with a single encoder. The model is initialized from Qwen3.5-9B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added. Ovis-VL-Embedding-9B is designed for high-quality multimodal search…

Open weights apache-2.0 262,144 tokens transformers

Model · Feature extraction

Ovis-VL-Embedding-2B

ATH-MaaS

Ovis-VL-Embedding-2B is a compact vision-language embedding model for text, images, visual documents, video, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling cross-modal retrieval with a single 2B-scale encoder. The model is initialized from Qwen3.5-2B. It retains the native text and vision encoders together with the shared multimodal language backbone, removes the language-modeling head, and directly uses the final-layer hidden state at the last non-padding token as the retrieval embedding. No modality-specific projection head is added. Ovis-VL-Embedding-2B is designed for efficient multimodal search, multimodal RAG…

Open weights apache-2.0 262,144 tokens transformers

Model · Feature extraction

Ovis-Omni-Embedding-3B

ATH-MaaS

Ovis-Omni-Embedding-3B is a 3B-parameter universal embedding model for text, images, visual documents, video, audio, and interleaved multimodal inputs. It maps every supported input type into one coherent representation space, enabling any-to-any retrieval with a single encoder. The model is initialized from Qwen2.5-Omni-3B. Rather than attaching separate modality-specific embedding towers, it retains the native text tokenizer, vision encoder, audio encoder, and shared Thinker backbone. The speech-generation Talker and language-modeling head are removed, and the final-layer hidden state at the last non-padding token is used directly as the retrieval embedding. Ovis-Omni-Embedding-3B is…

Open weights apache-2.0 transformers