SAVRN
Search Contact SAVRN

Research paper · 2025-01-02

KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

Xinshuo Hu, Zifei Shan, Xinping Zhao, Zetian Sun, Zhenyu Liu, Dongfang Li, Shaolin Ye, Xinyuan Wei, Qian Chen, Baotian Hu, Min Zhang

3 open models in the SAVRN Model Hub cite KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model (2025). Together they draw 81 downloads a month. The most downloaded is KaLM-Reranker-V1-Nano-R2 by KaLM-Embedding (text ranking, 786M parameters).

Published2025-01-02
Authors11
Citing Models3
arXiv2501.01028

Abstract

As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial. Despite the growing number of general embedding models, prior work often overlooks the critical role of training data quality. In this work, we introduce KaLM-Embedding, a general multilingual embedding model that leverages a large quantity of cleaner, more diverse, and domain-specific training data. Our model has been trained with key techniques proven to enhance performance: (1) persona-based synthetic data to create diversified examples distilled from LLMs, (2) ranking consistency filtering to remove less informative samples, and (3) semi-homogeneous task batch sampling to improve training efficacy. Departing from traditional BERT-like architectures, we adopt Qwen2-0.5B as the pre-trained model, facilitating the adaptation of auto-regressive language models for general embedding tasks. Extensive evaluations of the MTEB benchmark across multiple languages show that our model outperforms others of comparable size, setting a new standard for multilingual embedding models with <1B parameters.

Full paper on arXiv · Code

Details

arXiv identifier
2501.01028
Published
2025-01-02
Authors
Xinshuo Hu, Zifei Shan, Xinping Zhao, Zetian Sun, Zhenyu Liu, Dongfang Li, Shaolin Ye, Xinyuan Wei, Qian Chen, Baotian Hu, Min Zhang

Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

ModelTaskSizeLicenseMonthly downloadsCheapest setup at 16-bit
KaLM-Reranker-V1-Nano-R2
KaLM-Embedding
Text ranking 786M apache-2.0 43 1x MI300X $1.85/hr
KaLM-Reranker-V1-Small-R2
KaLM-Embedding
Text ranking 2.1B apache-2.0 22 1x MI300X $1.85/hr
KaLM-Reranker-V1-Large-R2
KaLM-Embedding
Text ranking 7.5B apache-2.0 16 1x MI300X $1.85/hr