Research paper · 2025-01-02
KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
Xinshuo Hu, Zifei Shan, Xinping Zhao, Zetian Sun, Zhenyu Liu, Dongfang Li, Shaolin Ye, Xinyuan Wei, Qian Chen, Baotian Hu, Min Zhang
3 open models in the SAVRN Model Hub cite KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model (2025). Together they draw 81 downloads a month. The most downloaded is KaLM-Reranker-V1-Nano-R2 by KaLM-Embedding (text ranking, 786M parameters).
Abstract
As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial. Despite the growing number of general embedding models, prior work often overlooks the critical role of training data quality. In this work, we introduce KaLM-Embedding, a general multilingual embedding model that leverages a large quantity of cleaner, more diverse, and domain-specific training data. Our model has been trained with key techniques proven to enhance performance: (1) persona-based synthetic data to create diversified examples distilled from LLMs, (2) ranking consistency filtering to remove less informative samples, and (3) semi-homogeneous task batch sampling to improve training efficacy. Departing from traditional BERT-like architectures, we adopt Qwen2-0.5B as the pre-trained model, facilitating the adaptation of auto-regressive language models for general embedding tasks. Extensive evaluations of the MTEB benchmark across multiple languages show that our model outperforms others of comparable size, setting a new standard for multilingual embedding models with <1B parameters.
Details
- arXiv identifier
- 2501.01028
- Published
- 2025-01-02
- Authors
- Xinshuo Hu, Zifei Shan, Xinping Zhao, Zetian Sun, Zhenyu Liu, Dongfang Li, Shaolin Ye, Xinyuan Wei, Qian Chen, Baotian Hu, Min Zhang
Open Models Built on This Paper
Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.
| Model | Task | Size | License | Monthly downloads | Cheapest setup at 16-bit |
|---|---|---|---|---|---|
| KaLM-Reranker-V1-Nano-R2 KaLM-Embedding |
Text ranking | 786M | apache-2.0 | 43 | 1x MI300X $1.85/hr |
| KaLM-Reranker-V1-Small-R2 KaLM-Embedding |
Text ranking | 2.1B | apache-2.0 | 22 | 1x MI300X $1.85/hr |
| KaLM-Reranker-V1-Large-R2 KaLM-Embedding |
Text ranking | 7.5B | apache-2.0 | 16 | 1x MI300X $1.85/hr |