Research paper · 2023-09-14
C-Pack: Packaged Resources To Advance General Chinese Embedding
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff
Abstract
We introduce C-Pack, a package of resources that significantly advance the field of general Chinese embeddings. C-Pack includes three critical resources. 1) C-MTEB is a comprehensive benchmark for Chinese text embeddings covering 6 tasks and 35 datasets. 2) C-MTP is a massive text embedding dataset curated from labeled and unlabeled Chinese corpora for training embedding models. 3) C-TEM is a family of embedding models covering multiple sizes. Our models outperform all prior Chinese text embeddings on C-MTEB by up to +10% upon the time of the release. We also integrate and optimize the entire suite of training methods for C-TEM. Along with our resources on general Chinese embedding, we release our data and models for English text embeddings. The English models achieve state-of-the-art performance on MTEB benchmark; meanwhile, our released English data is 2 times larger than the Chinese data. All these resources are made publicly available at https://github.com/FlagOpen/FlagEmbedding.
Details
- arXiv identifier
- 2309.07597
- Published
- 2023-09-14
- Authors
- Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff
Models That Cite This Paper
- Described bybge-small-en-v1.5
- Described bybge-large-en-v1.5
- Described bybge-base-en-v1.5
- Described bybge-small-zh-v1.5
- Described bybge-reranker-base
- Described bybge-reranker-large
- Described bybge-base-en
- Described bybge-large-zh-v1.5
- Described bybge-base-zh-v1.5