SAVRN
Search Contact SAVRN

Open-weight model · Sentence similarity

multilingual-e5-base

by Liang Wang intfloat/multilingual-e5-base

Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei, arXiv 2024 This model has 12 layers and the embedding size is 768. Below is an example to encode queries and passages from the MS-MARCO passage ranking dataset.

Parameters278M
Context514
Weights5.3 GB
Licensemit
AccessOpen weights
Monthly Downloads7.6M

Runs On

What it takes to serve multilingual-e5-base (278M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on multilingual-e5-base

Seven tenths of a gigabyte is the whole footprint at 16-bit, on a card that holds 192 GB. That is multilingual-e5-base, a 278M-parameter sentence-similarity model. Nobody dedicates an MI300X at $1.85 an hour to 0.7 GB; you run it beside other work on a card you already own, or drop to 8-bit and 0.3 GB. Disk is another matter: 23 files and 5.3 GB, because the repository carries safetensors, ONNX, OpenVINO and PyTorch copies, and you need one.

MIT is about as short as licenses get: commercial use, modification and redistribution, provided the copyright and permission notices travel with the files. Context is 514 tokens, so long passages get chunked before encoding. The publisher initialized it from xlm-roberta-base and continually trained it on multilingual data, covering 100 languages with degradation possible in low-resource ones, so test yours first. Weights are stored in float32; the 16-bit figure assumes you convert.

Model Card

By Liang Wang, published under mit, revision d12875059715.

Multilingual E5 Text Embeddings: A Technical Report. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei, arXiv 2024

This model has 12 layers and the embedding size is 768.

Usage

Below is an example to encode queries and passages from the MS-MARCO passage ranking dataset.

Read the full model card (900 words)

Configuration

Architecture
XLMRobertaModel
Context length (tokens)
514
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
250,002
Stored precision
float32
Model type
xlm-roberta

Identity and Version

Repository
intfloat/multilingual-e5-base
Publisher
Liang Wang
Task
Sentence similarity
Modality
Text
Library
sentence-transformers
Parameters
278M parameters
Languages
af, am, ar, as, az, be, bg, bn
Revision
d128750597153bb5987e10b1c3493a34e5a4502a
First published
2023-05-19
Last updated
2026-04-02

Files and Weights

23 files, 5.3 GB in total. The weights are 6 files totalling 5.3 GB in bin, onnx, safetensors.

Weights6 files · 5.3 GB
Configuration8 files · 3.2 KB
Tokenizer4 files · 34.2 MB
Documentation1 file · 179.3 KB
Other3 files · 10.5 MB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.1 GB a18a44fad1d0
onnx/model.onnxWeights1.1 GB 84a4d426f7e8
onnx/model_O4.onnxWeights554.9 MB f60256a833ca
onnx/model_qint8_avx512_vnni.onnxWeights278.7 MB 252355187865
openvino/openvino_model.binWeights1.1 GB 116c03a2814d
pytorch_model.binWeights1.1 GB f061cb764188
.eval_results/ArguAna.yamlConfiguration575 B
1_Pooling/config.jsonConfiguration200 B
config.jsonConfiguration694 B
modules.jsonConfiguration387 B
onnx/config.jsonConfiguration686 B
onnx/special_tokens_map.jsonConfiguration280 B
sentence_bert_config.jsonConfiguration57 B
special_tokens_map.jsonConfiguration280 B
README.mdDocumentation179.3 KB
onnx/sentencepiece.bpe.modelOther5.1 MB cfc8146abe2a
openvino/openvino_model.xmlOther367.7 KB
sentencepiece.bpe.modelOther5.1 MB cfc8146abe2a
.gitattributesRepository1.5 KB
onnx/tokenizer.jsonTokenizer17.1 MB 62c24cdc13d4
onnx/tokenizer_config.jsonTokenizer418 B
tokenizer.jsonTokenizer17.1 MB 62c24cdc13d4
tokenizer_config.jsonTokenizer418 B

License and Download

License
mit
Access
Open weights, no gate
Download size
5.3 GB
Download from Liang Wang

Released by Liang Wang through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
MTEB AmazonCounterfactualClassification (de) Configuration deTask ClassificationMetric accuracyComparison conditions not established 71.7238 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (de) Configuration deTask ClassificationMetric apComparison conditions not established 82.2209 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (de) Configuration deTask ClassificationMetric f1Comparison conditions not established 69.9553 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (en) Configuration enTask ClassificationMetric accuracyComparison conditions not established 78.9701 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (en) Configuration enTask ClassificationMetric apComparison conditions not established 43.6935 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (en) Configuration enTask ClassificationMetric f1Comparison conditions not established 73.3808 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (en-ext) Configuration en-extTask ClassificationMetric accuracyComparison conditions not established 79.6552 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (en-ext) Configuration en-extTask ClassificationMetric apComparison conditions not established 28.5079 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (en-ext) Configuration en-extTask ClassificationMetric f1Comparison conditions not established 66.8452 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (ja) Configuration jaTask ClassificationMetric accuracyComparison conditions not established 73.3298 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (ja) Configuration jaTask ClassificationMetric apComparison conditions not established 20.7205 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonCounterfactualClassification (ja) Configuration jaTask ClassificationMetric f1Comparison conditions not established 59.78 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonPolarityClassification Configuration defaultTask ClassificationMetric accuracyComparison conditions not established 90.6377 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonPolarityClassification Configuration defaultTask ClassificationMetric apComparison conditions not established 87.2228 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonPolarityClassification Configuration defaultTask ClassificationMetric f1Comparison conditions not established 90.6038 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (de) Configuration deTask ClassificationMetric accuracyComparison conditions not established 41.828 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (de) Configuration deTask ClassificationMetric f1Comparison conditions not established 41.271 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (en) Configuration enTask ClassificationMetric accuracyComparison conditions not established 44.546 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (en) Configuration enTask ClassificationMetric f1Comparison conditions not established 44.0567 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (es) Configuration esTask ClassificationMetric accuracyComparison conditions not established 40.534 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (es) Configuration esTask ClassificationMetric f1Comparison conditions not established 39.8207 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (fr) Configuration frTask ClassificationMetric accuracyComparison conditions not established 39.684 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (fr) Configuration frTask ClassificationMetric f1Comparison conditions not established 39.1105 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (ja) Configuration jaTask ClassificationMetric accuracyComparison conditions not established 37.436 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (ja) Configuration jaTask ClassificationMetric f1Comparison conditions not established 37.0708 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (zh) Configuration zhTask ClassificationMetric accuracyComparison conditions not established 37.226 intfloat
Publisher reported
Evaluated revision not stated
MTEB AmazonReviewsClassification (zh) Configuration zhTask ClassificationMetric f1Comparison conditions not established 36.6537 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric map_at_1Comparison conditions not established 22.831 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric map_at_10Comparison conditions not established 36.42 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric map_at_100Comparison conditions not established 37.699 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric map_at_1000Comparison conditions not established 37.724 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric map_at_3Comparison conditions not established 32.207 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric map_at_5Comparison conditions not established 34.312 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric mrr_at_1Comparison conditions not established 23.257 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric mrr_at_10Comparison conditions not established 36.574 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric mrr_at_100Comparison conditions not established 37.854 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric mrr_at_1000Comparison conditions not established 37.878 intfloat
Publisher reported
Evaluated revision not stated
MTEB ArguAna Configuration defaultTask RetrievalMetric mrr_at_3Comparison conditions not established 32.385 intfloat
Publisher reported
Evaluated revision not stated
mteb/arguana Task ArguAnaMetric ArguAnaSetup Obtained using MTEB v1.12.75Comparison conditions not established 44.206 Obtained using MTEB v1.12.75
Reported by a third party
Evaluated revision not stated 2026-04-02
mteb/arguana Task ArguAna_default_testMetric ArguAna_default_testSetup Obtained using MTEB v1.12.75Comparison conditions not established 44.206 Obtained using MTEB v1.12.75
Reported by a third party
Evaluated revision not stated 2026-04-02

Memory Requirements

PrecisionWeights in memory
As published5.3 GB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare multilingual-e5-base

Questions About multilingual-e5-base

How much GPU memory does multilingual-e5-base need?

About 0.7 GB at 16-bit and 0.2 GB at 4-bit: the weights (278M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run multilingual-e5-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use multilingual-e5-base commercially?

Yes. multilingual-e5-base is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is multilingual-e5-base's context length?

514 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. Using this model becomes easy when you have sentence-transformers installed: Then you can use the model like this: Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings. Text Embeddings Inference (TEI) is a blazing fast inference solution for text embedding models. Send a request to /v1/embeddings to generate embeddings via the OpenAI Embeddings API: Or check…

Open weights apache-2.0 278M parameters 514 tokens sentence-transformers

Model · Sentence similarity

lt-wikidata-comp-multi

Dell Research Harvard

This is a LinkTransformer model. At its core this model this is a sentence transformer model sentence-transformers model- it just wraps around the class. It is designed for quick and easy record linkage (entity-matching) through the LinkTransformer package. The tasks include clustering, deduplication, linking, aggregation and more. Notwithstanding that, it can be used for any sentence similarity task within the sentence-transformers framework as well. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. Take a look at the documentation of sentence-transformers if you want to use this model for more than what we…

Open weights 278M parameters 514 tokens sentence-transformers

This is a LinkTransformer model. At its core this model this is a sentence transformer model sentence-transformers model- it just wraps around the class. It is designed for quick and easy record linkage (entity-matching) through the LinkTransformer package. The tasks include clustering, deduplication, linking, aggregation and more. Notwithstanding that, it can be used for any sentence similarity task within the sentence-transformers framework as well. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. Take a look at the documentation of sentence-transformers if you want to use this model for more than what we…

Open weights 278M parameters 514 tokens sentence-transformers

Model · Sentence similarity

lt-un-data-fine-fine-multi

Dell Research Harvard

This is a LinkTransformer model. At its core this model this is a sentence transformer model sentence-transformers model- it just wraps around the class. It is designed for quick and easy record linkage (entity-matching) through the LinkTransformer package. The tasks include clustering, deduplication, linking, aggregation and more. Notwithstanding that, it can be used for any sentence similarity task within the sentence-transformers framework as well. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. Take a look at the documentation of sentence-transformers if you want to use this model for more than what we…

Open weights 278M parameters 514 tokens sentence-transformers

This is a LinkTransformer model. At its core this model this is a sentence transformer model sentence-transformers model- it just wraps around the class. It is designed for quick and easy record linkage (entity-matching) through the LinkTransformer package. The tasks include clustering, deduplication, linking, aggregation and more. Notwithstanding that, it can be used for any sentence similarity task within the sentence-transformers framework as well. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. Take a look at the documentation of sentence-transformers if you want to use this model for more than what we…

Open weights 278M parameters 514 tokens sentence-transformers

Model · Sentence similarity

embeddinggemma-300m

Google

EmbeddingGemma is a 300M parameter, state-of-the-art for its size, open embedding model from Google, built from Gemma 3 (with T5Gemma initialization) and the same research and technology used to create Gemini models. EmbeddingGemma produces vector representations of text, making it well-suited for search and retrieval tasks, including classification, clustering, and semantic similarity search. This model was trained with data in 100+ spoken languages. The small size and on-device focus makes it possible to deploy in environments with limited resources such as mobile phones, laptops, or desktops, democratizing access to state of the art AI models and helping foster innovation for everyone.…

Access requested at publisher gemma 303M parameters sentence-transformers