# splade-cocondenser-selfdistil by NAVER: Open-Weight Model
Source: https://savrn.com/models/splade-cocondenser-selfdistil
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Model Card

SPLADE model for passage retrieval. For additional details, please visit: This is a SPLADE Sparse Encoder model. It maps sentences & paragraphs to a 30522-dimensional sparse vector space and can be used for semantic search and sparse retrieval. First install the Sentence Transformers library: Then you can load this model and run inference. If you use our checkpoint, please cite our work

Excerpt from the card by NAVER, licensed cc-by-nc-sa-4.0.

## Configuration

Architecture

BertForMaskedLM

Context length (tokens)

512

Layers

12

Hidden size

768

Feed-forward size

3,072

Attention heads

12

Vocabulary size

30,522

Stored precision

float32

Model type

bert

## Identity and Version

Repository

naver/splade-cocondenser-selfdistil

Publisher

NAVER

Task

Feature extraction

Modality

Text

Library

sentence-transformers

Parameters

Not stated by the source

Languages

en

Revision

31f3eb4b5aae648cde974db3b7b44bf2d095106b

First published

2022-05-09

Last updated

2025-06-30

## Files and Weights

12 files, 438.8 MB in total. The weights are 1 file totalling 438.1 MB in bin.

Weights1 file · 438.1 MB

Configuration6 files · 1.5 KB

Tokenizer3 files · 698.1 KB

Documentation1 file · 3.9 KB

Repository1 file · 1.2 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| pytorch_model.bin | Weights | 438.1 MB | 0a44b51cf90c |
| 1_SpladePooling/config.json | Configuration | 106 B | — |
| config.json | Configuration | 670 B | — |
| config_sentence_transformers.json | Configuration | 274 B | — |
| modules.json | Configuration | 274 B | — |
| sentence_bert_config.json | Configuration | 57 B | — |
| special_tokens_map.json | Configuration | 112 B | — |
| README.md | Documentation | 3.9 KB | — |
| .gitattributes | Repository | 1.2 KB | — |
| tokenizer.json | Tokenizer | 466.1 KB | — |
| tokenizer_config.json | Tokenizer | 466 B | — |
| vocab.txt | Tokenizer | 231.5 KB | — |

## License and Download

License

cc-by-nc-sa-4.0

Access

Open weights, no gate

Download size

438.1 MB

[Download from NAVER](https://huggingface.co/naver/splade-cocondenser-selfdistil)

Released by NAVER through its official repository on Hugging Face.

## Built From

- Described by arXiv:2205.04733
- Trained on (disclosed) ms_marco

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 438.1 MB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About splade-cocondenser-selfdistil

### Can I use splade-cocondenser-selfdistil commercially?

Not without separate permission. splade-cocondenser-selfdistil is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0. CC BY-NC-SA 4.0 permits non-commercial sharing and adapting with credit, and requires adaptations to use the same license. Commercial use needs separate permission.

### What is splade-cocondenser-selfdistil's context length?

512 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Feature extraction

### [all-MiniLM-L6-v2](https://savrn.com/models/all-minilm-l6-v2-2)

[Joshua](https://savrn.com/model-publishers/xenova)

https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: You can then use the model to compute embeddings like this: You can convert this Tensor to a nested JavaScript array using.tolist(): Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights apache-2.0 512 tokens transformers.js

[View model](https://savrn.com/models/all-minilm-l6-v2-2)

Model · Feature extraction

### [bge-base-en-v1.5](https://savrn.com/models/bge-base-en-v1-5-2)

[Joshua](https://savrn.com/model-publishers/xenova)

https://huggingface.co/BAAI/bge-base-en-v1.5 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: You can then use the model to compute embeddings, as follows: You can also use the model for retrieval. For example: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights mit 512 tokens transformers.js

[View model](https://savrn.com/models/bge-base-en-v1-5-2)

Model · Feature extraction

### [clap-htsat-unfused](https://savrn.com/models/clap-htsat-unfused)

[LAION eV](https://savrn.com/model-publishers/laion)

The abstract of the paper states that: You can use this model for zero shot audio classification or extracting audio and/or textual features. You can also get the audio and text embeddings using ClapModel If you are using this model for your work, please consider citing the original paper

Open weights apache-2.0 514 tokens transformers

[View model](https://savrn.com/models/clap-htsat-unfused)

Model · Feature extraction

### [bge-large-zh-v1.5](https://savrn.com/models/bge-large-zh-v1-5)

[Beijing Academy of Artificial Intelligence](https://savrn.com/model-publishers/baai)

For more details please refer to our Github: FlagEmbedding. If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3. FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: - 1/30/2024: Release BGE-M3, a new member to BGE model series! M3 stands for Multi-linguality (100+ languages), Multi-granularities (input length up to 8192), Multi-Functionality (unification of dense, lexical, multi-vec/colbert retrieval). It is the first embedding model which supports all three retrieval methods, achieving new SOTA on multi-lingual (MIRACL) and cross-lingual (MKQA) benchmarks. Technical…

Open weights mit 512 tokens sentence-transformers

[View model](https://savrn.com/models/bge-large-zh-v1-5)

Model · Feature extraction

### [wavlm-large](https://savrn.com/models/wavlm-large)

[Microsoft](https://savrn.com/model-publishers/microsoft)

The large model pretrained on 16kHz sampled speech audio. When using the model, make sure that your speech input is also sampled at 16kHz. Note: This model does not have a tokenizer as it was pretrained on audio alone. In order to use this model speech recognition, a tokenizer should be created and the model should be fine-tuned on labeled text data. Check out this blog for more in-detail explanation of how to fine-tune the model. - 60,000 hours of Libri-Light - 10,000 hours of GigaSpeech - 24,000 hours of VoxPopuli Authors: Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin…

Open weights transformers

[View model](https://savrn.com/models/wavlm-large)

Model · Feature extraction

### [1](https://savrn.com/models/1)

[Unsloth Backup Account](https://savrn.com/model-publishers/unslothai)

Open weights 2,048 tokens transformers

[View model](https://savrn.com/models/1)

## NAVER

[All models and datasets](https://savrn.com/model-publishers/naver)

## Versions

- [31f3eb4b5aae](https://savrn.com/models/splade-cocondenser-selfdistil/versions/31f3eb4b5aae) · current 2026-10-06

## Explore More

- [All feature extraction models](https://savrn.com/models/tasks/feature-extraction)
- [All models under cc-by-nc-sa-4.0](https://savrn.com/models/licenses/cc-by-nc-sa-4-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-01.
- [Hugging Face record](https://huggingface.co/naver/splade-cocondenser-selfdistil)
- [How the hub is built](https://savrn.com/model-hub/methodology)
