CodeRankEmbed-flash-attn is an open-weight model from Jackson Davis, released under MIT License. It has 137M parameters. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 11.9k downloads a month.
A bf16 quantization of nomic-ai/CodeRankEmbed with a three-tier attention dispatch built into a custom modelinghfnomicbert.py shipped in this repo. It is not a finetune — the weights are the original CodeRankEmbed weights cast to bf16 (no further training).
Runs On
What it takes to serve CodeRankEmbed-flash-attn (137M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.3 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
CodeRankEmbed-flash-attn on every accelerator the SAVRN Index prices, at every precision
Model Card
By Jackson Davis, published under mit, revision 8187d6a65fa3.
A bf16 quantization of nomic-ai/CodeRankEmbed
with a three-tier attention dispatch built into a custom modeling_hf_nomic_bert.py shipped in
this repo. It is not a finetune — the weights are the original CodeRankEmbed weights cast to bf16
(no further training). Two of the three tiers replace the original eager O(seq²) attention with an
O(N) unpadded path; the third keeps the original eager algorithm as the correctness reference and
universal fallback.
Why
nomic-ai/CodeRankEmbed loads through trust_remote_code, and its attention path is eager
only — activation memory grows as batch × heads × seq², which OOMs at large batches even though
the model is only 137M params. This repo adds two attention paths that compute the same attention in
O(N) memory by packing unpadded sequences, so the large batches that OOM the eager path run
comfortably — with parity embeddings (no quality change):
Configuration
- Architecture
- NomicBertModel
- Vocabulary size
- 30,528
- Stored precision
- bfloat16
- Model type
- nomic_bert
Identity and Version
- Repository
- handwoven8588/CodeRankEmbed-flash-attn
- Publisher
- Jackson Davis
- Task
- Not stated by the source
- Modality
- Other
- Library
- sentence-transformers
- Parameters
- 137M parameters
- Languages
- en
- Revision
- 8187d6a65fa30b4d79004eebf7ca63977eed003a
- First published
- 2026-06-20
- Last updated
- 2026-09-26
Files and Weights
15 files, 274.5 MB in total. The weights are 1 file totalling 273.5 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 273.5 MB | 919b7d8b4352 |
| 1_Pooling/config.json | Configuration | 296 B | — |
| config.json | Configuration | 1.5 KB | — |
| config_sentence_transformers.json | Configuration | 245 B | — |
| configuration_hf_nomic_bert.py | Configuration | 2.0 KB | — |
| modeling_hf_nomic_bert.py | Configuration | 64.3 KB | — |
| modules.json | Configuration | 229 B | — |
| sentence_bert_config.json | Configuration | 140 B | — |
| special_tokens_map.json | Configuration | 695 B | — |
| NOTICE | Documentation | 1.3 KB | — |
| README.md | Documentation | 9.4 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 711.6 KB | — |
| tokenizer_config.json | Tokenizer | 1.4 KB | — |
| vocab.txt | Tokenizer | 231.5 KB | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 273.5 MB
Released by Jackson Davis through its official repository on Hugging Face. Read the license.
Built From
- Derived from nomic-ai/CodeRankEmbed
- Quantized from nomic-ai/CodeRankEmbed
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 273.5 MB |
| 16-bit | 0.3 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About CodeRankEmbed-flash-attn
How much GPU memory does CodeRankEmbed-flash-attn need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (137M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run CodeRankEmbed-flash-attn on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use CodeRankEmbed-flash-attn commercially?
Yes. CodeRankEmbed-flash-attn is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.