A bf16 quantization of nomic-ai/CodeRankEmbed with a three-tier attention dispatch built into a custom modelinghfnomicbert.py shipped in this repo. It is not a finetune — the weights are the original CodeRankEmbed weights cast to bf16 (no further training). Two of the three tiers replace the original eager O(seq²) attention with an O(N) unpadded path; the third keeps the original eager algorithm as the correctness reference and universal fallback. nomic-ai/CodeRankEmbed loads through trustremotecode, and its attention path is eager only — activation memory grows as batch × heads × seq², which OOMs at large batches even though the model is only 137M params. This repo adds two attention paths…
Independent publisher
Jackson Davis
handwoven8588
Models in Library1
Datasets in Library0
Models on Hugging Face2
Followers2