Open-weight model
KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32
by Xinping Zhao Yuki131/KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32
KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32 is an open-weight model from Xinping Zhao. It has 786M parameters. At 16-bit it needs about 1.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
We release the checkpoints from the three-stage training pipeline described in the third version of our paper.
Runs On
What it takes to serve KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32 (786M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 1.6 GB | 1.9 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.8 GB | 0.9 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.4 GB | 0.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 23, 2026.
Model Card
We release the checkpoints from the three-stage training pipeline described in the third version of our paper. Stage 1 uses supervised fine-tuning; Stage 2 produces two checkpoints through soft-label distillation; and Stage 3 combines them through model soup to produce the final R2 models.
Excerpt from the card by Xinping Zhao.
Configuration
- Architecture
- T5Gemma2ForConditionalGeneration
- Vocabulary size
- 262,144
- Model type
- t5gemma2
Identity and Version
- Repository
- Yuki131/KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32
- Publisher
- Xinping Zhao
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 786M parameters
- Languages
- Not stated by the source
- Revision
- 0ffd0d2b0fbb819044ef3388a445bef4b29c31d2
- First published
- 2026-09-03
- Last updated
- 2026-09-23
Files and Weights
7 files, 1.6 GB in total. The weights are 1 file totalling 1.6 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 1.6 GB | acabe84ebf4b |
| config.json | Configuration | 6.1 KB | — |
| generation_config.json | Configuration | 190 B | — |
| README.md | Documentation | 2.7 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 33.4 MB | f5b325224482 |
| tokenizer_config.json | Tokenizer | 772 B | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
- Download size
- 1.6 GB
Released by Xinping Zhao through its official repository on Hugging Face.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 1.6 GB |
| 16-bit | 1.6 GB |
| 8-bit | 0.8 GB |
| 4-bit | 0.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32
How much GPU memory does KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32 need?
About 1.9 GB at 16-bit and 0.5 GB at 4-bit: the weights (786M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run KaLM-Reranker-V1-Nano-R2-Stage2-r64-a32 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.