SAVRN
Search Contact SAVRN

Open-weight model · Text ranking

GemmaDecision-270M

by Rajan Shukla rajan2k/GemmaDecision-270M

GemmaDecision-270M is an open-weight model for text ranking from Rajan Shukla, released under Gemma Terms of Use. It has 268M parameters and a 32,768-token context. At 16-bit it needs about 0.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 39 downloads a month.

A fully tuned Gemma-270M candidate ranker: it scores each supplied choice together with the request, using a scalar decision head. Intended research tasks are banking-support topic selection and three-way evidence relations. It does not generate answers.

Parameters268M
Context32,768
Weights1.1 GB
Licensegemma
AccessOpen weights
Monthly Downloads39

Runs On

What it takes to serve GemmaDecision-270M (268M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.5 GB 0.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.

GemmaDecision-270M on every accelerator the SAVRN Index prices, at every precision

Model Card

A fully tuned Gemma-270M candidate ranker: it scores each supplied choice together with the request, using a scalar decision head. Intended research tasks are banking-support topic selection and three-way evidence relations. It does not generate answers. The complete modified encoder and head are included. On the private fresh-prompt test, accuracy is 96.50% for four-way banking topics and 78.00% for three-way NLI, with an equal-family mean of 87.25%. Both frozen task targets were met on this private test. See RESULTS.md for both baselines, paired uncertainty, coverage, synthetic challenge results and failures. Package 0.2.0 uses ONNX Runtime on CPU by default. The standard install requires…

Excerpt from the card by Rajan Shukla, licensed gemma.

Configuration

Architecture
Gemma3TextModel
Context length (tokens)
32,768
Layers
18
Hidden size
640
Feed-forward size
2,048
Attention heads
4
Key/value heads
1
Head dimension
256
Vocabulary size
262,144
Sliding window (tokens)
512
Model type
gemma3_text

Identity and Version

Repository
rajan2k/GemmaDecision-270M
Publisher
Rajan Shukla
Task
Text ranking
Modality
Other
Library
pytorch
Parameters
268M parameters
Languages
en
Revision
37fc181b6e8034807aff36f5246f6984e4f1e69e
First published
2026-09-27
Last updated
2026-10-01

Files and Weights

66 files, 1.2 GB in total. The weights are 4 files totalling 1.1 GB in onnx, safetensors.

Weights4 files · 1.1 GB
Configuration37 files · 8.1 MB
Tokenizer5 files · 73.8 MB
Documentation15 files · 112.4 KB
Other4 files · 36.2 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
joint_head.safetensorsWeights1.3 MB 72ec4e7d1f09
model.safetensorsWeights536.2 MB d7a3e291bfdf
onnx/model.onnxWeights538.5 MB 19b6718eb96d
reproduction/initial-v3-joint-head.safetensorsWeights1.3 MB fd2733b7cf13
RELEASE.jsonConfiguration502 B —
SHA256SUMS.jsonConfiguration7.4 KB —
SOURCE_HASHES.jsonConfiguration851 B —
added_tokens.jsonConfiguration35 B —
base-provenance.jsonConfiguration3.3 KB —
clm_heads.pyConfiguration7.5 KB —
clm_schema.pyConfiguration6.9 KB —
common.pyConfiguration3.5 KB —
config.jsonConfiguration1.5 KB —
evaluation.jsonConfiguration13.9 KB —
evidence/calibration.jsonConfiguration1.3 KB —
evidence/cloud-payload-hashes.jsonConfiguration3.7 KB —
evidence/data-audit.jsonConfiguration8.6 KB —
evidence/development-baselines.jsonConfiguration4.5 KB —
evidence/frozen-recipe.jsonConfiguration4.0 KB —
evidence/gpu-choice.jsonConfiguration2.4 KB —
evidence/predictions/calibration-gemma_likelihood-predictions.jsonConfiguration72.0 KB —
evidence/predictions/calibration-selected-predictions.jsonConfiguration71.9 KB —
evidence/predictions/calibration-v3_frozen_joint-predictions.jsonConfiguration72.1 KB —
evidence/predictions/final-gemma_likelihood-predictions.jsonConfiguration279.0 KB —
evidence/predictions/final-selected-predictions.jsonConfiguration278.6 KB —
evidence/predictions/final-v3_frozen_joint-predictions.jsonConfiguration279.6 KB —
evidence/split-identities.jsonConfiguration6.9 MB —
evidence/training.jsonConfiguration51.2 KB —
example_request.jsonConfiguration458 B —
frozen-recipe.jsonConfiguration4.0 KB —
joint_config.jsonConfiguration992 B —
joint_deployment.pyConfiguration8.9 KB —
offline-check.jsonConfiguration3.3 KB —
onnx/added_tokens.jsonConfiguration35 B —
onnx/config.jsonConfiguration1.5 KB —
onnx/joint_config.jsonConfiguration992 B —
onnx/onnx_manifest.jsonConfiguration2.3 KB —
onnx/special_tokens_map.jsonConfiguration662 B —
onnx/validation-report.jsonConfiguration17.9 KB —
special_tokens_map.jsonConfiguration662 B —
verification-examples.jsonConfiguration1.6 KB —
BASE_MODEL_CARD.mdDocumentation28.3 KB —
DATA_PROTOCOL.mdDocumentation7.0 KB —
DATA_PROVENANCE.mdDocumentation2.3 KB —
GPU_PROFILE_RESULTS.mdDocumentation3.3 KB —
LICENSEDocumentation10.1 KB —
LICENSE_CODEDocumentation11.4 KB —
NOTICEDocumentation919 B —
README.mdDocumentation10.8 KB —
RELEASE_ATTRIBUTION_AUDIT.mdDocumentation6.9 KB —
REPRODUCIBILITY.mdDocumentation4.1 KB —
RESULTS.mdDocumentation3.7 KB —
onnx/LICENSEDocumentation10.1 KB —
onnx/LICENSE_CODEDocumentation11.4 KB —
onnx/NOTICEDocumentation919 B —
onnx/README.mdDocumentation1.5 KB —
GEMMA_PROHIBITED_USE_POLICY.txtOther3.6 KB —
TRAINING_CURVE.svgOther28.9 KB —
onnx/GEMMA_PROHIBITED_USE_POLICY.txtOther3.6 KB —
requirements.txtOther91 B —
.gitattributesRepository1.6 KB —
onnx/tokenizer.jsonTokenizer33.4 MB 7d4046bf0505
onnx/tokenizer_config.jsonTokenizer1.2 MB —
tokenizer.jsonTokenizer33.4 MB 7d4046bf0505
tokenizer.modelTokenizer4.7 MB 1299c11d7cf6
tokenizer_config.jsonTokenizer1.2 MB —

License and Download

License
gemma
Access
Open weights, no gate
Download size
1.1 GB
Download from Rajan Shukla

Released by Rajan Shukla through its official repository on Hugging Face.

Built From

  • Derived from google/gemma-3-270m
  • Quantized from google/gemma-3-270m

Memory Requirements

PrecisionWeights in memory
As published1.1 GB
16-bit0.5 GB
8-bit0.3 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About GemmaDecision-270M

How much GPU memory does GemmaDecision-270M need?

About 0.6 GB at 16-bit and 0.2 GB at 4-bit: the weights (268M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run GemmaDecision-270M on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use GemmaDecision-270M commercially?

Yes, with conditions. GemmaDecision-270M is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.

What is GemmaDecision-270M's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text ranking

KielEmbed-Rerank

Tech

This is a Cross Encoder model finetuned from BAAI/bge-reranker-base using the sentence-transformers library. It computes scores for pairs of texts, which can be used for text reranking and semantic search. First install the Sentence Transformers library: Then you can load this model and run inference. Approximate statistics based on the first 100 samples: - perdevicetrainbatchsize: 4 - numtrainepochs: 1 - learningrate: 2e-05 - warmupsteps: 0.1 - gradientaccumulationsteps: 8 - fp16: True - perdevicetrainbatchsize: 4 - numtrainepochs: 1 - maxsteps: -1 - learningrate: 2e-05 - lrschedulertype: linear - lrschedulerkwargs: None - warmupsteps: 0.1 - optim: adamwtorchfused - optimargs: None…

Open weights 278M parameters 514 tokens sentence-transformers

The Jina Reranker v2 (jina-reranker-v2-base-multilingual) is a transformer-based model that has been fine-tuned for text reranking task, which is a crucial component in many information retrieval systems. It is a cross-encoder model that takes a query and a document pair as input and outputs a score indicating the relevance of the document to the query. The model is trained on a large dataset of query-document pairs and is capable of reranking documents in multiple languages with high accuracy. Compared with the state-of-the-art reranker models, including the previous released jina-reranker-v1-base-en, the Jina Reranker v2 model has demonstrated competitiveness across a series of benchmarks…

Open weights cc-by-nc-4.0 278M parameters 1,026 tokens transformers

This is a sentence-transformers model based on a pre-trained DeepPavlov/rubert-base-cased and finetuned with MS-MARCO Russian passage ranking dataset. The model can be used for Information Retrieval in the Russian language: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See SBERT.net Retrieve & Re-rank for more details. Using this model becomes easy when you have sentence-transformers installed: Then you can use the model like this: Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you need to get the logits from the…

Open weights mit 178M parameters 512 tokens sentence-transformers

We are excited to introduce the gte-modernbert series of models, which are built upon the latest modernBERT pre-trained encoder-only foundation models. The gte-modernbert series models include both text embedding models and rerank models. The gte-modernbert models demonstrates competitive performance in several text embedding and text retrieval evaluation tasks when compared to similar-scale models from the current open-source community. This includes assessments such as MTEB, LoCO, and COIR evaluation. Use with transformers Use with sentence-transformers: Before you start, install the sentence-transformers libraries: Use with transformers.js Additionally, you can also deploy…

Open weights apache-2.0 150M parameters 8,192 tokens transformers

This model was trained on the MMARCO dataset. It is a machine translated version of MS MARCO using Google Translate. It was translated to 14 languages. In our experiments, we observed that it performs also well for other languages. As a base model, we used the multilingual MiniLMv2 model. The model can be used for Information Retrieval: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See SBERT.net Retrieve & Re-rank for more details. The training code is available here: SBERT.net Training MS Marco The usage becomes easy when you have SentenceTransformers installed. Then, you can use the pre-trained…

Open weights apache-2.0 118M parameters 514 tokens sentence-transformers

This model was trained on the MS Marco Passage Ranking task. The model can be used for Information Retrieval: Given a query, encode the query will all possible passages (e.g. retrieved with ElasticSearch). Then sort the passages in a decreasing order. See SBERT.net Retrieve & Re-rank for more details. The training code is available here: SBERT.net Training MS Marco The usage is easy when you have SentenceTransformers installed. Then you can use the pre-trained models like this: In the following table, we provide various pre-trained Cross-Encoders together with their performance on the TREC Deep Learning 2019 and the MS Marco Passage Reranking dataset.

Open weights apache-2.0 109M parameters 512 tokens sentence-transformers