SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Gemma-4-31B-IT-NVFP4

by NVIDIA nvidia/Gemma-4-31B-IT-NVFP4

Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output.

Parameters20.9B
Context262,144
Weights32.6 GB
Licenseother
AccessOpen weights
Monthly Downloads1.6M

Runs On

What it takes to serve Gemma-4-31B-IT-NVFP4 (20.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 41.7 GB 50.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 20.9 GB 25.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 10.4 GB 12.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Gemma-4-31B-IT-NVFP4

Two numbers on this page pull against each other. The context window is 262,144 tokens, while the 4-bit row, the one that fits an NVFP4 quantization of google/gemma-4-31B-it, needs 12.5 GB of memory for 10.4 GB of weights. On the 192 GB MI300X our Index lists as the cheapest host at $1.85 per hour, the weights are a small tenant; the key-value cache for long inputs is what fills the card, so we size this one by cache, not by weights. The publisher lists text and image input and over 140 languages.

The license is recorded as other, with no commercial-use line on file, so legal reads the terms attached to the gemma-4-31B-it lineage before anything ships. No reported evaluations and no Index host prices exist for this quantization, so a buyer runs their own tests and weighs the 16-bit row's 50.1 GB against the 12.5 GB saved.

Model Card

Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding, and multimodal understanding on consumer GPUs and workstations, with a 256K-token context window and support for over 140 languages. The model uses a hybrid attention mechanism that interleaves local sliding-window and full global attention, with unified Keys and Values in global layers and Proportional RoPE (p-RoPE) to support long-context performance. The NVIDIA Gemma 4 31B IT NVFP4 model is quantized with NVIDIA Model…

Excerpt from the card by NVIDIA, licensed other.

Configuration

Architecture
Gemma4ForConditionalGeneration
Context length (tokens)
262,144
Layers
60
Hidden size
5,376
Feed-forward size
21,504
Attention heads
32
Key/value heads
16
Head dimension
256
Vocabulary size
262,144
Sliding window (tokens)
1,024
Model type
gemma4
Quantization
modelopt

Identity and Version

Repository
nvidia/Gemma-4-31B-IT-NVFP4
Publisher
NVIDIA
Task
Text generation
Modality
Text
Library
Model Optimizer
Parameters
20.9B parameters
Languages
Not stated by the source
Revision
4135a98a9b728a548947683219633b25682223ac
First published
2026-04-02
Last updated
2026-07-13

Files and Weights

15 files, 32.7 GB in total. The weights are 4 files totalling 32.6 GB in safetensors.

Weights4 files · 32.6 GB
Configuration5 files · 189.6 KB
Tokenizer2 files · 32.2 MB
Documentation1 file · 7.6 KB
Other1 file · 16.9 KB
Repository2 files · 280.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00004.safetensorsWeights10.0 GB 4d955de1f740
model-00002-of-00004.safetensorsWeights10.0 GB f8418783fa27
model-00003-of-00004.safetensorsWeights10.0 GB 5cfe8ee0f73c
model-00004-of-00004.safetensorsWeights2.7 GB 4dbb1ed31e59
config.jsonConfiguration9.5 KB
generation_config.jsonConfiguration208 B
hf_quant_config.jsonConfiguration3.7 KB
model.safetensors.index.jsonConfiguration174.5 KB
processor_config.jsonConfiguration1.7 KB
README.mdDocumentation7.6 KB
chat_template.jinjaOther16.9 KB
.gitattributesRepository1.6 KB
.quant_summary.txtRepository278.9 KB
tokenizer.jsonTokenizer32.2 MB cc8d3a0ce364
tokenizer_config.jsonTokenizer2.1 KB

License and Download

License
other
Access
Open weights, no gate
Download size
32.6 GB
Download from NVIDIA

Released by NVIDIA through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published32.6 GB
16-bit41.7 GB
8-bit20.9 GB
4-bit10.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare Gemma-4-31B-IT-NVFP4

Questions About Gemma-4-31B-IT-NVFP4

How much GPU memory does Gemma-4-31B-IT-NVFP4 need?

About 50.1 GB at 16-bit and 12.5 GB at 4-bit: the weights (20.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Gemma-4-31B-IT-NVFP4 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Gemma-4-31B-IT-NVFP4 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Gemma-4-31B-IT-NVFP4's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

gpt-oss-20b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 20.9B parameters 131,072 tokens transformers

Model · Text generation

gpt-neox-20b

EleutherAI

GPT-NeoX-20B is a 20 billion parameter autoregressive language model trained on the Pile using the GPT-NeoX library. Its architecture intentionally resembles that of GPT-3, and is almost identical to that of GPT-J- 6B. Its training dataset contains a multitude of English-language texts, reflecting the general-purpose nature of this model. See the accompanying paper for details about model architecture (including how it differs from GPT-3), training procedure, and additional evaluations. Model](https://arxiv.org/abs/2204.06745). For details about the training dataset, see the Pile paper, and its data sheet. Discord](https://discord.gg/zBGx3azzUn), and post them in #release-discussion. Please…

Open weights apache-2.0 20.7B parameters 2,048 tokens transformers

Model · Text generation

Ornith-1.5-35B-A3B-NVFP4

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit 19.5B parameters 262,144 tokens transformers

A refusal-removed (abliterated) build of Qwen's Qwen3.8-Flash-Next, quantized to EXL3 2.50 bpw so the full model runs on a single 24 GB card (RTX 3090 / 4090) using MoE CPU-offload — with the vision tower, MTP head, native 262,144-token context, and the PLE n-gram table all intact. Requires the same MoE CPU-offload setup as the stock 2.50bpw pack. Needs ~59 GB host RAM for the CPU expert tail and a fast NVMe for the streamed n-gram table. Expected on an RTX 3090: ~38 tok/s decode with MTP on (~28 without), ~20 tok/s at 175K depth, ~664 tok/s prefill. See the upstream repo for the full measured ledger; this quant uses the identical flags and layout, so numbers should track closely. Fired on…

Open weights other 22.3B parameters 262,144 tokens

Model · Text generation

Qwen3.6-35B-A3B-NVFP4

NVIDIA

The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.6-35B-A3B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots…

Open weights apache-2.0 18.7B parameters 262,144 tokens Model Optimizer

Model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

NVIDIA

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-Nano-3-30B-A3B-NVFP4 is a quantized version of Nemotron-Nano-3-30B-A3B and is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be…

Open weights other 18.2B parameters 262,144 tokens transformers