SAVRN
Search Contact SAVRN

Open-weight model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

by NVIDIA nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Parameters18.2B
Context262,144
Weights19.3 GB
Licenseother
AccessOpen weights
Monthly Downloads730.6k

Runs On

What it takes to serve NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 (18.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 36.5 GB 43.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 18.2 GB 21.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 9.1 GB 10.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

Start with the context: 262,144 tokens. NVIDIA built Nemotron 3 Nano as a single model for both reasoning and non-reasoning work: it writes a reasoning trace first, then the answer, and a flag in the chat template turns the reasoning on or off. This NVFP4 build is the quantized version of the BF16 release, with 128 routed experts. At 4-bit the weights are 9.1 GB and the run needs 10.9 GB, so the cheapest listed setup, one MI300X with 192 GB at $1.85 an hour, leaves most of the card for context.

The license is listed as other with no summary on the page, so read NVIDIA's terms yourself before this goes near a product. Check the dates too: pre-training cutoff June 25, 2025, post-training November 28, 2025. The seven Nemotron training sets are named on the page, so you can see what went in.

Model Card

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-Nano-3-30B-A3B-NVFP4 is a quantized version of Nemotron-Nano-3-30B-A3B and is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be…

Excerpt from the card by NVIDIA, licensed other.

Configuration

Architecture
NemotronHForCausalLM
Context length (tokens)
262,144
Layers
52
Hidden size
2,688
Feed-forward size
1,856
Attention heads
32
Key/value heads
2
Head dimension
128
Vocabulary size
131,072
Routed experts
128
Experts active per token
6
RoPE base
10,000
Stored precision
bfloat16
Model type
nemotron_h

Identity and Version

Repository
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
Publisher
NVIDIA
Task
Text generation
Modality
Text
Library
transformers
Parameters
18.2B parameters
Languages
en, es, fr, de, ja, it
Revision
6efb4a2a1c1fa277ce7b3df7a1416255011b1c99
First published
2025-12-20
Last updated
2026-08-24

Files and Weights

18 files, 19.4 GB in total. The weights are 5 files totalling 19.3 GB in safetensors.

Weights5 files · 19.3 GB
Configuration8 files · 2.6 MB
Tokenizer2 files · 17.3 MB
Documentation1 file · 73.7 KB
Other1 file · 10.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00005.safetensorsWeights4.0 GB 2fdac76b3e49
model-00002-of-00005.safetensorsWeights4.0 GB 559806ee0cb6
model-00003-of-00005.safetensorsWeights4.0 GB d82084978870
model-00004-of-00005.safetensorsWeights4.0 GB f5ccb7cfa787
model-00005-of-00005.safetensorsWeights3.3 GB c9dd91428393
config.jsonConfiguration1.8 KB
configuration_nemotron_h.pyConfiguration12.9 KB
generation_config.jsonConfiguration197 B
hf_quant_config.jsonConfiguration3.0 KB
model.safetensors.index.jsonConfiguration2.5 MB
modeling_nemotron_h.pyConfiguration83.8 KB
nano_v3_reasoning_parser.pyConfiguration798 B
special_tokens_map.jsonConfiguration420 B
README.mdDocumentation73.7 KB
chat_template.jinjaOther10.5 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer17.1 MB c6021eb6847e
tokenizer_config.jsonTokenizer188.0 KB

License and Download

License
other
Access
Open weights, no gate
Download size
19.3 GB
Download from NVIDIA

Released by NVIDIA through its official repository on Hugging Face.

Built From

  • Derived from nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
  • Described by arXiv:2512.20848
  • Described by arXiv:2512.20856
  • Described by arXiv:2601.20088
  • Quantized from nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
  • Trained on (disclosed) nvidia/Nemotron-3-Nano-RL-Training-Blend
  • Trained on (disclosed) nvidia/Nemotron-Agentic-v1
  • Trained on (disclosed) nvidia/Nemotron-CC-Code-v1
  • Trained on (disclosed) nvidia/Nemotron-CC-Math-v1
  • Trained on (disclosed) nvidia/Nemotron-CC-v2
  • Trained on (disclosed) nvidia/Nemotron-CC-v2.1
  • Trained on (disclosed) nvidia/Nemotron-Competitive-Programming-v1
  • Trained on (disclosed) nvidia/Nemotron-Instruction-Following-Chat-v1
  • Trained on (disclosed) nvidia/Nemotron-Math-Proofs-v1
  • Trained on (disclosed) nvidia/Nemotron-Math-v2
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Code-v1
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Code-v2
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Dataset-sample
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-SFT-v1
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Specialized-v1
  • Trained on (disclosed) nvidia/Nemotron-Science-v1

Memory Requirements

PrecisionWeights in memory
As published19.3 GB
16-bit36.5 GB
8-bit18.2 GB
4-bit9.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

Questions About NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

How much GPU memory does NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 need?

About 43.8 GB at 16-bit and 10.9 GB at 4-bit: the weights (18.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of May 2026. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.…

Open weights other 17.8B parameters 1,048,576 tokens transformers

Model · Text generation

Qwen3.6-35B-A3B-NVFP4

NVIDIA

The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.6-35B-A3B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots…

Open weights apache-2.0 18.7B parameters 262,144 tokens Model Optimizer

Model · Text generation

Ornith-1.5-35B-A3B-NVFP4

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit 19.5B parameters 262,144 tokens transformers

Model · Text generation

gpt-neox-20b

EleutherAI

GPT-NeoX-20B is a 20 billion parameter autoregressive language model trained on the Pile using the GPT-NeoX library. Its architecture intentionally resembles that of GPT-3, and is almost identical to that of GPT-J- 6B. Its training dataset contains a multitude of English-language texts, reflecting the general-purpose nature of this model. See the accompanying paper for details about model architecture (including how it differs from GPT-3), training procedure, and additional evaluations. Model](https://arxiv.org/abs/2204.06745). For details about the training dataset, see the Pile paper, and its data sheet. Discord](https://discord.gg/zBGx3azzUn), and post them in #release-discussion. Please…

Open weights apache-2.0 20.7B parameters 2,048 tokens transformers

Model · Text generation

DeepSeek-Coder-V2-Lite-Instruct

DeepSeek

We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, DeepSeek-Coder-V2 is further pre-trained from an intermediate checkpoint of DeepSeek-V2 with additional 6 trillion tokens. Through this continued pre-training, DeepSeek-Coder-V2 substantially enhances the coding and mathematical reasoning capabilities of DeepSeek-V2, while maintaining comparable performance in general language tasks. Compared to DeepSeek-Coder-33B, DeepSeek-Coder-V2 demonstrates significant advancements in various aspects of code-related tasks, as well as reasoning and general capabilities.…

Open weights other 15.7B parameters 163,840 tokens transformers

Model · Text generation

Gemma-4-31B-IT-NVFP4

NVIDIA

Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding, and multimodal understanding on consumer GPUs and workstations, with a 256K-token context window and support for over 140 languages. The model uses a hybrid attention mechanism that interleaves local sliding-window and full global attention, with unified Keys and Values in global layers and Proportional RoPE (p-RoPE) to support long-context performance. The NVIDIA Gemma 4 31B IT NVFP4 model is quantized with NVIDIA Model…

Open weights other 20.9B parameters 262,144 tokens Model Optimizer