SAVRN
Search Contact SAVRN

Open-weight model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

by NVIDIA nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Parameters31.6B
Context262,144
Weights63.2 GB
Licenseother
AccessOpen weights
Monthly Downloads704.5k

Runs On

What it takes to serve NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (31.6B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 63.2 GB 75.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 31.6 GB 37.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 15.8 GB 18.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

Seventy-six gigabytes is the number that decides the hardware here. At 16-bit the 31.6B parameters need 75.8 GB, which fits on a single MI300X with 192 GB at $1.85 per hour, the cheapest setup the Index prices, with room left for the 262,144-token context. At 8-bit the need drops to 37.9 GB and at 4-bit to 18.9 GB, and all 128 routed experts are in that bill. NVIDIA trained it for reasoning and non-reasoning work: it writes a reasoning trace before its answer, and a chat template flag turns the trace off.

The license field says other and we have no summary of its terms on file, so read NVIDIA's license in full before any commercial deployment. Two more checks: the one reported score, 86.8 on Liquid AI's IFStruct v1.0, is third-party reported, and pre-training data stops at June 25, 2025. No Index host serves it by the token yet.

Model Card

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-3-Nano-30B-A3B-BF16 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in…

Excerpt from the card by NVIDIA, licensed other.

Configuration

Architecture
NemotronHForCausalLM
Context length (tokens)
262,144
Layers
52
Hidden size
2,688
Feed-forward size
1,856
Attention heads
32
Key/value heads
2
Head dimension
128
Vocabulary size
131,072
Routed experts
128
Experts active per token
6
RoPE base
10,000
Stored precision
bfloat16
Model type
nemotron_h

Identity and Version

Repository
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Publisher
NVIDIA
Task
Text generation
Modality
Text
Library
transformers
Parameters
31.6B parameters
Languages
en, es, fr, de, ja, it
Revision
bf77c3174f68ad409e1c2aa60daeb46e32d1c606
First published
2025-12-04
Last updated
2026-08-24

Files and Weights

35 files, 63.2 GB in total. The weights are 13 files totalling 63.2 GB in safetensors.

Weights13 files · 63.2 GB
Configuration11 files · 717.9 KB
Tokenizer2 files · 17.3 MB
Documentation5 files · 83.1 KB
Other3 files · 204.7 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00013.safetensorsWeights5.0 GB 4c77b0f1717f
model-00002-of-00013.safetensorsWeights5.0 GB 2e3de804d8c8
model-00003-of-00013.safetensorsWeights5.0 GB e113d2a3f815
model-00004-of-00013.safetensorsWeights5.0 GB c64af3570422
model-00005-of-00013.safetensorsWeights5.0 GB 1aa5867c6483
model-00006-of-00013.safetensorsWeights5.0 GB fd411a714fad
model-00007-of-00013.safetensorsWeights5.0 GB d16bc0bd0521
model-00008-of-00013.safetensorsWeights5.0 GB fc0aea38d897
model-00009-of-00013.safetensorsWeights5.0 GB 34b4715b5765
model-00010-of-00013.safetensorsWeights5.0 GB 4124abfaa922
model-00011-of-00013.safetensorsWeights5.0 GB d3bf1c127982
model-00012-of-00013.safetensorsWeights5.0 GB 4abbf8125860
model-00013-of-00013.safetensorsWeights3.2 GB 9458d10c7e99
.eval_results/gpqa.yamlConfiguration153 B
.eval_results/hle.yamlConfiguration148 B
.eval_results/mmlu-pro.yamlConfiguration158 B
config.jsonConfiguration1.8 KB
configuration_nemotron_h.pyConfiguration12.9 KB
generation_config.jsonConfiguration197 B
model.safetensors.index.jsonConfiguration613.3 KB
modeling_nemotron_h.pyConfiguration83.8 KB
nano_v3_reasoning_parser.pyConfiguration798 B
nemo-evaluator-launcher-configs/local_nvidia_nemotron_3_nano_30b_a3b.yamlConfiguration4.2 KB
special_tokens_map.jsonConfiguration420 B
README.mdDocumentation73.4 KB
bias.mdDocumentation2.3 KB
explainability.mdDocumentation3.0 KB
privacy.mdDocumentation2.3 KB
safety.mdDocumentation2.1 KB
accuracy_chart.pngOther191.0 KB 5fc15c897292
chat_template.jinjaOther10.5 KB
notebook.ipynbOther3.2 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer17.1 MB c6021eb6847e
tokenizer_config.jsonTokenizer188.0 KB

License and Download

License
other
Access
Open weights, no gate
Download size
63.2 GB
Download from NVIDIA

Released by NVIDIA through its official repository on Hugging Face.

Built From

  • Described by arXiv:2512.20848
  • Described by arXiv:2512.20856
  • Trained on (disclosed) nvidia/Nemotron-3-Nano-RL-Training-Blend
  • Trained on (disclosed) nvidia/Nemotron-Agentic-v1
  • Trained on (disclosed) nvidia/Nemotron-CC-Code-v1
  • Trained on (disclosed) nvidia/Nemotron-CC-Math-v1
  • Trained on (disclosed) nvidia/Nemotron-CC-v2
  • Trained on (disclosed) nvidia/Nemotron-CC-v2.1
  • Trained on (disclosed) nvidia/Nemotron-Competitive-Programming-v1
  • Trained on (disclosed) nvidia/Nemotron-Instruction-Following-Chat-v1
  • Trained on (disclosed) nvidia/Nemotron-Math-Proofs-v1
  • Trained on (disclosed) nvidia/Nemotron-Math-v2
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Code-v1
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Code-v2
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Dataset-sample
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-SFT-v1
  • Trained on (disclosed) nvidia/Nemotron-Pretraining-Specialized-v1
  • Trained on (disclosed) nvidia/Nemotron-Science-v1

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
LiquidAI/ifstruct-v1.0 Task ifstruct_v1Metric ifstruct_v1Comparison conditions not established 86.8 Liquid AI — IFStruct v1.0 blog (Nemotron-3-Nano-30B-A3B)
Reported by a third party
Evaluated revision not stated 2026-06-30

Memory Requirements

PrecisionWeights in memory
As published63.2 GB
16-bit63.2 GB
8-bit31.6 GB
4-bit15.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Compare NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

Questions About NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

How much GPU memory does NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 need?

About 75.8 GB at 16-bit and 18.9 GB at 4-bit: the weights (31.6B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is NVIDIA-Nemotron-3-Nano-30B-A3B-BF16's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Fastino-Nemotron-3.5-Lightning-Finance is a 30B-parameter, 3B-active mixture-of-experts model specialized for financial reasoning, extraction, and research fine-tuned on LoRA with the Fastino Fine-Tuning Agent. The model targets financial document reasoning, numerical question answering over filings and tables, numeric span extraction, financial entity recognition, conversational analysis, and source-grounded financial research. The evaluation suite includes FinQA, TAT-QA, SEC-Num, FinEntity, BizFinBench, BigFinanceBench, ConvFinQA, and FiQA. The published weights are BF16 and require about 66 GB before runtime overhead. An 80 GB or larger GPU, or tensor parallelism across multiple GPUs, is…

Open weights apache-2.0 31.6B parameters 262,144 tokens transformers

Model · Text generation

OTel-2.0-LLM-31B-IT

Farbod Tavakkoli

OTel-2.0-LLM-31B-IT is a telecom-specialized instruction model post-trained from Gemma 4 31B-IT on approximately 440 billion telecom training tokens. It is the first release in the OTel 2.0 family and is designed to support telco-grade AI workflows across network operations, standards interpretation, product development, network configuration assistance, RAG, and telecom-specific question answering. OTel 2.0 extends the original OTel effort from a RAG-oriented telecom fine-tuning release into a larger domain-adapted training program. The model was trained from a much larger standards and telecom corpus, with new data preparation coverage for direct telecom QnA, abstention, RAG…

Open weights apache-2.0 31.3B parameters 262,144 tokens transformers

Model · Text generation

GLM-4.7-Flash

Z.ai

Join our Discord community. Check out the GLM-4.7 technical blog, technical report(GLM-4.5). Use GLM-4.7-Flash API services on Z.ai API Platform. One click to GLM-4.7. GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency. Default Settings (Most Tasks) For multi-turn agentic tasks (τ²-Bench and Terminal Bench 2), please turn on Preserved Thinking mode. Terminal Bench, SWE Bench Verified τ^2-Bench For τ^2-Bench evaluation, we added an additional prompt to the Retail and Telecom user interaction to avoid failure modes caused by users ending the interaction…

Open weights mit 31.2B parameters 202,752 tokens transformers

Model · Text generation

Qwen3-VL-30B-A3B-Instruct-AWQ

QuantTrio

As of 2025-10-08, create a fresh Python environment and run: For more details, refer to vLLM Official Qwen3-VL Guide Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision…

Open weights apache-2.0 31.1B parameters 262,144 tokens transformers

Model · Text generation

granite-4.0-h-small-w8a8-llmcompressor

AMD

ZenDNN v6.1.0 - ZenTorch v2.13.0.0 - PyTorch v2.13.0.0 - LLM Compressor v0.13.0 - vLLM v0.29.0 This is a quantized version of granite-4.0-h-small created by AMD using LLM Compressor (compressed-tensors) for ZenDNN-optimized CPU inference. The model was quantized from granite-4.0-h-small using LLM Compressor via the Round-to-Nearest (RTN) algorithm. This reduces the model weights from 60.0 GiB to 30.4 GiB on disk (~49% reduction). granite-4.0-h-small is a hybrid Mamba-MoE model: of its 40 layers, 4 are full-attention blocks and the other 36 are Mamba (linear-attention) blocks, and every layer carries a 72-expert MoE block (top-10 routing) alongside a shared MLP. The recipe only needs two…

Open weights apache-2.0 32.2B parameters 131,072 tokens transformers