SAVRN
Search Contact SAVRN

Open-weight model · Text generation

NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

by NVIDIA nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025.

Parameters67.2B
Context262,144
Weights80.3 GB
Licenseother
AccessOpen weights
Monthly Downloads710.1k

Runs On

What it takes to serve NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 (67.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 134.5 GB 161.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 67.2 GB 80.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 33.6 GB 40.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Even at 16-bit this one fits on a single card: 161.3 GB needed against the 192 GB of the MI300X in the cheapest listed setup, $1.85 per hour. At 8-bit the need drops to 80.7 GB, at 4-bit to 40.3 GB, which opens smaller accelerators or several copies per card. NVIDIA built it for agentic and conversational work at high volume, with 512 routed experts across 88 layers and a 262,144-token context, so long agent sessions are the intended load.

The license reads other, with no summary in our facts, so the terms are whatever NVIDIA publishes with the files: read them before any commercial deployment. Check the precision: the files total 80.4 GB and the name carries NVFP4, so confirm what you are sizing memory for. The training data is named, nvidia/nemotron-pre-training-datasets and nvidia/nemotron-post-training-v3, with cutoffs of June 2025 and February 2026, and two papers describe the design.

Model Card

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-NVFP4 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Excerpt from the card by NVIDIA, licensed other.

Configuration

Architecture
NemotronHForCausalLM
Context length (tokens)
262,144
Layers
88
Hidden size
4,096
Feed-forward size
2,688
Attention heads
32
Key/value heads
2
Head dimension
128
Vocabulary size
131,072
Routed experts
512
Experts active per token
22
RoPE base
10,000
Model type
nemotron_h
Quantization
modelopt

Identity and Version

Repository
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
Publisher
NVIDIA
Task
Text generation
Modality
Text
Library
transformers
Parameters
67.2B parameters
Languages
en, fr, es, it, de, ja, zh
Revision
ff433f5493e25d631c9f12b5d55c674229923d02
First published
2026-03-10
Last updated
2026-08-24

Files and Weights

36 files, 80.4 GB in total. The weights are 17 files totalling 80.3 GB in safetensors.

Weights17 files · 80.3 GB
Configuration9 files · 30.3 MB
Tokenizer2 files · 17.3 MB
Documentation5 files · 92.2 KB
Other2 files · 89.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00017.safetensorsWeights5.0 GB 3d31c241fb77
model-00002-of-00017.safetensorsWeights5.0 GB 78d4a5ad2684
model-00003-of-00017.safetensorsWeights5.0 GB ecb72debb5d2
model-00004-of-00017.safetensorsWeights5.0 GB a7d1c68e8272
model-00005-of-00017.safetensorsWeights5.0 GB 1d432c9c7834
model-00006-of-00017.safetensorsWeights5.0 GB 037ea60b87cb
model-00007-of-00017.safetensorsWeights5.0 GB 3c0527dac96f
model-00008-of-00017.safetensorsWeights5.0 GB ef6e73f64c3d
model-00009-of-00017.safetensorsWeights5.0 GB b4f598068350
model-00010-of-00017.safetensorsWeights5.0 GB bf8bf1f4e2e0
model-00011-of-00017.safetensorsWeights5.0 GB ac436f2e1af8
model-00012-of-00017.safetensorsWeights5.0 GB 7814c9ae0dc3
model-00013-of-00017.safetensorsWeights5.0 GB b13aa79c0879
model-00014-of-00017.safetensorsWeights5.0 GB 1d8781b2643c
model-00015-of-00017.safetensorsWeights5.0 GB 5f6c3c0f8420
model-00016-of-00017.safetensorsWeights5.0 GB a5f795056a1a
model-00017-of-00017.safetensorsWeights345.8 MB 218355cf69f1
__init__.pyConfiguration
config.jsonConfiguration7.4 MB
configuration_nemotron_h.pyConfiguration19.8 KB
generation_config.jsonConfiguration210 B
hf_quant_config.jsonConfiguration6.1 MB
model.safetensors.index.jsonConfiguration16.6 MB e0ca77ceed3a
modeling_nemotron_h.pyConfiguration82.3 KB
special_tokens_map.jsonConfiguration563 B
super_v3_reasoning_parser.pyConfiguration1.9 KB
README.mdDocumentation81.6 KB
bias.mdDocumentation2.6 KB
explainability.mdDocumentation3.2 KB
privacy.mdDocumentation2.7 KB
safety.mdDocumentation2.1 KB
accuracy_chart.pngOther79.2 KB
chat_template.jinjaOther10.8 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer17.1 MB 623c34567aeb
tokenizer_config.jsonTokenizer177.2 KB

License and Download

License
other
Access
Open weights, no gate
Download size
80.3 GB
Download from NVIDIA

Released by NVIDIA through its official repository on Hugging Face.

Built From

  • Described by arXiv:2512.20848
  • Described by arXiv:2512.20856
  • Trained on (disclosed) nvidia/nemotron-post-training-v3
  • Trained on (disclosed) nvidia/nemotron-pre-training-datasets

Memory Requirements

PrecisionWeights in memory
As published80.3 GB
16-bit134.5 GB
8-bit67.2 GB
4-bit33.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Questions About NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

How much GPU memory does NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 need?

About 161.3 GB at 16-bit and 40.3 GB at 4-bit: the weights (67.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

Qwen3.5-122B-A10B-NVFP4

NVIDIA

The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.5-122B-A10B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems…

Open weights apache-2.0 64.6B parameters 262,144 tokens Model Optimizer

Model · Text generation

Llama-3.3-70B-Instruct

Meta Llama

The Meta Llama 3.3 multilingual large language model (LLM) is an instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.3 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and…

Access requested at publisher llama3.3 70.6B parameters transformers

Model · Text generation

Qwen-72B

Qwen

通义千问-72B(Qwen-72B)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B的仓库。 通义千问-72B(Qwen-72B)主要有以下特点: 1. 大规模高质量训练语料:使用超过3万亿tokens的数据进行预训练,包含高质量中、英、多语言、代码、数学等数据,涵盖通用及专业领域的训练语料。通过大量对比实验对预训练语料分布进行了优化。 2. 强大的性能:Qwen-72B在多个中英文下游评测任务上(涵盖常识推理、代码、数学、翻译等),效果显著超越现有的开源模型。具体评测结果请详见下文。 3. 覆盖更全面的词表:相比目前以中英词表为主的开源模型,Qwen-72B使用了约15万大小的词表。该词表对多语言更加友好,方便用户在不扩展词表的情况下对部分语种进行能力增强和扩展。 4. 较长的上下文支持:Qwen-72B支持32k的上下文长度。 Qwen-72B is the 72B-parameter version of the large language model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-72B is a Transformer-based large…

Open weights other 72.3B parameters 32,768 tokens transformers

Model · Text generation

Qwen3-Coder-Next-FP8

Qwen

Today, we're announcing Qwen3-Coder-Next-FP8, an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements: Qwen3-Coder-Next-FP8 has the following features: NOTE: This model supports only non-thinking mode and does not generate blocks in its output. Meanwhile, specifying enablethinking=False is no longer required. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. We advise you to use the latest version of transformers. The following contains a code snippet illustrating how to use the model generate content based on…

Open weights apache-2.0 79.7B parameters 262,144 tokens transformers

Model · Text generation

Qwen3.6-35B-A3B-abliterated-v4

CS

Uncensored version of Qwen/Qwen3.6-35B-A3B with refusal behavior removed via abliteration (norm-preserving orthogonalization). Zero refusals on harmful prompts. No false refusals on harmless prompts. Abliteration identifies the "refusal direction" in the model's residual stream — the linear direction that activates when the model decides to refuse — and surgically removes it from all output projection weights using norm-preserving orthogonalization. 1. Collect residual stream activations (last token position) for 512 harmful + 512 harmless prompts across all 40 layers 2. Compute mean difference vector per layer → this is the "refusal direction" candidate 3. Score layers by…

Open weights apache-2.0 34.7B parameters 262,144 tokens transformers

Model · Text generation

dolphin-2.9.1-yi-1.5-34b

Dolphin

Curated and trained by Eric Hartford, Lucas Atkins, and Fernando Fernandes, and Cognitive Computations This is our most spectacular outcome ever. FFT, all parameters, 16bit. 77.4 MMLU on 34b. And it talks like a dream. Although the max positional embeddings is 4k, we used rope theta of 1000000.0 and we trained with sequence length 8k. We plan to train on the upcoming 32k version as well. Our appreciation for the sponsors of Dolphin 2.9.1: - Crusoe Cloud - provided excellent on-demand 8xH100 node - OnDemand - provided inference sponsorship This model is based on Yi-1.5-34b, and is governed by apache 2.0 license. The base model has 4k context, but we used rope theta of 1000000.0 and the…

Open weights apache-2.0 34.4B parameters 8,192 tokens transformers