SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 vs Qwen-72B

NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 has 67.2B parameters and Qwen-72B has 72.3B parameters; both are released under other; at 16-bit, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 needs about 161.3 GB (1x MI300X from $1.85 an hour) and Qwen-72B about 173.5 GB (1x MI300X from $1.85 an hour).

Published metadata for 2 models, each read from its own repository.
Field NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
Qwen-72B
Qwen/Qwen-72B
Publisher NVIDIA Qwen
Task Text generation Text generation
Modality Text Text
Parameters, as reported 67.2B parameters 72.3B parameters
Architecture NemotronHForCausalLM QWenLMHeadModel
Library transformers transformers
Context length 262,144 tokens 32,768 tokens
Repository size 80.4 GB 144.6 GB
Artifact formats safetensors, pytorch safetensors
License other other
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 161.3 GB 173.5 GB
Cheapest GPUs at 16-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Memory at 4-bit (weights and margin) 40.3 GB 43.4 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed ff433f5493e2 b8e18ac61df6
Downloads reported by the hub 710.1k 4M
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

SAVRN's Notes on NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Even at 16-bit this one fits on a single card: 161.3 GB needed against the 192 GB of the MI300X in the cheapest listed setup, $1.85 per hour. At 8-bit the need drops to 80.7 GB, at 4-bit to 40.3 GB, which opens smaller accelerators or several copies per card. NVIDIA built it for agentic and conversational work at high volume, with 512 routed experts across 88 layers and a 262,144-token context, so long agent sessions are the intended load.

The license reads other, with no summary in our facts, so the terms are whatever NVIDIA publishes with the files: read them before any commercial deployment. Check the precision: the files total 80.4 GB and the name carries NVFP4, so confirm what you are sizing memory for. The training data is named, nvidia/nemotron-pre-training-datasets and nvidia/nemotron-post-training-v3, with cutoffs of June 2025 and February 2026, and two papers describe the design.

Questions

Which is larger, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 or Qwen-72B?

Qwen-72B (72.3B parameters) is larger than NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 (67.2B parameters), by the parameter counts their publishers report.

Which is cheaper to run, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 or Qwen-72B?

At 4-bit, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 fits on 1x MI300X from $1.85 an hour and Qwen-72B on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Related Comparisons