SAVRN Model Hub · Comparisons
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 vs Qwen-72B
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 has 67.2B parameters and Qwen-72B has 72.3B parameters; both are released under other; at 16-bit, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 needs about 161.3 GB (1x MI300X from $1.85 an hour) and Qwen-72B about 173.5 GB (1x MI300X from $1.85 an hour).
| Field | NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 | Qwen-72B Qwen/Qwen-72B |
|---|---|---|
| Publisher | NVIDIA | Qwen |
| Task | Text generation | Text generation |
| Modality | Text | Text |
| Parameters, as reported | 67.2B parameters | 72.3B parameters |
| Architecture | NemotronHForCausalLM | QWenLMHeadModel |
| Library | transformers | transformers |
| Context length | 262,144 tokens | 32,768 tokens |
| Repository size | 80.4 GB | 144.6 GB |
| Artifact formats | safetensors, pytorch | safetensors |
| License | other | other |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 161.3 GB | 173.5 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 40.3 GB | 43.4 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | ff433f5493e2 | b8e18ac61df6 |
| Downloads reported by the hub | 710.1k | 4M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
SAVRN's Notes on NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
Even at 16-bit this one fits on a single card: 161.3 GB needed against the 192 GB of the MI300X in the cheapest listed setup, $1.85 per hour. At 8-bit the need drops to 80.7 GB, at 4-bit to 40.3 GB, which opens smaller accelerators or several copies per card. NVIDIA built it for agentic and conversational work at high volume, with 512 routed experts across 88 layers and a 262,144-token context, so long agent sessions are the intended load.
The license reads other, with no summary in our facts, so the terms are whatever NVIDIA publishes with the files: read them before any commercial deployment. Check the precision: the files total 80.4 GB and the name carries NVFP4, so confirm what you are sizing memory for. The training data is named, nvidia/nemotron-pre-training-datasets and nvidia/nemotron-post-training-v3, with cutoffs of June 2025 and February 2026, and two papers describe the design.
Questions
Which is larger, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 or Qwen-72B?
Qwen-72B (72.3B parameters) is larger than NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 (67.2B parameters), by the parameter counts their publishers report.
Which is cheaper to run, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 or Qwen-72B?
At 4-bit, NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 fits on 1x MI300X from $1.85 an hour and Qwen-72B on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.