SAVRN
Search Contact SAVRN

Open-weight model · Text generation

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

by NVIDIA nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of May 2026.

Parameters17.8B
Context1,048,576
Weights21.6 GB
Licenseother
AccessOpen weights
Monthly Downloads1.2M

Runs On

What it takes to serve NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 (17.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 35.6 GB 42.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 17.8 GB 21.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 8.9 GB 10.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

A context window of 1,048,576 tokens is what an operator should plan around. NVIDIA built it as a hybrid of interleaved Mamba-2 and mixture-of-experts layers with select attention layers, 52 layers deep with 128 routed experts. The weights count 17.8 billion parameters. At 16-bit that is 35.6 GB of weights needing 42.8 GB; at 8-bit, 17.8 GB needing 21.4 GB; at 4-bit, 8.9 GB needing 10.7 GB. One MI300X with 192 GB at $1.85 an hour on demand covers all three with memory to spare.

The license is listed only as other, with no summary in our file, so read NVIDIA's terms before this goes near production. The model card reports 52.8 percent resolved on SWE-bench Verified and 81.62 on MMLU-Pro; those are NVIDIA's figures, not ours. Training data stops at September 2025 for pre-training and May 2026 for post-training. No Index host prices it per token.

Model Card

The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of May 2026. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.…

Excerpt from the card by NVIDIA, licensed other.

Configuration

Architecture
NemotronHForCausalLM
Context length (tokens)
1,048,576
Layers
52
Hidden size
2,688
Feed-forward size
1,856
Attention heads
32
Key/value heads
2
Head dimension
128
Vocabulary size
131,072
Routed experts
128
Experts active per token
6
RoPE base
10,000
Model type
nemotron_h
Quantization
modelopt

Identity and Version

Repository
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
Publisher
NVIDIA
Task
Text generation
Modality
Text
Library
transformers
Parameters
17.8B parameters
Languages
en, es, fr, de, it, ja
Revision
bee7596271d1495f6992ae224aefde4410e816b8
First published
2026-08-04
Last updated
2026-09-10

Files and Weights

70 files, 21.6 GB in total. The weights are 52 files totalling 21.6 GB in safetensors.

Weights52 files · 21.6 GB
Configuration6 files · 4.2 MB
Tokenizer2 files · 17.3 MB
Documentation6 files · 101.4 KB
Other3 files · 371.4 KB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00052.safetensorsWeights743.4 MB 672c8bda10fd
model-00002-of-00052.safetensorsWeights731.1 MB b365ac815ea7
model-00003-of-00052.safetensorsWeights38.8 MB 8fbad8247cd4
model-00004-of-00052.safetensorsWeights731.1 MB 243a94b1f2b3
model-00005-of-00052.safetensorsWeights38.8 MB e675767a4747
model-00006-of-00052.safetensorsWeights46.8 MB 8f8e9623934b
model-00007-of-00052.safetensorsWeights731.1 MB 306aee6514a7
model-00008-of-00052.safetensorsWeights38.8 MB 31235989ff64
model-00009-of-00052.safetensorsWeights731.1 MB bcd6b2645a6b
model-00010-of-00052.safetensorsWeights38.8 MB 5fae7d86c343
model-00011-of-00052.safetensorsWeights731.1 MB 10237b58cd28
model-00012-of-00052.safetensorsWeights38.8 MB dfa4a2312fc9
model-00013-of-00052.safetensorsWeights46.8 MB 41feb27e3376
model-00014-of-00052.safetensorsWeights731.1 MB 16653126b271
model-00015-of-00052.safetensorsWeights38.8 MB 1016d76b1b0d
model-00016-of-00052.safetensorsWeights731.1 MB 70aaec01374f
model-00017-of-00052.safetensorsWeights38.8 MB 67763384429d
model-00018-of-00052.safetensorsWeights731.1 MB 252e3fb2e6c3
model-00019-of-00052.safetensorsWeights38.8 MB 3fd080a21df9
model-00020-of-00052.safetensorsWeights46.8 MB 31dada97bfd5
model-00021-of-00052.safetensorsWeights731.1 MB 601a05eb0b22
model-00022-of-00052.safetensorsWeights38.8 MB 4b3eb401a315
model-00023-of-00052.safetensorsWeights731.1 MB 1f92f9b4bc06
model-00024-of-00052.safetensorsWeights38.8 MB 64d36b45162f
model-00025-of-00052.safetensorsWeights731.1 MB 1d8f5436a1e7
model-00026-of-00052.safetensorsWeights38.8 MB 97e410e49188
model-00027-of-00052.safetensorsWeights46.8 MB e67fb58638bf
model-00028-of-00052.safetensorsWeights731.1 MB 1800ab07214a
model-00029-of-00052.safetensorsWeights38.8 MB 045dc61dc334
model-00030-of-00052.safetensorsWeights731.1 MB 794679784bde
model-00031-of-00052.safetensorsWeights38.8 MB 727bb1bf4829
model-00032-of-00052.safetensorsWeights731.1 MB cc68cfd7f1e8
model-00033-of-00052.safetensorsWeights38.8 MB ee8d3fc76401
model-00034-of-00052.safetensorsWeights46.8 MB 90c579b134e5
model-00035-of-00052.safetensorsWeights731.1 MB d6b9441c3fb7
model-00036-of-00052.safetensorsWeights38.8 MB 130ee06f0ba0
model-00037-of-00052.safetensorsWeights731.1 MB 7464b7b51237
model-00038-of-00052.safetensorsWeights38.8 MB 1a0ab61ef65a
model-00039-of-00052.safetensorsWeights731.1 MB 9f31f2fa9502
model-00040-of-00052.safetensorsWeights38.8 MB 407f0e74455e
model-00041-of-00052.safetensorsWeights731.1 MB 4e8442377f26
model-00042-of-00052.safetensorsWeights38.8 MB 85ca012bff38
model-00043-of-00052.safetensorsWeights46.8 MB 01c56b23b7ea
model-00044-of-00052.safetensorsWeights731.1 MB a3ebcaf00a08
model-00045-of-00052.safetensorsWeights38.8 MB e9d559259cb0
model-00046-of-00052.safetensorsWeights731.1 MB 5e5cfe7a6e50
model-00047-of-00052.safetensorsWeights38.8 MB 96ffe8358c3e
model-00048-of-00052.safetensorsWeights731.1 MB 1dcdf08d0f47
model-00049-of-00052.safetensorsWeights38.8 MB 25da4593995d
model-00050-of-00052.safetensorsWeights731.1 MB ef5b1baa53ab
model-00051-of-00052.safetensorsWeights38.8 MB 2da1f6cbbcfb
model-00052-of-00052.safetensorsWeights3.6 GB 85db447be6ac
.eval_results/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4.yamlConfiguration1.4 KB
config.jsonConfiguration1.3 MB
generation_config.jsonConfiguration209 B
hf_quant_config.jsonConfiguration928.1 KB
model.safetensors.index.jsonConfiguration1.9 MB
special_tokens_map.jsonConfiguration563 B
LICENSEDocumentation2.7 KB
README.mdDocumentation89.8 KB
bias.mdDocumentation3.3 KB
explainability.mdDocumentation2.7 KB
privacy.mdDocumentation828 B
safety.mdDocumentation2.1 KB
accuracy_plot.pngOther141.5 KB 1397995f8a5d
agentic_coding_benchmarks.pngOther220.0 KB 06cf5486ae94
chat_template.jinjaOther9.9 KB
.gitattributesRepository1.7 KB
tokenizer.jsonTokenizer17.1 MB 623c34567aeb
tokenizer_config.jsonTokenizer177.2 KB

License and Download

License
other
Access
Open weights, no gate
Download size
21.6 GB
Download from NVIDIA

Released by NVIDIA through its official repository on Hugging Face.

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 75.57 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
SWE-bench/SWE-bench_Multilingual Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established 36.47 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 52.8 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 81.62 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
cais/hle Task hleMetric hleSetup Text-only, no tools — HLE's default includes image questions, so this excludes the multimodal subset.Comparison conditions not established 10.47 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13

Memory Requirements

PrecisionWeights in memory
As published21.6 GB
16-bit35.6 GB
8-bit17.8 GB
4-bit8.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Questions About NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

How much GPU memory does NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 need?

About 42.8 GB at 16-bit and 10.7 GB at 4-bit: the weights (17.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

NVIDIA

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-Nano-3-30B-A3B-NVFP4 is a quantized version of Nemotron-Nano-3-30B-A3B and is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be…

Open weights other 18.2B parameters 262,144 tokens transformers

Model · Text generation

Qwen3.6-35B-A3B-NVFP4

NVIDIA

The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.6-35B-A3B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots…

Open weights apache-2.0 18.7B parameters 262,144 tokens Model Optimizer

Model · Text generation

Ornith-1.5-35B-A3B-NVFP4

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit 19.5B parameters 262,144 tokens transformers

Model · Text generation

DeepSeek-Coder-V2-Lite-Instruct

DeepSeek

We present DeepSeek-Coder-V2, an open-source Mixture-of-Experts (MoE) code language model that achieves performance comparable to GPT4-Turbo in code-specific tasks. Specifically, DeepSeek-Coder-V2 is further pre-trained from an intermediate checkpoint of DeepSeek-V2 with additional 6 trillion tokens. Through this continued pre-training, DeepSeek-Coder-V2 substantially enhances the coding and mathematical reasoning capabilities of DeepSeek-V2, while maintaining comparable performance in general language tasks. Compared to DeepSeek-Coder-33B, DeepSeek-Coder-V2 demonstrates significant advancements in various aspects of code-related tasks, as well as reasoning and general capabilities.…

Open weights other 15.7B parameters 163,840 tokens transformers

Model · Text generation

gpt-neox-20b

EleutherAI

GPT-NeoX-20B is a 20 billion parameter autoregressive language model trained on the Pile using the GPT-NeoX library. Its architecture intentionally resembles that of GPT-3, and is almost identical to that of GPT-J- 6B. Its training dataset contains a multitude of English-language texts, reflecting the general-purpose nature of this model. See the accompanying paper for details about model architecture (including how it differs from GPT-3), training procedure, and additional evaluations. Model](https://arxiv.org/abs/2204.06745). For details about the training dataset, see the Pile paper, and its data sheet. Discord](https://discord.gg/zBGx3azzUn), and post them in #release-discussion. Please…

Open weights apache-2.0 20.7B parameters 2,048 tokens transformers

Model · Text generation

Gemma-4-31B-IT-NVFP4

NVIDIA

Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding, and multimodal understanding on consumer GPUs and workstations, with a 256K-token context window and support for over 140 languages. The model uses a hybrid attention mechanism that interleaves local sliding-window and full global attention, with unified Keys and Values in global layers and Proportional RoPE (p-RoPE) to support long-context performance. The NVIDIA Gemma 4 31B IT NVFP4 model is quantized with NVIDIA Model…

Open weights other 20.9B parameters 262,144 tokens Model Optimizer