SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 vs Qwen3.6-35B-A3B-NVFP4

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 has 17.8B parameters and Qwen3.6-35B-A3B-NVFP4 has 18.7B parameters; NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is released under other and Qwen3.6-35B-A3B-NVFP4 under Apache License 2.0; at 16-bit, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 needs about 42.8 GB (1x MI300X from $1.85 an hour) and Qwen3.6-35B-A3B-NVFP4 about 44.8 GB (1x MI300X from $1.85 an hour).

Published metadata for 2 models, each read from its own repository.
Field NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
Qwen3.6-35B-A3B-NVFP4
nvidia/Qwen3.6-35B-A3B-NVFP4
Publisher NVIDIA NVIDIA
Task Text generation Text generation
Modality Text Text
Parameters, as reported 17.8B parameters 18.7B parameters
Architecture NemotronHForCausalLM Qwen3_5MoeForConditionalGeneration
Library transformers Model Optimizer
Context length 1,048,576 tokens 262,144 tokens
Repository size 21.6 GB 23.5 GB
Artifact formats safetensors, pytorch safetensors
License other apache-2.0
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 42.8 GB 44.8 GB
Cheapest GPUs at 16-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Memory at 4-bit (weights and margin) 10.7 GB 11.2 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed bee7596271d1 1355db6a0524
Downloads reported by the hub 1.2M 8.4M
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

Other Reported Results

These results are listed for each model on its own, because the conditions needed to compare them are not stated or do not match. Two results that leave a condition blank are not assumed to share it.

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 75.57 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
SWE-bench/SWE-bench_Multilingual Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established 36.47 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 52.8 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 81.62 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13
cais/hle Task hleMetric hleSetup Text-only, no tools — HLE's default includes image questions, so this excludes the multimodal subset.Comparison conditions not established 10.47 NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model card
Reported by a third party
Evaluated revision not stated 2026-08-13

SAVRN's Notes on NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

A context window of 1,048,576 tokens is what an operator should plan around. NVIDIA built it as a hybrid of interleaved Mamba-2 and mixture-of-experts layers with select attention layers, 52 layers deep with 128 routed experts. The weights count 17.8 billion parameters. At 16-bit that is 35.6 GB of weights needing 42.8 GB; at 8-bit, 17.8 GB needing 21.4 GB; at 4-bit, 8.9 GB needing 10.7 GB. One MI300X with 192 GB at $1.85 an hour on demand covers all three with memory to spare.

The license is listed only as other, with no summary in our file, so read NVIDIA's terms before this goes near production. The model card reports 52.8 percent resolved on SWE-bench Verified and 81.62 on MMLU-Pro; those are NVIDIA's figures, not ours. Training data stops at September 2025 for pre-training and May 2026 for post-training. No Index host prices it per token.

SAVRN's Notes on Qwen3.6-35B-A3B-NVFP4

NVIDIA's part in this one is the quantization, not the model. It ran Alibaba's Qwen3.6-35B-A3B through Model Optimizer, and its own card says NVIDIA neither owns nor developed the result: 18.7B parameters as a mixture of 256 experts, 8 active per token. Our sizing wants 44.8 GB of memory at 16-bit and 11.2 GB at 4-bit; the cheapest setup we list is one 192 GB MI300X at $1.85 per hour on-demand, and the 262,144 token context is where the rest of that card goes on real jobs.

Before committing, read Alibaba's card for the parent Qwen3.6-35B-A3B, since that is what you are deploying. Apache 2.0 covers commercial use, modification and redistribution, provided the license and NOTICE file stay attached and significant changes are stated. Files were last updated August 29, 2026, three months after the May 27 release, so confirm which revision you pulled.

Questions

Which is larger, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 or Qwen3.6-35B-A3B-NVFP4?

Qwen3.6-35B-A3B-NVFP4 (18.7B parameters) is larger than NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 (17.8B parameters), by the parameter counts their publishers report.

Which is cheaper to run, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 or Qwen3.6-35B-A3B-NVFP4?

At 4-bit, NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 fits on 1x MI300X from $1.85 an hour and Qwen3.6-35B-A3B-NVFP4 on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.6-35B-A3B-NVFP4 commercially?

Yes. Qwen3.6-35B-A3B-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Related Comparisons