SAVRN Model Hub · Comparisons
Gemma-4-26B-A4B-NVFP4 vs Qwen2.5-14B-Instruct
Gemma-4-26B-A4B-NVFP4 has 14.4B parameters and Qwen2.5-14B-Instruct has 14.8B parameters; both are released under Apache License 2.0; at 16-bit, Gemma-4-26B-A4B-NVFP4 needs about 34.5 GB (1x MI300X from $1.85 an hour) and Qwen2.5-14B-Instruct about 35.4 GB (1x MI300X from $1.85 an hour).
| Field | Gemma-4-26B-A4B-NVFP4 nvidia/Gemma-4-26B-A4B-NVFP4 | Qwen2.5-14B-Instruct Qwen/Qwen2.5-14B-Instruct |
|---|---|---|
| Publisher | NVIDIA | Qwen |
| Task | Text generation | Text generation |
| Modality | Text | Text |
| Parameters, as reported | 14.4B parameters | 14.8B parameters |
| Architecture | Gemma4ForConditionalGeneration | Qwen2ForCausalLM |
| Library | Model Optimizer | transformers |
| Context length | 262,144 tokens | 32,768 tokens |
| Repository size | 18.8 GB | 29.6 GB |
| Artifact formats | safetensors | safetensors |
| License | apache-2.0 | apache-2.0 |
| Access | Open weights, no gate | Open weights, no gate |
| Memory at 16-bit (weights and margin) | 34.5 GB | 35.4 GB |
| Cheapest GPUs at 16-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Memory at 4-bit (weights and margin) | 8.6 GB | 8.9 GB |
| Cheapest GPUs at 4-bit, per hour | 1x MI300X, $1.85 | 1x MI300X, $1.85 |
| Revision viewed | a19cfe00be84 | cf98f3b3bbb4 |
| Downloads reported by the hub | 1.8M | 2.4M |
| Last observed | 2026-09-18 | 2026-09-18 |
An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.
SAVRN's Notes on Gemma-4-26B-A4B-NVFP4
NVIDIA cut this build from Google's gemma-4-26B-A4B-it with its Model Optimizer tooling. It generates text from text, image and frame-by-frame video input over a 262,144 token context. The 4-bit row needs 8.6 GB of memory and the 16-bit row 34.5 GB, and the cheapest host for either is one 192 GB MI300X at $1.85 per hour on demand, which leaves most of the card free for that context.
Nothing in Apache 2.0 blocks a commercial deployment or a modified fork; keep the license and NOTICE file with it and state significant changes. Reconcile the parameter count first: the page counts 14.4 billion while the name carries 26B and A4B, and the base model page is where to settle it. Released May 1, 2026, this file has no per-token host price in the SAVRN Index yet, so $1.85 an hour is the cost figure to plan on.
SAVRN's Notes on Qwen2.5-14B-Instruct
Load the 16-bit weights and Qwen2.5-14B-Instruct wants 35.4 GB of memory. The cheapest slot we track is one MI300X with 192 GB at $1.85 an hour on demand, so a single card carries it with most of its memory unused. At 8-bit the need drops to 17.7 GB and at 4-bit to 8.9 GB, where this 14.8 billion parameter text generation build fits beside other work on the same card.
Apache License 2.0 lets us run it commercially, modify it and redistribute it, provided the notices travel with every copy and significant changes are stated. Two checks before committing: the configuration lists a 32,768-token context beside a 131,072-token sliding window and cites arXiv:2309.00071 on context extension, so pin down which limit your serving stack honors; and it derives from Qwen/Qwen2.5-14B, so confirm instruct tuning suits the workload. The SAVRN Index lists no per-token host price for it.
Questions
Which is larger, Gemma-4-26B-A4B-NVFP4 or Qwen2.5-14B-Instruct?
Qwen2.5-14B-Instruct (14.8B parameters) is larger than Gemma-4-26B-A4B-NVFP4 (14.4B parameters), by the parameter counts their publishers report.
Which is cheaper to run, Gemma-4-26B-A4B-NVFP4 or Qwen2.5-14B-Instruct?
At 4-bit, Gemma-4-26B-A4B-NVFP4 fits on 1x MI300X from $1.85 an hour and Qwen2.5-14B-Instruct on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Gemma-4-26B-A4B-NVFP4 commercially?
Yes. Gemma-4-26B-A4B-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Can I use Qwen2.5-14B-Instruct commercially?
Yes. Qwen2.5-14B-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.