SAVRN
Search Contact SAVRN

SAVRN Model Hub · Comparisons

Gemma-4-26B-A4B-NVFP4 vs Qwen2.5-14B-Instruct

Gemma-4-26B-A4B-NVFP4 has 14.4B parameters and Qwen2.5-14B-Instruct has 14.8B parameters; both are released under Apache License 2.0; at 16-bit, Gemma-4-26B-A4B-NVFP4 needs about 34.5 GB (1x MI300X from $1.85 an hour) and Qwen2.5-14B-Instruct about 35.4 GB (1x MI300X from $1.85 an hour).

Published metadata for 2 models, each read from its own repository.
Field Gemma-4-26B-A4B-NVFP4
nvidia/Gemma-4-26B-A4B-NVFP4
Qwen2.5-14B-Instruct
Qwen/Qwen2.5-14B-Instruct
Publisher NVIDIA Qwen
Task Text generation Text generation
Modality Text Text
Parameters, as reported 14.4B parameters 14.8B parameters
Architecture Gemma4ForConditionalGeneration Qwen2ForCausalLM
Library Model Optimizer transformers
Context length 262,144 tokens 32,768 tokens
Repository size 18.8 GB 29.6 GB
Artifact formats safetensors safetensors
License apache-2.0 apache-2.0
Access Open weights, no gate Open weights, no gate
Memory at 16-bit (weights and margin) 34.5 GB 35.4 GB
Cheapest GPUs at 16-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Memory at 4-bit (weights and margin) 8.6 GB 8.9 GB
Cheapest GPUs at 4-bit, per hour 1x MI300X, $1.85 1x MI300X, $1.85
Revision viewed a19cfe00be84 cf98f3b3bbb4
Downloads reported by the hub 1.8M 2.4M
Last observed 2026-09-18 2026-09-18

An evaluation row appears only where at least two of these models report the same benchmark with the same stated configuration, metric, unit and setup. Different evaluators stay named in each cell. Values are shown as reported: no unit conversion, no ranking.

SAVRN's Notes on Gemma-4-26B-A4B-NVFP4

NVIDIA cut this build from Google's gemma-4-26B-A4B-it with its Model Optimizer tooling. It generates text from text, image and frame-by-frame video input over a 262,144 token context. The 4-bit row needs 8.6 GB of memory and the 16-bit row 34.5 GB, and the cheapest host for either is one 192 GB MI300X at $1.85 per hour on demand, which leaves most of the card free for that context.

Nothing in Apache 2.0 blocks a commercial deployment or a modified fork; keep the license and NOTICE file with it and state significant changes. Reconcile the parameter count first: the page counts 14.4 billion while the name carries 26B and A4B, and the base model page is where to settle it. Released May 1, 2026, this file has no per-token host price in the SAVRN Index yet, so $1.85 an hour is the cost figure to plan on.

SAVRN's Notes on Qwen2.5-14B-Instruct

Load the 16-bit weights and Qwen2.5-14B-Instruct wants 35.4 GB of memory. The cheapest slot we track is one MI300X with 192 GB at $1.85 an hour on demand, so a single card carries it with most of its memory unused. At 8-bit the need drops to 17.7 GB and at 4-bit to 8.9 GB, where this 14.8 billion parameter text generation build fits beside other work on the same card.

Apache License 2.0 lets us run it commercially, modify it and redistribute it, provided the notices travel with every copy and significant changes are stated. Two checks before committing: the configuration lists a 32,768-token context beside a 131,072-token sliding window and cites arXiv:2309.00071 on context extension, so pin down which limit your serving stack honors; and it derives from Qwen/Qwen2.5-14B, so confirm instruct tuning suits the workload. The SAVRN Index lists no per-token host price for it.

Questions

Which is larger, Gemma-4-26B-A4B-NVFP4 or Qwen2.5-14B-Instruct?

Qwen2.5-14B-Instruct (14.8B parameters) is larger than Gemma-4-26B-A4B-NVFP4 (14.4B parameters), by the parameter counts their publishers report.

Which is cheaper to run, Gemma-4-26B-A4B-NVFP4 or Qwen2.5-14B-Instruct?

At 4-bit, Gemma-4-26B-A4B-NVFP4 fits on 1x MI300X from $1.85 an hour and Qwen2.5-14B-Instruct on 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Gemma-4-26B-A4B-NVFP4 commercially?

Yes. Gemma-4-26B-A4B-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Can I use Qwen2.5-14B-Instruct commercially?

Yes. Qwen2.5-14B-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Related Comparisons