SAVRN
Search Contact SAVRN

Open-weight model · Text generation

gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound

by INC Optimized Models 4 INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound

gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound is an open-weight model for text generation from INC Optimized Models 4, released under Apache License 2.0. It has 25.8B parameters and a 262,144-token context. At 16-bit it needs about 61.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

This checkpoint is an AutoRound model-free MXFP8 RTN quantization of exported in llmcompressor / compressed-tensors format.

Parameters25.8B
Context262,144
Weights28.4 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound (25.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 51.6 GB 61.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 25.8 GB 31.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 12.9 GB 15.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound on every accelerator the SAVRN Index prices, at every precision

Model Card

By INC Optimized Models 4, published under apache-2.0, revision dc8c7c2ba6d6.

This checkpoint is an AutoRound model-free MXFP8 RTN quantization of exported in llmcompressor / compressed-tensors format. Routed experts and the self-attention projections present in the source are stored as F8E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV scales were calibrated with AutoRound using the text dataset NeelNanda/pile-10k; the vision tower was not quantized. On the paired repository lmeval protocol, the four primary metrics were non-decreasing relative to the BF16 baseline using the same vLLM FP8 KV cache. The AQA gate was GO. This is a result for those tasks and settings only, not a claim of lossless quantization or general quality improvement. The…

Read INC Optimized Models 4's full model card

This checkpoint is an AutoRound model-free MXFP8 RTN quantization of google/gemma-4-26B-A4B, exported in llm_compressor / compressed-tensors format. Routed experts and the self-attention projections present in the source are stored as F8_E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV scales were calibrated with AutoRound using the text dataset NeelNanda/pile-10k; the vision tower was not quantized.

On the paired repository lm_eval protocol, the four primary metrics were non-decreasing relative to the BF16 baseline using the same vLLM FP8 KV cache. The AQA gate was GO. This is a result for those tasks and settings only, not a claim of lossless quantization or general quality improvement.

Read this first: max_model_len=131072 was configured, but the benchmark prompts were not a 128K long-context stress test. This run evaluated text tasks only; vision quality, multimodal prompts, thinking-on, throughput and latency were not measured. The tokenizer configuration used by the run had no chat_template, so GSM8K used the repository's bare-prompt path.

1. Model summary

Base model This checkpoint
Name / architecture / parameters Gemma 4 26B-A4B; Gemma4ForConditionalGeneration; nominal count 26,544,131,376 (26.544B; tied-weight counting convention) gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound; 25,805,936,206 unique serialized weight elements
Layer structure 30 text layers; 128 routed experts/layer, top-8; 25 sliding-attention + 5 full-attention layers; vision tower Routed-expert and source-present text self-attention projection weights MXFP8; protected modules BF16
Weight tensors 1,013 source index entries, BF16 24,222 exported tensor entries: 11,635 F8_E4M3 weights, 11,635 U8 block scales, 838 BF16 weights and 114 F32 static KV scales
Size on disk 51,611,872,412 bytes, source safetensors index total_size 28,412,087,908 bytes, exported safetensors index total_size
Compression / effective bits per parameter BF16 source 1.8165× smaller; 8.8079 effective bits per unique serialized weight element (8 × export bytes / 25,805,936,206)
Quantization / format / context / license Gemma 4; Apache 2.0 MXFP8 weights + static FP8 KV; compressed-tensors; evaluated with a 131,072-token maximum length; Apache 2.0 terms apply

The model configuration sets tie_word_embeddings: true. The nominal 26.544B model count and the index-derived unique serialized element count are different counting bases; the size and dtype accounting below uses the latter. The A4B designation is the upstream model's active-parameter naming, not a measurement made by this quantization run.

Serving was validated through the vLLM 0.29.0 backend on three RTX 5090 GPUs with TP=1 / PP=3. That is the tested topology, not a claim that other hardware or parallel layouts are equivalent.

2. Precision plan

Percentages below are relative to the 25,805,936,206 unique stored model-weight elements. Scale tensors are metadata and are not included in this denominator.

Bit width / stored dtype Module (layers) Weight tensors Weight elements Share
8-bit float (F8_E4M3) Routed expert gate/up/down weights (30 × 128 experts) 11,520 22,837,985,280 88.4990%
8-bit float (F8_E4M3) Existing text self-attention projections (q/k/o in 30 layers; v in 25) 115 1,110,179,840 4.3020%
16-bit (BF16) Shared MLP branch 90 535,265,280 2.0742%
16-bit (BF16) Router 90 10,901,760 0.0422%
16-bit (BF16) Vision tower 355 569,550,384 2.2071%
16-bit (BF16) Vision embedding projection 1 3,244,032 0.0126%
16-bit (BF16) Token embedding 1 738,197,504 2.8606%
16-bit (BF16) Other retained tensors, including norms and patch/position tensors 301 612,126 0.0024%
Total model weights F8_E4M3 + BF16 12,473 25,805,936,206 100.0000%

Shares are rounded independently; the unrounded category shares sum to 100%. The exported weight configuration is Linear, 8-bit float, symmetric, group-wise, group size 32, static weights, with U8 scales and memoryless_minmax observer. The config also records dynamic 8-bit float input-activation quantization metadata. Static KV is 8-bit float, tensor-wise and dynamic=false.

The exported checkpoint also has 11,635 U8 block-scale tensors and 114 F32 static KV scales. Of the latter, the 60 text K/V scales (30 layers × K/V) are finite and nonzero; the 54 vision scales are zero because calibration and evaluation were text-only.

Ignore groups

Group Reason / evidence
router Preserve routing decisions; retained as BF16 in the header audit.
mlp This is the always-on shared branch added to routed-expert output; retained as BF16.
vision_tower, embed_vision Preserve the multimodal tower and projection; not part of this text-only evaluation.
embed_tokens Embedding is not a supported AutoRound Linear target; source/export key is present and BF16. The logged unmatched-module warning is benign, not a missing checkpoint key.

Norms and other non-target tensors remain BF16 by the model-free defaults. The checkpoint contains no multi_modal_projector key, so it is not included in the ignore list.

2.1 Quantization plan and artifact deviations

Audited plan Export observation Difference / consequence
Quantize all routed experts 11,520 F8_E4M3 tensors across 30 layers × 128 experts Matches plan; 60 fused source expert keys were expanded into per-expert tensors.
Quantize source-present text self-attention projections 115 F8_E4M3 projection tensors Matches source index: q/k/o exist in all 30 layers, while v exists only in 25 sliding-attention layers. The five full-attention source layers have no v_proj.
Keep shared MLP, router, vision, embeddings and other non-target tensors BF16 838 BF16 tensors, 1,857,771,086 elements Matches plan; no non-fused source keys were absent from the exported index.
Add static FP8 KV scales 114 F32 scales; all finite; 60 text scales nonzero Text scales cover the evaluated workload. Vision scales are zero after text-only calibration and were not exercised because language_model_only=true.

There was no unplanned weight-precision change. The embed_tokens warning reflects that embeddings are not Linear quantization targets; the key remains present in BF16.

2.2 Size accounting

Artifact Header/index bytes Comparison
BF16 source checkpoint 51,611,872,412 Source index total_size; 1,013 BF16 entries
MXFP8 + BF16 export 28,412,087,908 Export index total_size; 24,222 weight/scale entries

The ratio is 51,611,872,412 / 28,412,087,908 = 1.8165×. Effective bits per unique serialized model-weight element are 28,412,087,908 × 8 / 25,805,936,206 = 8.8079; this includes block/KV scale storage represented in the exported index and is not the nominal MXFP8 weight bit width.

3. Evaluation results

Measured 2026-09-28 with lm_eval 0.4.13 using its vLLM backend (vLLM 0.29.0), TP=1 / PP=3, seed 42, BF16 compute dtype and FP8 KV cache for both baseline and candidate. Values are means ± lm_eval standard error (SE); N is the evaluated sample count. The final column is the score difference divided by sqrt(SE_baseline² + SE_MXFP8²); it is an approximate independent-SE scale, not a paired significance test.

Task / primary metric N BF16 baseline MXFP8 + FP8 KV Δ (pp) Δ / combined SE
PIQA (acc) 1,838 0.8226 ± 0.0089 0.8286 ± 0.0088 +0.60 +0.48
MMLU (acc) 14,042 0.7434 ± 0.0034 0.7441 ± 0.0034 +0.07 +0.15
HellaSwag (acc) 10,042 0.6325 ± 0.0048 0.6347 ± 0.0048 +0.22 +0.32
GSM8K (exact_match,strict-match) 1,319 0.7180 ± 0.0124 0.7324 ± 0.0122 +1.44 +0.83

The arithmetic mean of the four primary scores is 0.7291 (baseline) vs 0.7350 (MXFP8), a +0.58 pp difference; this mean is descriptive and has no combined SE. PIQA acc_norm was 0.8395 ± 0.0086 vs 0.8341 ± 0.0087 (−0.54 pp); HellaSwag acc_norm was 0.8305 ± 0.0037 vs 0.8317 ± 0.0037 (+0.12 pp). The AQA gate uses the primary metrics above and reports GO (max_accuracy_drop=0.02). The benchmark differences are within one approximate combined SE; this run does not establish statistical superiority.

3.2 What was held fixed

Both measured runs used max_model_len=131072, kv_cache_dtype=fp8, tensor_parallel_size=1, pipeline_parallel_size=3, gpu_memory_utilization=0.85, dtype=bfloat16, attention_backend=FLASHINFER, language_model_only=true, max_num_seqs=1, trust_remote_code=true, add_bos_token=true, prefix caching disabled, thinking disabled and seed 42. PIQA/MMLU/HellaSwag used the repository task group with batch_size=auto; GSM8K used 5-shot and batch_size=64, with a 2,048-token generation budget. Inference-engine internal scheduling details not exposed by lm_eval are not claimed as controlled measurements.

3.3 Versus baseline

All four AQA primary scores are non-decreasing in this paired run. The differences are small relative to their reported task-level uncertainty, so the result supports the stated gate decision but does not prove that MXFP8 improves accuracy. An earlier failed run had a different repeated BF16 GSM8K strict-match score (0.7286 vs 0.7180); its cause is unknown and it is not mixed into this gate comparison.

3.4 Functional evidence and provenance

Direct evidence: the exported checkpoint loaded in vLLM 0.29.0 with FP8 KV cache and completed all four listed lm_eval tasks. Its quantization_config.json records compressed-tensors MXFP8 weights and a static 8-bit float KV scheme. The validation run ID is gemma-4-26B-A4B_MXFP8_20260927-220842_a0c394.

Not measured: full 128K prompts, multimodal/vision quality, thinking-on, agentic tasks, throughput or latency. Environment support probes and sibling model runs are not substituted for this checkpoint's measured evidence.

4. Usage

4.1 vLLM through the validated lm_eval path

The run validated vLLM loading/evaluation via lm_eval. After the repository is available, the following uses its Hub ID in place of the original local checkpoint path:

MODEL_ID=INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
MODEL_ARGS="pretrained=${MODEL_ID},tensor_parallel_size=1,pipeline_parallel_size=3,max_model_len=131072,gpu_memory_utilization=0.85,dtype=bfloat16,trust_remote_code=True,add_bos_token=True,enable_prefix_caching=False,max_gen_toks=2048,attention_backend=FLASHINFER,language_model_only=True,kv_cache_dtype=fp8,max_num_seqs=1,enable_thinking=False"

# Select three idle GPUs before running; this example assumes physical GPUs 1-3 are idle.
export CUDA_VISIBLE_DEVICES=1,2,3
lm_eval --model vllm --model_args "$MODEL_ARGS" \
  --tasks piqa,mmlu,hellaswag --batch_size auto --log_samples --seed 42
lm_eval --model vllm --model_args "$MODEL_ARGS" \
  --tasks gsm8k --batch_size 64 --log_samples --seed 42 --num_fewshot 5

Hard constraints and evidence:

Setting Requirement / evidence
KV cache Keep kv_cache_dtype=fp8 to match the evaluated quantized model protocol; confirmed in both baseline/candidate lm_eval metadata and the export KV scheme.
Context max_model_len=131072 is the measured configuration. This is not evidence of 128K prompt throughput or quality.
Parallel layout TP=1 / PP=3 on three GPUs is the tested layout. PP=2 failed CUDA graph warmup with an additional 636 MiB allocation in logs/aqa-run-retry-01.log; other layouts are unverified.
Backend vLLM 0.29.0 with attention_backend=FLASHINFER was exercised by the four evaluation tasks. For head dimension 512, the XQA decode path logged a fallback to native FlashInfer decode; KV remained FP8 and evaluation completed.
Prompt formatting The run's local/upstream tokenizer configuration had no chat_template. GSM8K therefore used apply_chat_template=false and fewshot_as_multiturn=false; multimodal/chat formatting was not evaluated.

For text inference, supply a plain prompt compatible with the task rather than assuming a chat template. A direct production vllm serve deployment was not separately benchmarked by this run.

4.2 Transformers inspection path

The following reads model configuration only; it is not a weight-loading or serving validation path:

from transformers import AutoConfig

model_id = "INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound"
config = AutoConfig.from_pretrained(model_id, trust_remote_code=True)
print(config.model_type)

The checkpoint contains about 28.41 GB of indexed tensors. Loading full weights through an inspection-oriented Transformers path can use substantially more host/GPU memory; use the tested vLLM backend for inference.

5. Reproducibility

5.1 Hardware / OS

  • GPUs: 3 × NVIDIA GeForce RTX 5090 used (compute capability 12.0; physical devices 1, 2 and 3); 32,607 MiB total memory reported per device.
  • Host probe on 2026-09-28: NVIDIA driver 610.57.04; Intel Xeon 6767P (256 logical CPUs); 188.4 GiB RAM.
  • OS: Ubuntu 24.04.4 LTS, kernel 6.8.0-139-generic, x86_64.
  • CUDA: PyTorch build CUDA 13.0; runtime 13.0.88.
  • Storage: model and run were on local filesystem paths; filesystem type was not recorded.

5.2 Environment

Package Version
Python 3.12.14
auto-round 0.16.0.dev193+g22d6e28c
transformers 5.17.0
torch 2.13.0+cu130
vllm 0.29.0
flashinfer-python 0.7.0
compressed-tensors 0.17.0
lm-eval 0.4.13
autoquant-agent 0.1.0

Versions are from the run manifest / frozen Conda environment, not a statement of minimum compatible versions.

5.3 Version selection

No packages were installed or upgraded for this run. These exact versions were used for the successful quantization and evaluation. Compatibility with other AutoRound, Transformers, vLLM or FlashInfer versions was not established here.

5.4 Artifact-level dependency warning

AutoRound 0.16.0.dev193 with Transformers 5.17.0 initially failed when its fused-MoE meta skeleton could not materialize the new self_attn.k_scale and self_attn.v_scale parameters. The successful run set AR_DISABLE_META_LOAD=1, which disabled that meta-load path and allowed CPU loading/initialization. Peak host RAM was 64.14 GiB. No source-code patch was applied.

5.5 Environment variables

For quantization, set the tested workaround and map three verified idle physical devices:

export AR_DISABLE_META_LOAD=1
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>

AR_DISABLE_META_LOAD is a quantization workaround; it is not required for inference. Do not put Hub tokens in this README or in a run YAML.

6. Reproduce the artifact

The original run was orchestrated with AQA; do not bypass it with a manually assembled AutoRound or lm_eval pipeline. Use fresh local paths and select three idle physical GPUs before starting.

6.1 Quantization YAML

The following is the core of the validated run.yaml; adapt only model, environment and output paths:

model: /path/to/gemma-4-26B-A4B

quant:
  scheme: MXFP8
  method: model_free
  export_format: llm_compressor
  ignore_layers: "router,vision_tower,embed_vision,embed_tokens,mlp"
  static_kv_fp8: true
  device_map: "0,1,2"

eval:
  backend: vllm
  harness: lm_eval
  benchmarks: [piqa, mmlu, hellaswag, gsm8k]
  baseline: true
  thinking: "off"
  gen_max_toks: 2048
  lm_eval:
    group: true
    seed: 42
    apply_chat_template: false
    fewshot_as_multiturn: false
    model_args:
      attention_backend: FLASHINFER
      language_model_only: true
      kv_cache_dtype: fp8
      max_model_len: 131072
      max_num_seqs: 1
      gpu_memory_utilization: 0.85
      tensor_parallel_size: 1
      pipeline_parallel_size: 3
      enable_thinking: false
      enable_prefix_caching: false

gate:
  max_accuracy_drop: 0.02

model_free uses the model-free RTN weight path, but static_kv_fp8 still loads the checkpoint and performs calibration forwards for KV scales. The successful log used the standard text calibration dataset NeelNanda/pile-10k; the vision tower was not quantized.

6.2 Exact quantization arguments

These are the exact AutoRound flags recorded from the successful AQA builder run; only output_dir is parameterized here. Reproduce the run through AQA as shown below:

auto-round --model_name /models/gemma-4-26B-A4B --scheme MXFP8 --model_free \
  --static_kv_dtype fp8 \
  --ignore_layers router,vision_tower,embed_vision,embed_tokens,mlp \
  --device_map 0,1,2 --format llm_compressor \
  --output_dir <fresh-run-dir>/quantized

6.3 Validate and run

export AR_DISABLE_META_LOAD=1
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>
aqa validate run.yaml
aqa run run.yaml --dry-run
aqa run run.yaml

The tested run completed the full AQA baseline → quantization → evaluation pipeline. Its output had six safetensors shards and an index size of 28,412,087,908 bytes.

6.4 Evaluate

The measured commands used the same vLLM model arguments shown in §4.1. The GSM8K task used five few-shot examples, seed 42 and a 2,048-token generation budget, with chat-template and multiturn formatting disabled because of the tokenizer configuration observed in this run.

7. Known issues and caveats

  • Do not infer long-context behavior from max_model_len=131072; no 128K prompt was measured.
  • Do not infer vision/multimodal quality from the text-only results. The vision weights remain BF16, but their quality was not evaluated.
  • PP=2 at the same length hit a CUDA graph warmup OOM on this host; the tested topology is TP=1 / PP=3 on three RTX 5090 GPUs.
  • FlashInfer XQA did not support this model's 512 head dimension in the tested path; vLLM fell back to native FlashInfer decode. FP8 KV evaluation still completed.
  • The token embedding warning is benign: it is not an AutoRound Linear target, its key is present, and it remains BF16.
  • A previous BF16 GSM8K baseline repeat scored 0.7286 vs 0.7180 in the successful run. The cause is unknown; the reported gate compares baseline and MXFP8 only within the successful run.
  • PIQA acc_norm decreased by 0.54 pp even though the AQA primary acc increased. Consult both metric and uncertainty tables; do not call the quantization lossless.

8. License and attribution

The base model metadata identifies Apache 2.0 and links to the Gemma 4 license information. Users must comply with the base-model license and preserve its attribution. This quantized checkpoint is not an official Google release.

Quantization: Intel AutoRound. Inference validation: vLLM and FlashInfer. Evaluation: lm-evaluation-harness. Text calibration used NeelNanda/pile-10k; calibration samples are not included in this repository. This repository contains the quantized weights, tokenizer/processor configuration and model configuration; it does not redistribute benchmark datasets.

Configuration

Architecture
Gemma4ForConditionalGeneration
Context length (tokens)
262,144
Layers
30
Hidden size
2,816
Feed-forward size
2,112
Attention heads
16
Key/value heads
8
Head dimension
256
Vocabulary size
262,144
Experts
128
Sliding window (tokens)
1,024
Model type
gemma4
Quantization
compressed-tensors

Identity and Version

Repository
INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
Publisher
INC Optimized Models 4
Task
Text generation
Modality
Text
Library
transformers
Parameters
25.8B parameters
Languages
Not stated by the source
Revision
dc8c7c2ba6d6d0fb803f8966052e594ae45f52b7
First published
2026-09-28
Last updated
2026-09-28

Files and Weights

15 files, 28.5 GB in total. The weights are 6 files totalling 28.4 GB in safetensors.

Weights6 files · 28.4 GB
Configuration5 files · 2.6 MB
Tokenizer2 files · 32.2 MB
Documentation1 file · 19.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00006.safetensorsWeights5.4 GB 88c867b5b419
model-00002-of-00006.safetensorsWeights5.4 GB bb608c31a522
model-00003-of-00006.safetensorsWeights5.4 GB a0b9b9e9fb8c
model-00004-of-00006.safetensorsWeights5.4 GB 83ab3c5a52d8
model-00005-of-00006.safetensorsWeights5.4 GB 96c5c01a5726
model-00006-of-00006.safetensorsWeights1.6 GB 4fcc361091d7
config.jsonConfiguration25.1 KB —
generation_config.jsonConfiguration177 B —
model.safetensors.index.jsonConfiguration2.5 MB —
processor_config.jsonConfiguration1.7 KB —
quantization_config.jsonConfiguration20.2 KB —
README.mdDocumentation19.9 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer32.2 MB 12bac982b793
tokenizer_config.jsonTokenizer1.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
28.4 GB
Download from INC Optimized Models 4

Released by INC Optimized Models 4 through its official repository on Hugging Face. Read the license.

Built From

  • Derived from google/gemma-4-26B-A4B
  • Quantized from google/gemma-4-26B-A4B

Memory Requirements

PrecisionWeights in memory
As published28.4 GB
16-bit51.6 GB
8-bit25.8 GB
4-bit12.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound

How much GPU memory does gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound need?

About 61.9 GB at 16-bit and 15.5 GB at 4-bit: the weights (25.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound commercially?

Yes. gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This repository contains an MXFP8-weight checkpoint derived from exported in compressed-tensors format. The checkpoint retains the source multimodal components, but the evaluation reported here covers text tasks only. The exported checkpoint contains 11,635 F8E4M3 weight tensors, 838 BF16 weight tensors, 11,635 U8 block-scale tensors, and 171 FP32 scale/metadata tensors. Routed-expert weights and source-present text self-attention projections are MXFP8; the router, shared MLP, vision tower, and embeddings remain BF16. Measured on 2026-09-30 with lm-eval 0.4.13 and vLLM 0.29.0. The tested settings were TRITONATTN, tensor parallelism 2, pipeline parallelism 1, batch size 64, maxnumseqs=64…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

Model · Text generation

WaifuGemma4-26b-a4b-v1

HiWaifu Research

Gemma 4 26B-A4B, post-trained with GRPO against a reward model learned from 1.2 million double-blind votes cast by HiWaifu users inside their own role-play conversations. Put back into the same arena, blind, it met GLM-5.1 in 1,430 battles and won 49.6% of the decided votes; against a 13-model field including Gemini, DeepSeek-v4 and Qwen's character models it won 54.7%. Most open role-play models are tuned on preferences that come from an LLM judge, from a handful of annotators, or from synthetic pairs. We had something rarer: a live arena where, inside ordinary chats on our platform, a user is occasionally shown two candidate replies and asked which one they want to continue with. Those…

Open weights gemma 25.8B parameters 262,144 tokens transformers

Model · Text generation

Darwin-27B-RSI

FINAL_Bench

Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution). Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests. As the reasoning engine of Darwin-27B-JEV on…

Open weights apache-2.0 26.9B parameters 262,144 tokens transformers

Prism ML's ternary Ternary-Bonsai-2-27B build of Qwen/Qwen3.8-27B, repacked for chad, a Claude-Code-style local coding agent for Apple Silicon, with its speculative decoder bundled in. This is chad's default model. Created using Bonsai by Prism ML. with, already quantized. Nothing is built on first run. Every projection of Qwen3.8-27B (a dense qwen35 hybrid: 64 layers, 48 GatedDeltaNet + 16 full attention) is stored in a Hadamard-rotated basis: multiplied by a fixed sign vector and put through a blockwise Walsh-Hadamard transform offline, then quantized to 2-bit affine group-128 whose three levels reproduce the ternary set {−s, 0, +s}. The rotation costs no extra bits and no extra weight…

Open weights apache-2.0 26.9B parameters 262,144 tokens mlx

Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via JiRackDeltaNetTokenizer - JiRack DeltaNet understand video and images that best for Robotics also A fast and efficient 27B model optimized for CPU inference. Built on a Qwen3.8-style DeltaNet architecture (hybrid attention + SSM), with an updated tokenizer that includes Routing, Media, Vision, Sound, Tool call, and Robotics tags. Ready-to-run GGUF quantizations, and native Ollama support with reasoning disabled by default for fast, direct responses. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert…

Open weights mit 27.3B parameters 262,144 tokens

Model · Text generation

Ternary-Bonsai-27B-mlx-2bit

Prism ML

Full 27B-class reasoning in ternary transformer weights — on everyday laptops - \~7.2 GB deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 95% of FP16 intelligence retained: 80.49 average across 15 thinking-mode benchmarks — a higher score than the conventional IQ2XXS build (72.73) at less than two-thirds of its footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 - End-to-end ternary language weights across embeddings, attention projections, MLP…

Open weights apache-2.0 27.4B parameters 262,144 tokens mlx