SAVRN
Search Contact SAVRN

Open-weight model · Text generation

ACE-3-26B-A4B-Preview

by APMIC APMIC/ACE-3-26B-A4B-Preview

ACE-3-26B-A4B-Preview is a model for text generation from APMIC, released under Gemma Terms of Use (access requested at publisher). It has 25.8B parameters. At 16-bit it needs about 61.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

ACE-3-26B-A4B-Preview-260910 is a preview release of APMIC's ACE-3 model family, built for Traditional Chinese (Taiwan) enterprise scenarios and agentic workflows.

Parameters25.8B
Context—
Weights51.6 GB
Licensegemma
AccessAccess requested at publisher
Monthly Downloads—

Runs On

What it takes to serve ACE-3-26B-A4B-Preview (25.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 51.6 GB 61.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 25.8 GB 31.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 12.9 GB 15.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

ACE-3-26B-A4B-Preview on every accelerator the SAVRN Index prices, at every precision

Model Card

ACE-3-26B-A4B-Preview-260910 is a preview release of APMIC's ACE-3 model family, built for Traditional Chinese (Taiwan) enterprise scenarios and agentic workflows. The model is based on google/gemma-4-26B-A4B-it, a Mixture-of-Experts model with 26B total parameters and about 4B active parameters per token (the "A4B" in the name). APMIC has further optimized it to strengthen: - Traditional Chinese output in Taiwan usage (terminology, phrasing and orthography) The bundled generationconfig.json uses temperature=1.0, topp=0.95, topk=64. These follow the base model's recommended settings. The model can be served with any inference engine that supports Gemma 4 (for example vLLM). Because the…

Excerpt from the card by APMIC, licensed gemma.

Identity and Version

Repository
APMIC/ACE-3-26B-A4B-Preview
Publisher
APMIC
Task
Text generation
Modality
Text
Library
transformers
Parameters
25.8B parameters
Languages
zh, en
Revision
28ef6fbd2557ab28ebb8ee22d4a5cceb7f9f852e
First published
2026-10-07
Last updated
2026-10-07

Files and Weights

11 files, 51.7 GB in total. The weights are 2 files totalling 51.6 GB in safetensors.

Weights2 files · 51.6 GB
Configuration4 files · 108.9 KB
Tokenizer2 files · 32.2 MB
Documentation1 file · 4.2 KB
Other1 file · 18.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights49.3 GB —
model-00002-of-00002.safetensorsWeights2.3 GB —
config.jsonConfiguration3.9 KB —
generation_config.jsonConfiguration204 B —
model.safetensors.index.jsonConfiguration103.2 KB —
processor_config.jsonConfiguration1.7 KB —
README.mdDocumentation4.2 KB —
chat_template.jinjaOther18.9 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer32.2 MB —
tokenizer_config.jsonTokenizer3.7 KB —

License and Download

License
gemma
Access
Access requested at publisher
Download size
51.6 GB
Request access from APMIC

APMIC grants access through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published51.6 GB
16-bit51.6 GB
8-bit25.8 GB
4-bit12.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ACE-3-26B-A4B-Preview

How much GPU memory does ACE-3-26B-A4B-Preview need?

About 61.9 GB at 16-bit and 15.5 GB at 4-bit: the weights (25.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run ACE-3-26B-A4B-Preview on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use ACE-3-26B-A4B-Preview commercially?

Yes, with conditions. ACE-3-26B-A4B-Preview is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.

Similar Models

Model · Text generation

WaifuGemma4-26b-a4b-v1

HiWaifu Research

Gemma 4 26B-A4B, post-trained with GRPO against a reward model learned from 1.2 million double-blind votes cast by HiWaifu users inside their own role-play conversations. Put back into the same arena, blind, it met GLM-5.1 in 1,430 battles and won 49.6% of the decided votes; against a 13-model field including Gemini, DeepSeek-v4 and Qwen's character models it won 54.7%. Most open role-play models are tuned on preferences that come from an LLM judge, from a handful of annotators, or from synthetic pairs. We had something rarer: a live arena where, inside ordinary chats on our platform, a user is occasionally shown two candidate replies and asked which one they want to continue with. Those…

Open weights gemma 25.8B parameters 262,144 tokens transformers

This checkpoint is an AutoRound model-free MXFP8 RTN quantization of exported in llmcompressor / compressed-tensors format. Routed experts and the self-attention projections present in the source are stored as F8E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV scales were calibrated with AutoRound using the text dataset NeelNanda/pile-10k; the vision tower was not quantized. On the paired repository lmeval protocol, the four primary metrics were non-decreasing relative to the BF16 baseline using the same vLLM FP8 KV cache. The AQA gate was GO. This is a result for those tasks and settings only, not a claim of lossless quantization or general quality improvement. The…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

This repository contains an MXFP8-weight checkpoint derived from exported in compressed-tensors format. The checkpoint retains the source multimodal components, but the evaluation reported here covers text tasks only. The exported checkpoint contains 11,635 F8E4M3 weight tensors, 838 BF16 weight tensors, 11,635 U8 block-scale tensors, and 171 FP32 scale/metadata tensors. Routed-expert weights and source-present text self-attention projections are MXFP8; the router, shared MLP, vision tower, and embeddings remain BF16. Measured on 2026-09-30 with lm-eval 0.4.13 and vLLM 0.29.0. The tested settings were TRITONATTN, tensor parallelism 2, pipeline parallelism 1, batch size 64, maxnumseqs=64…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

Model · Text generation

Darwin-27B-RSI

FINAL_Bench

Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution). Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests. As the reasoning engine of Darwin-27B-JEV on…

Open weights apache-2.0 26.9B parameters 262,144 tokens transformers

Prism ML's ternary Ternary-Bonsai-2-27B build of Qwen/Qwen3.8-27B, repacked for chad, a Claude-Code-style local coding agent for Apple Silicon, with its speculative decoder bundled in. This is chad's default model. Created using Bonsai by Prism ML. with, already quantized. Nothing is built on first run. Every projection of Qwen3.8-27B (a dense qwen35 hybrid: 64 layers, 48 GatedDeltaNet + 16 full attention) is stored in a Hadamard-rotated basis: multiplied by a fixed sign vector and put through a blockwise Walsh-Hadamard transform offline, then quantized to 2-bit affine group-128 whose three levels reproduce the ternary set {−s, 0, +s}. The rotation costs no extra bits and no extra weight…

Open weights apache-2.0 26.9B parameters 262,144 tokens mlx

Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via JiRackDeltaNetTokenizer - JiRack DeltaNet understand video and images that best for Robotics also A fast and efficient 27B model optimized for CPU inference. Built on a Qwen3.8-style DeltaNet architecture (hybrid attention + SSM), with an updated tokenizer that includes Routing, Media, Vision, Sound, Tool call, and Robotics tags. Ready-to-run GGUF quantizations, and native Ollama support with reasoning disabled by default for fast, direct responses. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert…

Open weights mit 27.3B parameters 262,144 tokens