SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

ACE-3-F-26B-A4B-260804

by APMIC APMIC/ACE-3-F-26B-A4B-260804

ACE-3-F-26B-A4B-260804 is a model for image and text to text from APMIC, released under Gemma Terms of Use (access requested at publisher). It has 25.8B parameters. At 16-bit it needs about 61.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Parameters25.8B
Context—
Weights—
Licensegemma
AccessAccess requested at publisher
Monthly Downloads—

Runs On

What it takes to serve ACE-3-F-26B-A4B-260804 (25.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 51.6 GB 61.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 25.8 GB 31.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 12.9 GB 15.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

ACE-3-F-26B-A4B-260804 on every accelerator the SAVRN Index prices, at every precision

Model Card

Identity and Version

Repository
APMIC/ACE-3-F-26B-A4B-260804
Publisher
APMIC
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
25.8B parameters
Languages
zh, en
Revision
d9172ae23626953f05d9387da9e25b5cd7fda824
First published
2026-08-20
Last updated
2026-10-07

License and Download

License
gemma
Access
Access requested at publisher
Request access from APMIC

APMIC grants access through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
16-bit51.6 GB
8-bit25.8 GB
4-bit12.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ACE-3-F-26B-A4B-260804

How much GPU memory does ACE-3-F-26B-A4B-260804 need?

About 61.9 GB at 16-bit and 15.5 GB at 4-bit: the weights (25.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run ACE-3-F-26B-A4B-260804 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use ACE-3-F-26B-A4B-260804 commercially?

Yes, with conditions. ACE-3-F-26B-A4B-260804 is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.

Similar Models

Model · Image and text to text

gemma-4-26B-A4B-it

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

Model · Image and text to text

gemma-4-26B-A4B-it-AWQ-4bit

Cyankiwi

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in four distinct sizes: E2B, E4B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from high-end phones…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

Model · Image and text to text

Aura-Prototype-26B-A4B

EldritchLabs

This is a merge of pre-trained language models created using mergekit. This model was merged using the aura merge method. Aura is an experimental method with a live heatmap visualizer. This model took 10 hours to merge using graphv18.py The following models were included in the merge: - TheDrummer/Orion-26B-A4B-v1.1 - Gryphe/Pantheon-Reasoning-26B-A4B-1.1-V2 - electroglyph/gemma4-26b-fiction-bf16 The following YAML configuration was used to produce this model

Open weights apache-2.0 26B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-mlx-6Bit

Vy Thông Nguyễn

The Model nvythong/Qwen3.8-27B-mlx-6Bit was converted to MLX format from Qwen/Qwen3.8-27B using mlx-lm version 0.31.2.

Open weights apache-2.0 26.9B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-35B-A3B-NVFP4

Unsloth AI

1.56x faster throughput than other NVFP4 quants. This is an Unsloth NVFP4 quantized checkpoint calibrated on a mixture of our Unsloth dataset + UltraChat dataset. Works on a 32GB VRAM GPU. Benchmarks on 1xB200 128 concurrency. Use the 35B NVFP4 Fast version for 1.79x faster at a little less accuracy For accuracy benchmarks, we conducted MMLU-Pro, AIME 2025, GPQA for FP8, BF16, NVIDIA's NVFP4 and our NVFP4s - we show our faster quants do similarly on all: Read all benchmarks in our NVFP4 blog To install vLLM in a separate venv: Then to serve the 35B variant: You must use the below or you will get 2x slower inference! Also do NOT use the Marlin backend since it's 2x slower - use the native…

Open weights apache-2.0 24.6B parameters 262,144 tokens transformers