SAVRN
Search Contact SAVRN

Open-weight model · Text generation

AliceAI-Foundation-80B-A3B-BF16-GGUF

by Amaimediacom AMAImedia/AliceAI-Foundation-80B-A3B-BF16-GGUF

AliceAI-Foundation-80B-A3B-BF16-GGUF is an open-weight model for text generation from Amaimediacom, released under Apache License 2.0. It has 81.3B parameters and a 262,144-token context. At 16-bit it needs about 195.1 GB of GPU memory, which fits on 1x MI325X from $2.00 an hour; at 4-bit, 48.8 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index.

Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.

Parameters81.3B
Context262,144
Weights165.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve AliceAI-Foundation-80B-A3B-BF16-GGUF (81.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 162.6 GB 195.1 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70
8-bit 81.3 GB 97.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 40.6 GB 48.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

AliceAI-Foundation-80B-A3B-BF16-GGUF on every accelerator the SAVRN Index prices, at every precision

Model Card

By Amaimediacom, published under apache-2.0, revision a6019ebfd33c.

Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: Deterministic Hybrid Control Framework for Frozen Neural Operators — DHCF-FNO). Repack (BF16 safetensors, shards ≤50 GB) and community GGUF quantizations of (modeltype: aliceai, 80B total / 3B active MoE with KDA layers). Исходные 49 шардов Yandex слиты в один стриминговый файл и заново нарезаны SHA256 каждого тензора, 0 расхождений (см.…

Read Amaimediacom's full model card

Each donation funds the next large quant.

I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.

Boosty | Buy Me a Coffee | DonationAlerts

Thanks to Hugging Face for extra storage.


NOESIS / AMAImedia

Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: Deterministic Hybrid Control Framework for Frozen Neural Operators — DHCF-FNO).

AMAImedia


AliceAI-Foundation-80B-A3B — BF16 repack + GGUF quantizations

Repack (BF16 safetensors, shards ≤50 GB) and community GGUF quantizations of yandex/AliceAI-Foundation-80B-A3B-Base (model_type: alice_ai, 80B total / 3B active MoE with KDA layers).

Совместимость. alice_ai — гибридная архитектура (Kimi Delta Attention + gated attention + MoE + Attention Residuals). На сентябрь 2026 upstream llama.cpp не имеет графа инференса для alice_ai: все GGUF в этом репозитории собраны и проверены, но запускаются только на AliceAI-aware runtime. Для инференса используйте Transformers 5.16.1 + flash-linear-attention==0.5.0 (как в карточке оригинала) или AliceAI-aware vLLM-образ. Обычный llama-cli для этой архитектуры пока неприменим.

Compatibility. alice_ai is a hybrid architecture (Kimi Delta Attention + gated attention + MoE + Attention Residuals). As of September 2026 upstream llama.cpp has no inference graph for alice_ai: every GGUF here is converted and validated, but needs an AliceAI-aware runtime. Use Transformers 5.16.1 + flash-linear-attention==0.5.0 (upstream card) or an AliceAI-aware vLLM image. Stock llama-cli does not apply to this architecture yet.

О модели / About the model

Параметр Значение
Параметры 80B всего, 3B активных на токен
Hidden size 2048
Словарь 129024 (SentencePiece, tokenizer.model)
Слоёв 48 · 12 × (3 × (KDA → MoE) → 1 × (Gated Attention → MoE))
KDA 32 Q-головы / 32 KV-головы, head dim 128, conv kernel 4
Gated Attention 16 Q-голов / 2 KV-головы, head dim 256
MoE 512 экспертов, top-10 + 1 shared, expert FFN 512
MTP 1 слой (публикуется отдельно: mtp.safetensors)
Контекст 262 144 токена
Лицензия Apache-2.0

Файлы / Files

BF16 safetensors (repack)

File Size (bytes) Notes
models-00001-of-00004.safetensors 40 161 159 351 shard 1/4 (≤50 GB, HF limit)
models-00002-of-00004.safetensors 40 627 521 805 shard 2/4
models-00003-of-00004.safetensors 39 553 779 850 shard 3/4
models-00004-of-00004.safetensors 39 546 206 588 shard 4/4
mtp.safetensors 3 300 946 285 MTP head, 20 tensors
models.safetensors.index.json, model.safetensors.index.json — tensor → shard map (1307 tensors)
config.json, configuration_alice_ai.py, modeling_alice_ai.py, tokenizer.model, tokenizer_config.json, special_tokens_map.json — оригинальные файлы модели

Исходные 49 шардов Yandex слиты в один стриминговый файл и заново нарезаны на 4 части ≤50 ГБ; корректность подтверждена побайтово — 1307 тензоров, SHA256 каждого тензора, 0 расхождений (см. code/compare_tensors.py).

GGUF

File Notes
GGUF/AliceAI-Foundation-80B-A3B-F16.gguf F16, 1287 тензоров (MTP исключён)
GGUF/AliceAI-Foundation-80B-A3B-Q8_0.gguf Q8_0 — near-lossless reference
GGUF/AliceAI-Foundation-80B-A3B-Q6_K.gguf Q6_K — near-lossless K-quant
GGUF/AliceAI-Foundation-80B-A3B-Q4_K_M.gguf Q4_K_M — recommended general use
GGUF/AliceAI-Foundation-80B-A3B-Q2_K.gguf Q2_K — experimental low-memory tier

Как это сделано / How it was made

Весь пайплайн лежит в code/:

  1. download.py — снапшот оригинального чекпоинта (49 шардов).
  2. split_merge.py — стриминговое слияние в один файл + выделение MTP.
  3. compare_tensors.py — сверка заголовков и по-тензорный SHA256 (PASS, 0 потерь).
  4. shard4.py — нарезка на 4 валидных safetensors-шарда ≤50 ГБ + индексы.
  5. upload_shards.py — публикация BF16.
  6. patch_alice_ai.py + alice_ai.cpp + alice-ai-quant-stub.patch — quant-only стаб alice_ai для llama.cpp (регистрирует архитектуру для llama-quantize).
  7. convert_alice_foundation.py — собственный HF → GGUF F16 конвертер (маппинг тензоров, SentencePiece-словарь, метаданные; без mmap).
  8. quantize_alice.sh — 4 параллельных кванта (Q8_0, Q6_K, Q4_K_M, Q2_K).

Быстрый старт (инференс) / Quick start (inference)

# transformers==5.16.1, accelerate==1.14.0, flash-linear-attention==0.5.0
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AMAImedia/AliceAI-Foundation-80B-A3B-BF16-GGUF"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, dtype="bfloat16", device_map="auto",
)

Оригинальная документация и бенчмарки — в карточке yandex/AliceAI-Foundation-80B-A3B-Base.

Лицензия и атрибуция / License and attribution

Веса распространяются под Apache-2.0 (как у оригинала). Это community-repack и community-кванты, не официальный релиз Yandex. Инференс-код модели — modeling_alice_ai.py — включён в репозиторий без изменений.

Configuration

Architecture
AliceAIForCausalLM
Context length (tokens)
262,144
Layers
48
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
129,024
Experts
512
Experts active per token
10
RoPE base
1e+06
Model type
alice_ai

Identity and Version

Repository
AMAImedia/AliceAI-Foundation-80B-A3B-BF16-GGUF
Publisher
Amaimediacom
Task
Text generation
Modality
Text
Library
transformers
Parameters
81.3B parameters
Languages
ru, en
Revision
a6019ebfd33c963208b119a5e702175435e7e1bd
First published
2026-09-23
Last updated
2026-09-23

Files and Weights

34 files, 165.9 GB in total. The weights are 5 files totalling 165.9 GB in safetensors.

Weights5 files · 165.9 GB
Configuration17 files · 303.7 KB
Tokenizer2 files · 2.6 MB
Documentation4 files · 19.4 KB
Other5 files · 9.4 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
models-00001-of-00004.safetensorsWeights40.2 GB e0e6d7ec0f37
models-00002-of-00004.safetensorsWeights40.6 GB e3f809591c61
models-00003-of-00004.safetensorsWeights39.6 GB 8226902ea65c
models-00004-of-00004.safetensorsWeights42.2 GB 722ec735deea
mtp.safetensorsWeights3.3 GB 7910473d8449
code/check_arch.pyConfiguration1.2 KB —
code/compare_tensors.pyConfiguration4.5 KB —
code/convert_alice_foundation.pyConfiguration10.8 KB —
code/download.pyConfiguration417 B —
code/finalize_gguf.pyConfiguration3.6 KB —
code/patch_alice_ai.pyConfiguration3.8 KB —
code/rename_repo.pyConfiguration1.0 KB —
code/shard4.pyConfiguration3.8 KB —
code/split_merge.pyConfiguration4.2 KB —
code/test_infer_alice.pyConfiguration2.1 KB —
code/upload_shards.pyConfiguration1.6 KB —
config.jsonConfiguration2.2 KB —
configuration_alice_ai.pyConfiguration4.9 KB —
model.safetensors.index.jsonConfiguration112.2 KB —
modeling_alice_ai.pyConfiguration35.1 KB —
models.safetensors.index.jsonConfiguration112.2 KB —
special_tokens_map.jsonConfiguration72 B —
LICENSEDocumentation577 B —
README.mdDocumentation8.3 KB —
README_en.mdDocumentation7.2 KB —
code/README.mdDocumentation3.3 KB —
NOTICESOther480 B —
chat_template.jinjaOther3.6 KB —
code/alice-ai-quant-stub.patchOther3.5 KB —
code/alice_ai.cppOther674 B —
code/quantize_alice.shOther1.1 KB —
.gitattributesRepository1.5 KB —
tokenizer.modelTokenizer2.6 MB aed6fe7cbde4
tokenizer_config.jsonTokenizer236 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
165.9 GB
Download from Amaimediacom

Released by Amaimediacom through its official repository on Hugging Face. Read the license.

Built From

  • Derived from yandex/AliceAI-Foundation-80B-A3B-Base
  • Quantized from yandex/AliceAI-Foundation-80B-A3B-Base

Memory Requirements

PrecisionWeights in memory
As published165.9 GB
16-bit162.6 GB
8-bit81.3 GB
4-bit40.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About AliceAI-Foundation-80B-A3B-BF16-GGUF

How much GPU memory does AliceAI-Foundation-80B-A3B-BF16-GGUF need?

About 195.1 GB at 16-bit and 48.8 GB at 4-bit: the weights (81.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run AliceAI-Foundation-80B-A3B-BF16-GGUF on?

At 16-bit, 1x MI325X from $2.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use AliceAI-Foundation-80B-A3B-BF16-GGUF commercially?

Yes. AliceAI-Foundation-80B-A3B-BF16-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is AliceAI-Foundation-80B-A3B-BF16-GGUF's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

Qwen3-Coder-Next-FP8

Qwen

Today, we're announcing Qwen3-Coder-Next-FP8, an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements: Qwen3-Coder-Next-FP8 has the following features: NOTE: This model supports only non-thinking mode and does not generate blocks in its output. Meanwhile, specifying enablethinking=False is no longer required. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. We advise you to use the latest version of transformers. The following contains a code snippet illustrating how to use the model generate content based on…

Open weights apache-2.0 79.7B parameters 262,144 tokens transformers

Model · Text generation

Qwen-72B

Qwen

通义千问-72B(Qwen-72B)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B的仓库。 通义千问-72B(Qwen-72B)主要有以下特点: 1. 大规模高质量训练语料:使用超过3万亿tokens的数据进行预训练,包含高质量中、英、多语言、代码、数学等数据,涵盖通用及专业领域的训练语料。通过大量对比实验对预训练语料分布进行了优化。 2. 强大的性能:Qwen-72B在多个中英文下游评测任务上(涵盖常识推理、代码、数学、翻译等),效果显著超越现有的开源模型。具体评测结果请详见下文。 3. 覆盖更全面的词表:相比目前以中英词表为主的开源模型,Qwen-72B使用了约15万大小的词表。该词表对多语言更加友好,方便用户在不扩展词表的情况下对部分语种进行能力增强和扩展。 4. 较长的上下文支持:Qwen-72B支持32k的上下文长度。 Qwen-72B is the 72B-parameter version of the large language model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-72B is a Transformer-based large…

Open weights other 72.3B parameters 32,768 tokens transformers

Model · Text generation

Llama-3.3-70B-Instruct

Meta Llama

The Meta Llama 3.3 multilingual large language model (LLM) is an instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.3 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and…

Access requested at publisher llama3.3 70.6B parameters transformers

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-NVFP4 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 67.2B parameters 262,144 tokens transformers

Model · Text generation

Qwen3.5-122B-A10B-NVFP4

NVIDIA

The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.5-122B-A10B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems…

Open weights apache-2.0 64.6B parameters 262,144 tokens Model Optimizer

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens