SAVRN
Search Contact SAVRN

Open-weight model · Text generation

CAT-YOKO

by Avrova Donz AvrovaDonz/CAT-YOKO

CAT-YOKO is an open-weight model for text generation from Avrova Donz, released under Apache License 2.0. Its published files total 1.8 GB.

YOCO 式因果 encoder-decoder MoE。从 MiniCPM5-2B 上采样:openbmb/MiniCPM5-2B-Base(Apache-2.0,Llama GQA)。 libraryname: transformers 只表示 tokenizer / 分片约定走 HuggingFace 生态。当前图是仓库里的 catyoko PyTorch 实现,不是 Hub 上可 AutoModelForCausalLM 直接加载的架构。 50B token 信封上的理论墙钟,不是实测。B200 /…

Parameters—
Context—
Weights1.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By Avrova Donz, published under apache-2.0, revision 10e9d939aa67.

YOCO 式因果 encoder-decoder MoE。从 MiniCPM5-2B 上采样:openbmb/MiniCPM5-2B-Base(Apache-2.0,Llama GQA)。 libraryname: transformers 只表示 tokenizer / 分片约定走 HuggingFace 生态。当前图是仓库里的 catyoko PyTorch 实现,不是 Hub 上可 AutoModelForCausalLM 直接加载的架构。 50B token 信封上的理论墙钟,不是实测。B200 / SM100 允许的线性 GEMM 走 TeNvfp4Linear(TE NVFP4BlockScaling;B0 冻 encoder 走 FPROP,WGRAD 留给 B1/B2)。无 TE / sm120 时 Nvfp4Linear E2M1/16 仿真。attn softmax / SDPA 仍 fp32。 发布信封 进行中(DummyStream,尚未跑完 8e9)。Vast B200 已回收(2026-09-19)。本文件是释放前快照,不是终局。GitHub 口径:docs/STATUS.md。 代码与指针:GitHub AvrovaDonz2026/CAT-YOKO(不用 LFS)。 这些 overlay 里,checkpoints/b0/ 与 checkpoints/b0-nvfp4-try/ 不是 8B token 信封(只是 --try)。checkpoints/b0-full/trainable.pt 是发布信封 进行中 的 B0…

Read Avrova Donz's full model card

CAT-YOKO-12B

YOCO 式因果 encoder-decoder MoE。从 MiniCPM5-2B 上采样:openbmb/MiniCPM5-2B-Base(Apache-2.0,Llama GQA)。

本卡对应 AvrovaDonz/CAT-YOKO。训练代码:AvrovaDonz2026/CAT-YOKO。

library_name: transformers 只表示 tokenizer / 分片约定走 HuggingFace 生态。当前图是仓库里的 cat_yoko PyTorch 实现,不是 Hub 上可 AutoModelForCausalLM 直接加载的架构。

规格

项 值
名称 CAT-YOKO-12B
架构 YOCO 式因果 encoder-decoder MoE
底座 MiniCPM5-2B(Llama GQA,untied)
(d) 2048
(V) 130560
(L) 42(16 encoder + 26 decoder)
注意力 16 Q / 2 KV,head_dim=128
FFN SwiGLU 中间维 6144
MoE 1 shared + 20 routed
top-(k) encoder 7 / decoder 10
存储参数 12,250,381,312(≈12.25B)
Encoder 活跃 ≈2.03B / 输入 token
Decoder 活跃 ≈4.33B / 输出 token

课程 C1

子阶段 token Encoder 可训练
B0 8B 冻结 仅新模块
B1 27B 冻结 decoder + lm_head + 最终 RMSNorm
B2 15B 可训练 全部

墙钟(发布账本)

50B token 信封上的理论墙钟,不是实测。B200 / SM100 允许的线性 GEMM 走 TeNvfp4Linear(TE NVFP4BlockScaling;B0 冻 encoder 走 FPROP,WGRAD 留给 B1/B2)。无 TE / sm_120 时 Nvfp4Linear E2M1/16 仿真。attn softmax / SDPA 仍 fp32。

配方 H100-h 角色
C1+NVFP4 571 发布墙钟。B0 student bf16;B1/B2 允许的 GEMM 走 NVFP4
C1+FP8 729 Hopper / Ada 回退
联合 bf16 1325 100% 对照基线

数据与分词

项 值
语料 Ultra-FineWeb en 55% / zh 30% + UltraData-Math 10% + StarCoder 5%(思考 mix;不拉 50B)
Tokenizer openbmb/MiniCPM5-2B

当前 B0 快照

发布信封 进行中(DummyStream,尚未跑完 8e9)。Vast B200 已回收(2026-09-19)。本文件是释放前快照,不是终局。GitHub 口径:docs/STATUS.md。

项 值
文件 checkpoints/b0-full/trainable.pt
文件夹说明 checkpoints/b0-full/README.md
step 26940
tokens_in_phase 130,041,856(信封 8e9 的 ≈1.63%)
sha256 7eebc9a4da78d79be71bbe52881f2a0eaffd899f58ada3a3325f410eca181955
张量 132,无 Adam
机器 Vast NVIDIA B200 SM 10.0(已回收)
运行时 torch 2.11+cu128 + TE nvcc 12.9 SM100
吞吐 / 显存 micro-batch=2,~15.7k tok/s,Trainer ~138GiB
续训 MiniCPM5 上采样后 overlay 本文件;同阶段 resume 保留 tokens_in_phase。未知下一张卡:GitHub python3 -m cat_yoko.hw_recipe + scripts/run_b0_next.sh

代码与指针:GitHub AvrovaDonz2026/CAT-YOKO(不用 LFS)。

当前权重

路径 来源 说明
checkpoints/b0/trainable.pt 6000D --try 32 步,MiniCPM5 上采样 gate 0.301;peak 24244 MiB;sha256 9012e5ac55c2f59ef7cacc34d5769444413d070116dbff0696c7b258b9aa0636
checkpoints/b0-nvfp4-try/trainable.pt 6000D NVFP4 wrap --try 2 步 nvfp4_n=2815;gate 0.301;peak 34442 MiB;sha256 461b4ffc05fd46e2668448393789764ccf9dd673644040fe4527259b176a510e
checkpoints/b0-full/trainable.pt Vast B200 发布档 B0 进行中(8e9 信封,seq=4096) 见上一节「当前 B0 快照」。step 26940,sha256 7eebc9a4da78d79be71bbe52881f2a0eaffd899f58ada3a3325f410eca181955。
checkpoints/b0-3090-bf16/trainable.pt RTX 3090 BF16 sibling(同阶段 resume,--no-nvfp4) step 27400,sha256 4951c637…。不是发布口径,不覆盖 b0-full。
checkpoints/b1/trainable.pt 6000D B1 --try(等 GPU) decoder + lm_head + 最终 RMSNorm;resume B0 overlay + MiniCPM5。尚未上传
checkpoints/b2/ 6000D B2 --try(待 GPU) 全模型 overlay;resume B1 + MiniCPM5 encoder/embed。指针 checkpoints/b2/README.md

这些 overlay 里,checkpoints/b0/ 与 checkpoints/b0-nvfp4-try/ 不是 8B token 信封(只是 --try)。checkpoints/b0-full/trainable.pt 是发布信封 进行中 的 B0 overlay(尚未跑完 8e9)。尚未上传 23GiB 全图。GitHub 不存权重、不用 Git LFS。日志在 GitHub artifacts/vast-b200/、artifacts/autodl-rtx6000d/。

无公开评测分数。

许可

产物 许可
本仓库代码与派生权重 Apache-2.0
底座 MiniCPM5-2B Apache-2.0

English (short)

CAT-YOKO-12B is a YOCO-style causal encoder-decoder MoE upcycled from openbmb/MiniCPM5-2B-Base (Apache-2.0, Llama GQA). Code: GitHub. The transformers tag is for tokenizer / shard convention; the graph lives in cat_yoko, not AutoModelForCausalLM.

Item Value
(d) / (V) / (L) 2048 / 130560 / 42 (16 encoder + 26 decoder)
Attention 16 Q / 2 KV, head_dim=128
FFN / MoE SwiGLU 6144; 1+20 routed; top-(k) 7/10 (enc/dec)
Stored params 12,250,381,312 (≈12.25B)
Active encoder ≈2.03B/in, decoder ≈4.33B/out

C1: B0/B1 freeze the encoder. B0 trains new modules only. B1 trains decoder + lm_head + final RMSNorm. B2 trains all. Tokens B0/B1/B2 = 8/27/15B.

Published wall-clock (50B-token envelope, not measured). B200 / SM100 allowed linear GEMMs use TeNvfp4Linear (TE NVFP4BlockScaling; B0 frozen encoder is FPROP, WGRAD is B1/B2). Without TE or on sm_120, Nvfp4Linear E2M1/16 emulation. Attn softmax / SDPA stay fp32.

Recipe H100-h Role
C1+NVFP4 571 published; B0 student bf16; B1/B2 allowed GEMMs NVFP4
C1+FP8 729 Hopper / Ada fallback
joint bf16 1325 100% baseline

Data: Ultra-FineWeb en 55% / zh 30% + UltraData-Math 10% + StarCoder 5% (thinking mix; no 50B download). Tokenizer: openbmb/MiniCPM5-2B.

This Hub has RTX 6000D MiniCPM5-upcycle overlays (not the 8B/27B/15B-token envelopes): checkpoints/b0/trainable.pt (32 steps), checkpoints/b0-nvfp4-try/trainable.pt (2 steps, NVFP4 wrap), checkpoints/b1/ (B1 --try, pending GPU), and checkpoints/b2/ (B2 --try, pending GPU).

Current published B0 overlay (in progress, DummyStream, not finished 8e9; Vast B200 recycled 2026-09-19): checkpoints/b0-full/trainable.pt. Step 26940, tokens_in_phase=130,041,856 (≈1.63% of 8e9), sha256 7eebc9a4da78d79be71bbe52881f2a0eaffd899f58ada3a3325f410eca181955. Last run: micro-batch=2, ~15.7k tok/s. Folder card: checkpoints/b0-full/README.md. Ampere BF16 sibling snapshot (does not replace the published pin): checkpoints/b0-3090-bf16/trainable.pt step 27400. Weights do not live on GitHub. Logs: GitHub artifacts/vast-b200/, artifacts/autodl-rtx6000d/, artifacts/autodl-rtx3090/.

License: Apache-2.0 (this repo, derived weights, and MiniCPM5 base). No eval scores.

Identity and Version

Repository
AvrovaDonz/CAT-YOKO
Publisher
Avrova Donz
Task
Text generation
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
en, zh
Revision
10e9d939aa671a78d4ceb729cc1812559bbc63ef
First published
2026-09-18
Last updated
2026-09-20

Files and Weights

11 files, 1.8 GB in total. The weights are 4 files totalling 1.8 GB in pt.

Weights4 files · 1.8 GB
Configuration1 file · 934 B
Documentation4 files · 22.9 KB
Other1 file · 20.6 KB
Repository1 file · 176 B
Every file
FileTypeSizeSHA-256
checkpoints/b0-3090-bf16/trainable.ptWeights438.5 MB 4951c6373f34
checkpoints/b0-full/trainable.ptWeights438.5 MB 7eebc9a4da78
checkpoints/b0-nvfp4-try/trainable.ptWeights438.5 MB 461b4ffc05fd
checkpoints/b0/trainable.ptWeights438.5 MB 9012e5ac55c2
checkpoints/b0/meta.jsonConfiguration934 B —
LICENSEDocumentation11.3 KB —
README.mdDocumentation8.3 KB —
checkpoints/b0-3090-bf16/README.mdDocumentation1.4 KB —
checkpoints/b0-full/README.mdDocumentation1.9 KB —
checkpoints/b0/metrics.jsonlOther20.6 KB —
.gitattributesRepository176 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.8 GB
Download from Avrova Donz

Released by Avrova Donz through its official repository on Hugging Face. Read the license.

Built From

  • Derived from openbmb/MiniCPM5-2B-Base

Memory Requirements

PrecisionWeights in memory
As published1.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About CAT-YOKO

Can I use CAT-YOKO commercially?

Yes. CAT-YOKO is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers