YOCO 式因果 encoder-decoder MoE。从 MiniCPM5-2B 上采样:openbmb/MiniCPM5-2B-Base(Apache-2.0,Llama GQA)。 libraryname: transformers 只表示 tokenizer / 分片约定走 HuggingFace 生态。当前图是仓库里的 catyoko PyTorch 实现,不是 Hub 上可 AutoModelForCausalLM 直接加载的架构。 50B token 信封上的理论墙钟,不是实测。B200 / SM100 允许的线性 GEMM 走 TeNvfp4Linear(TE NVFP4BlockScaling;B0 冻 encoder 走 FPROP,WGRAD 留给 B1/B2)。无 TE / sm120 时 Nvfp4Linear E2M1/16 仿真。attn softmax / SDPA 仍 fp32。 发布信封 进行中(DummyStream,尚未跑完 8e9)。Vast B200 已回收(2026-09-19)。本文件是释放前快照,不是终局。GitHub 口径:docs/STATUS.md。 代码与指针:GitHub AvrovaDonz2026/CAT-YOKO(不用 LFS)。 这些 overlay 里,checkpoints/b0/ 与 checkpoints/b0-nvfp4-try/ 不是 8B token 信封(只是 --try)。checkpoints/b0-full/trainable.pt 是发布信封 进行中 的 B0…
Open weights
apache-2.0
transformers
This repository contains QPR artifacts for the dense, unpruned openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with group size 128 and GPTQ calibration; it does not apply 2:4 pruning. The recommended file is dense-int4-distill2000-v2-best.qpr. The minicpm5-2b-dense-int4-g128.qpr file is the matching dense-int4 baseline. Each QPR file is about 1.30 GB and is stored with Git LFS. The candidate was selected from a 2,000-update distillation run starting at the dense-int4 baseline. It was selected at update 775 and independently reloaded before scoring. On 100 held-out Chinese and English prompts with 512 new tokens, direct gpt-6-sol judging gave 39 baseline wins, 58 candidate…
Open weights
mpl-2.0