SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Hy3-Razor-226B-A18B-E144of192

by Mingyang Song Nickyang/Hy3-Razor-226B-A18B-E144of192

Hy3-Razor-226B-A18B-E144of192 is an open-weight model for text generation from Mingyang Song, released under Apache License 2.0. It has 226.3B parameters and a 262,144-token context. At 16-bit it needs about 543.1 GB of GPU memory, which fits on 2x MI355X from $5.18 an hour; at 4-bit, 135.8 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index.

Hy3 with a quarter of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 144 of its original 192 routed experts.

Parameters226.3B
Context262,144
Weights452.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Hy3-Razor-226B-A18B-E144of192 (226.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 452.6 GB 543.1 GB 2x MI355X (288 GB)
Vultr
$5.18 3x MI300X $5.55 · 3x MI325X $6.00
8-bit 226.3 GB 271.6 GB 1x MI355X (288 GB)
Vultr
$2.59 2x MI300X $3.70 · 2x MI325X $4.00
4-bit 113.1 GB 135.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Hy3-Razor-226B-A18B-E144of192 on every accelerator the SAVRN Index prices, at every precision

Model Card

By Mingyang Song, published under apache-2.0, revision 9d19a386b125.

Hy3 with a quarter of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 144 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function…

Read Mingyang Song's full model card

Hy3 with a quarter of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 144 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights.

Base This model
Routed experts per MoE layer 192 144
Total parameters 295B 226B
Active parameters per token ~21B ~18.4B
Active experts per token (top-k) 8 8 (unchanged)
MoE decoder layers 80 80 (unchanged)

Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed.

The other budget is Hy3-Razor-154B-A18B-E96of192.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Nickyang/Hy3-Razor-226B-A18B-E144of192"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, dtype="bfloat16", device_map="auto",
)

messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))

Requires a Transformers build containing the native hy_v3 implementation. Weights are bfloat16.

How the experts were selected

RAZOR asks whether the surviving computation can replace an expert's function, rather than how often or how strongly the expert fires. For a token routed to the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$. Deleting a selected expert $i$ makes the router promote its highest-ranked unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the original selected score sum. With $\lambda$ the routed-output scale, the exact local output change is

$$ \delta_i=\lambda\lVert c-\tilde c^{-i}\rVert_2 =\lambda\,\frac{\lVert w_i r_i - w_r r_r\rVert_2}{1-w_i+w_r}. $$

Scores are aggregated by conditional root mean square over the calibration tokens routed to each expert, and the highest-scoring experts are retained per layer. See the paper for the derivation and the code for the implementation.

Reproducing this checkpoint

Calibration used 32,768-token rows from RazorCal, the 2,048-sample multi-domain corpus released with RAZOR.

pip install razor-moe

razor saliency --model <path-to-Hy3> \
               --data data/RazorCal.json --max-len 32768 --out out/sal
razor prune    --model <path-to-Hy3> --saliency out/sal \
               --method razor --ratio 0.25 --out out/pruned

kept_expert_indices.json is the keep-set manifest this checkpoint was built from, in the format razor verify expects:

razor verify --pruned . --model <path-to-Hy3>

Selection depends on the calibration draw, so an independent run reproduces the procedure rather than this exact expert set.

Limitations

Expert pruning is lossy. Benchmark retention and predictive fidelity do not guarantee stable generation: the paper reports that responses change in diversity, formatting and termination behaviour even where task accuracy is largely preserved. Evaluate on your own workload before deploying.

The calibration corpus is multi-domain but finite, so behaviour on domains far from it is not characterised by the released measurements.

License and attribution

This is a derivative of Hy3, released under the Apache-2.0 license, and is distributed under the same license. The retained weights are the base model's own weights; RAZOR contributes the expert selection, not new parameters.

The RAZOR code is Apache-2.0. RazorCal records remain subject to their applicable upstream terms; see data/LICENSE-DATA.

Citation

@misc{song2026razorpruningreplaceableexperts,
      title={RAZOR: Pruning Replaceable Experts in LLMs},
      author={Mingyang Song and Mao Zheng},
      year={2026},
      eprint={2609.30465},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.30465},
}

Configuration

Architecture
HYV3ForCausalLM
Context length (tokens)
262,144
Layers
80
Hidden size
4,096
Feed-forward size
13,312
Attention heads
64
Key/value heads
8
Head dimension
128
Vocabulary size
120,832
Experts
144
Experts active per token
8
RoPE base
1.11588e+07
Model type
hy_v3

Identity and Version

Repository
Nickyang/Hy3-Razor-226B-A18B-E144of192
Publisher
Mingyang Song
Task
Text generation
Modality
Text
Library
transformers
Parameters
226.3B parameters
Languages
moe
Revision
9d19a386b125f3d528999aea34f5cc4f2123fec4
First published
2026-09-28
Last updated
2026-09-29

Files and Weights

66 files, 452.6 GB in total. The weights are 57 files totalling 452.6 GB in safetensors.

Weights57 files · 452.6 GB
Configuration4 files · 3.0 MB
Tokenizer2 files · 9.7 MB
Documentation1 file · 4.9 KB
Other1 file · 10.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001.safetensorsWeights8.0 GB 6bd395943085
model-00002.safetensorsWeights8.0 GB 466d8b7b045f
model-00003.safetensorsWeights8.0 GB 954853c337a6
model-00004.safetensorsWeights8.0 GB 22582b276ac4
model-00005.safetensorsWeights8.0 GB ee24dac4e971
model-00006.safetensorsWeights8.0 GB 6cdd00d9967f
model-00007.safetensorsWeights8.4 GB 48ccf48e5f9e
model-00008.safetensorsWeights8.0 GB 36af335c7ff2
model-00009.safetensorsWeights8.0 GB 766825678ed7
model-00010.safetensorsWeights8.0 GB c977aaeaf7b9
model-00011.safetensorsWeights8.0 GB d2611dccdddd
model-00012.safetensorsWeights8.0 GB 6a2fa0cf7eec
model-00013.safetensorsWeights8.0 GB 78be62dc0ff5
model-00014.safetensorsWeights8.0 GB 4c9371bf474e
model-00015.safetensorsWeights8.0 GB b3cf965d7424
model-00016.safetensorsWeights8.0 GB 4e9e26d2cc4f
model-00017.safetensorsWeights8.0 GB 6af043c35213
model-00018.safetensorsWeights8.0 GB ec61efb035a1
model-00019.safetensorsWeights8.0 GB efb72a534387
model-00020.safetensorsWeights8.0 GB fab1d649b0b1
model-00021.safetensorsWeights8.0 GB 61636d125fa1
model-00022.safetensorsWeights8.0 GB 97ed11a5d517
model-00023.safetensorsWeights8.0 GB edd811cb9caf
model-00024.safetensorsWeights8.0 GB 48e998290b38
model-00025.safetensorsWeights8.0 GB a8d7f682474e
model-00026.safetensorsWeights8.0 GB a581f7a15d91
model-00027.safetensorsWeights8.0 GB cc4215ba9acd
model-00028.safetensorsWeights8.0 GB 7951e20c1a0d
model-00029.safetensorsWeights8.0 GB 56652f464e48
model-00030.safetensorsWeights8.0 GB b215c1e50dc3
model-00031.safetensorsWeights8.0 GB e8390a38537c
model-00032.safetensorsWeights8.0 GB b346c1ed91cb
model-00033.safetensorsWeights8.0 GB 98162d57c39c
model-00034.safetensorsWeights8.0 GB 3aeb0475b5a2
model-00035.safetensorsWeights8.0 GB dd0b6abd4bde
model-00036.safetensorsWeights8.0 GB 78580178b7c6
model-00037.safetensorsWeights8.0 GB 5f43771e831e
model-00038.safetensorsWeights8.0 GB 9127122fe1eb
model-00039.safetensorsWeights8.0 GB a8887aa43aa1
model-00040.safetensorsWeights8.0 GB 976b66bca67e
model-00041.safetensorsWeights8.0 GB 4e78c0bb4f02
model-00042.safetensorsWeights8.0 GB 7451f19f9cf0
model-00043.safetensorsWeights8.0 GB 6e76fdad4866
model-00044.safetensorsWeights8.0 GB 6c102d5a7450
model-00045.safetensorsWeights8.0 GB 8884d59bb482
model-00046.safetensorsWeights8.0 GB 74a52864fb76
model-00047.safetensorsWeights8.0 GB d8a07640b9bf
model-00048.safetensorsWeights8.0 GB 8774f43ed132
model-00049.safetensorsWeights8.0 GB 8093ad57f83c
model-00050.safetensorsWeights8.0 GB 5a783fde625d
model-00051.safetensorsWeights8.0 GB 4eb095a2c2e7
model-00052.safetensorsWeights8.0 GB cf905d25d3ac
model-00053.safetensorsWeights8.0 GB ae726a47a759
model-00054.safetensorsWeights8.0 GB 7be4f8ff72aa
model-00055.safetensorsWeights8.0 GB 886bb4051f51
model-00056.safetensorsWeights8.0 GB fe9cbcdf2b89
model-00057.safetensorsWeights4.0 GB 423484aa7508
config.jsonConfiguration2.3 KB —
kept_expert_indices.jsonConfiguration98.6 KB —
model.safetensors.index.jsonConfiguration2.9 MB —
special_tokens_map.jsonConfiguration491 B —
README.mdDocumentation4.9 KB —
chat_template.jinjaOther10.2 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer9.5 MB —
tokenizer_config.jsonTokenizer166.0 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
452.6 GB
Download from Mingyang Song

Released by Mingyang Song through its official repository on Hugging Face. Read the license.

Built From

  • Derived from tencent/Hy3
  • Described by arXiv:2609.30465

Memory Requirements

PrecisionWeights in memory
As published452.6 GB
16-bit452.6 GB
8-bit226.3 GB
4-bit113.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Hy3-Razor-226B-A18B-E144of192

How much GPU memory does Hy3-Razor-226B-A18B-E144of192 need?

About 543.1 GB at 16-bit and 135.8 GB at 4-bit: the weights (226.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Hy3-Razor-226B-A18B-E144of192 on?

At 16-bit, 2x MI355X from $5.18 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Hy3-Razor-226B-A18B-E144of192 commercially?

Yes. Hy3-Razor-226B-A18B-E144of192 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Hy3-Razor-226B-A18B-E144of192's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

MiniMax-M2.7

MiniMax

Join Our WeChat Discord community. MiniMax Agent API CLI MiniMax Website Hugging Face GitHub ModelScope LICENSE MiniMax-M2.7 is our first model deeply participating in its own evolution. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search. For more details, see our blog post. M2.7 initiates a cycle of model self-evolution: during development, we let the model update its own memory, build dozens of complex skills for RL experiments, and improve its own learning process based on experiment results. An internal version of M2.7 autonomously optimized a programming scaffold over…

Open weights other 228.7B parameters 204,800 tokens transformers

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

Model · Text generation

DeepSeek-V4-Flash

DeepSeek

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache…

Open weights mit 290.9B parameters 1,048,576 tokens transformers

Model · Text generation

MiMo-V2.6-Flash-RL

Xiaomi MiMo

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers

Model · Text generation

MiMo-V2.6-Flash-RL

Zachary Howard

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers

Model · Text generation

Hy3-Razor-154B-A18B-E96of192

Mingyang Song

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…

Open weights apache-2.0 153.8B parameters 262,144 tokens transformers