Hy3 with a quarter of its routed experts removed by
RAZOR, a training-free expert pruning
method. Every MoE layer keeps 144 of its original 192 routed experts. No
gradient updates or recovery training were applied: the retained weights are
the base model's own weights.
|
Base |
This model |
| Routed experts per MoE layer |
192 |
144 |
| Total parameters |
295B |
226B |
| Active parameters per token |
~21B |
~18.4B |
| Active experts per token (top-k) |
8 |
8 (unchanged) |
| MoE decoder layers |
80 |
80 (unchanged) |
Pruning touches only the routed expert pool. Attention, shared experts, the
embedding and the LM head are untouched, so the compute per token drops only by
the share of expert FLOPs that the removed experts would have contributed.
The other budget is
Hy3-Razor-154B-A18B-E96of192.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Nickyang/Hy3-Razor-226B-A18B-E144of192"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype="bfloat16", device_map="auto",
)
messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))
Requires a Transformers build containing the native hy_v3 implementation.
Weights are bfloat16.
How the experts were selected
RAZOR asks whether the surviving computation can replace an expert's function,
rather than how often or how strongly the expert fires. For a token routed to
the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be
the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$.
Deleting a selected expert $i$ makes the router promote its highest-ranked
unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the
original selected score sum. With $\lambda$ the routed-output scale, the exact
local output change is
$$
\delta_i=\lambda\lVert c-\tilde c^{-i}\rVert_2
=\lambda\,\frac{\lVert w_i r_i - w_r r_r\rVert_2}{1-w_i+w_r}.
$$
Scores are aggregated by conditional root mean square over the calibration
tokens routed to each expert, and the highest-scoring experts are retained per
layer. See the paper for the derivation and
the code for the implementation.
Reproducing this checkpoint
Calibration used 32,768-token rows from
RazorCal, the 2,048-sample
multi-domain corpus released with RAZOR.
pip install razor-moe
razor saliency --model <path-to-Hy3> \
--data data/RazorCal.json --max-len 32768 --out out/sal
razor prune --model <path-to-Hy3> --saliency out/sal \
--method razor --ratio 0.25 --out out/pruned
kept_expert_indices.json is the keep-set manifest this checkpoint was built
from, in the format razor verify expects:
razor verify --pruned . --model <path-to-Hy3>
Selection depends on the calibration draw, so an independent run reproduces the
procedure rather than this exact expert set.
Limitations
Expert pruning is lossy. Benchmark retention and predictive fidelity do not
guarantee stable generation: the paper reports that responses change in
diversity, formatting and termination behaviour even where task accuracy is
largely preserved. Evaluate on your own workload before deploying.
The calibration corpus is multi-domain but finite, so behaviour on domains far
from it is not characterised by the released measurements.
License and attribution
This is a derivative of Hy3, released
under the Apache-2.0 license, and is distributed under the same license. The
retained weights are the base model's own weights; RAZOR contributes the expert
selection, not new parameters.
The RAZOR code is Apache-2.0. RazorCal records remain subject to their
applicable upstream terms; see
data/LICENSE-DATA.
Citation
@misc{song2026razorpruningreplaceableexperts,
title={RAZOR: Pruning Replaceable Experts in LLMs},
author={Mingyang Song and Mao Zheng},
year={2026},
eprint={2609.30465},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.30465},
}