SAVRN
Search Contact SAVRN

Independent publisher

Mingyang Song

Nickyang

LRMs, Long-Context LLMs, LLM Judges, Many-Shot ICL

Models in Library2
Datasets in Library0
Models on Hugging Face13
Followers2

Models

Model · Text generation

Hy3-Razor-226B-A18B-E144of192

Mingyang Song

Hy3 with a quarter of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 144 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function…

Open weights apache-2.0 226.3B parameters 262,144 tokens transformers

Model · Text generation

Hy3-Razor-154B-A18B-E96of192

Mingyang Song

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…

Open weights apache-2.0 153.8B parameters 262,144 tokens transformers