SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Hy3-Razor-154B-A18B-E96of192

by Mingyang Song Nickyang/Hy3-Razor-154B-A18B-E96of192

Hy3-Razor-154B-A18B-E96of192 is an open-weight model for text generation from Mingyang Song, released under Apache License 2.0. It has 153.8B parameters and a 262,144-token context. At 16-bit it needs about 369.1 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 92.3 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index.

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts.

Parameters153.8B
Context262,144
Weights307.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Hy3-Razor-154B-A18B-E96of192 (153.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 307.6 GB 369.1 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
8-bit 153.8 GB 184.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 76.9 GB 92.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Hy3-Razor-154B-A18B-E96of192 on every accelerator the SAVRN Index prices, at every precision

Model Card

By Mingyang Song, published under apache-2.0, revision 6b1193aa6873.

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…

Read Mingyang Song's full model card

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights.

Base This model
Routed experts per MoE layer 192 96
Total parameters 295B 154B
Active parameters per token ~21B ~18.4B
Active experts per token (top-k) 8 8 (unchanged)
MoE decoder layers 80 80 (unchanged)

Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed.

The other budget is Hy3-Razor-226B-A18B-E144of192.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Nickyang/Hy3-Razor-154B-A18B-E96of192"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, dtype="bfloat16", device_map="auto",
)

messages = [{"role": "user", "content": "Explain mixture-of-experts routing."}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=256)[0]))

Requires a Transformers build containing the native hy_v3 implementation. Weights are bfloat16.

How the experts were selected

RAZOR asks whether the surviving computation can replace an expert's function, rather than how often or how strongly the expert fires. For a token routed to the set $S$ with normalized weights $w_j$, let $c=\sum_{j\in S} w_j f_j$ be the routed mixture and $r_j=f_j-c$ the consensus residual of expert $j$. Deleting a selected expert $i$ makes the router promote its highest-ranked unselected expert $r$, whose pseudo-weight $w_r$ is its score divided by the original selected score sum. With $\lambda$ the routed-output scale, the exact local output change is

$$ \delta_i=\lambda\lVert c-\tilde c^{-i}\rVert_2 =\lambda\,\frac{\lVert w_i r_i - w_r r_r\rVert_2}{1-w_i+w_r}. $$

Scores are aggregated by conditional root mean square over the calibration tokens routed to each expert, and the highest-scoring experts are retained per layer. See the paper for the derivation and the code for the implementation.

Reproducing this checkpoint

Calibration used 32,768-token rows from RazorCal, the 2,048-sample multi-domain corpus released with RAZOR.

pip install razor-moe

razor saliency --model <path-to-Hy3> \
               --data data/RazorCal.json --max-len 32768 --out out/sal
razor prune    --model <path-to-Hy3> --saliency out/sal \
               --method razor --ratio 0.5 --out out/pruned

kept_expert_indices.json is the keep-set manifest this checkpoint was built from, in the format razor verify expects:

razor verify --pruned . --model <path-to-Hy3>

Selection depends on the calibration draw, so an independent run reproduces the procedure rather than this exact expert set.

Limitations

Expert pruning is lossy. Benchmark retention and predictive fidelity do not guarantee stable generation: the paper reports that responses change in diversity, formatting and termination behaviour even where task accuracy is largely preserved. Evaluate on your own workload before deploying.

The calibration corpus is multi-domain but finite, so behaviour on domains far from it is not characterised by the released measurements.

License and attribution

This is a derivative of Hy3, released under the Apache-2.0 license, and is distributed under the same license. The retained weights are the base model's own weights; RAZOR contributes the expert selection, not new parameters.

The RAZOR code is Apache-2.0. RazorCal records remain subject to their applicable upstream terms; see data/LICENSE-DATA.

Citation

@misc{song2026razorpruningreplaceableexperts,
      title={RAZOR: Pruning Replaceable Experts in LLMs},
      author={Mingyang Song and Mao Zheng},
      year={2026},
      eprint={2609.30465},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.30465},
}

Configuration

Architecture
HYV3ForCausalLM
Context length (tokens)
262,144
Layers
80
Hidden size
4,096
Feed-forward size
13,312
Attention heads
64
Key/value heads
8
Head dimension
128
Vocabulary size
120,832
Experts
96
Experts active per token
8
RoPE base
1.11588e+07
Model type
hy_v3

Identity and Version

Repository
Nickyang/Hy3-Razor-154B-A18B-E96of192
Publisher
Mingyang Song
Task
Text generation
Modality
Text
Library
transformers
Parameters
153.8B parameters
Languages
moe
Revision
6b1193aa6873dd5a99e0dea2742cad08c0c4c4ab
First published
2026-09-28
Last updated
2026-09-29

Files and Weights

48 files, 307.6 GB in total. The weights are 39 files totalling 307.6 GB in safetensors.

Weights39 files · 307.6 GB
Configuration4 files · 2.0 MB
Tokenizer2 files · 9.7 MB
Documentation1 file · 4.9 KB
Other1 file · 10.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001.safetensorsWeights8.0 GB c9973de2f98c
model-00002.safetensorsWeights8.0 GB 02bc9f925fdc
model-00003.safetensorsWeights8.0 GB 70bcfea8a001
model-00004.safetensorsWeights8.0 GB c8e5b57b6e5d
model-00005.safetensorsWeights8.0 GB 55f6670f3e18
model-00006.safetensorsWeights8.0 GB e14b8cc6b3d9
model-00007.safetensorsWeights8.0 GB 510be105240d
model-00008.safetensorsWeights8.0 GB 3a728ef2602a
model-00009.safetensorsWeights8.0 GB 47aa80633d51
model-00010.safetensorsWeights8.0 GB 5106a0224377
model-00011.safetensorsWeights8.0 GB cf8233b5abd4
model-00012.safetensorsWeights8.0 GB a8a0a154d67d
model-00013.safetensorsWeights8.0 GB cdb0c45e98aa
model-00014.safetensorsWeights8.0 GB 37bf35f286c4
model-00015.safetensorsWeights8.0 GB 74daacbca08c
model-00016.safetensorsWeights8.0 GB b9b2f306773a
model-00017.safetensorsWeights8.0 GB a005f4da7eb4
model-00018.safetensorsWeights8.0 GB 0248f89abc04
model-00019.safetensorsWeights8.0 GB 08339c601d07
model-00020.safetensorsWeights8.0 GB fb2f06b381c0
model-00021.safetensorsWeights8.0 GB 682816dad667
model-00022.safetensorsWeights8.0 GB 6a49293ddcb6
model-00023.safetensorsWeights8.0 GB 751873c617f9
model-00024.safetensorsWeights8.0 GB c98e582642c3
model-00025.safetensorsWeights8.0 GB ee143b9ced6a
model-00026.safetensorsWeights8.0 GB a2915bacbd47
model-00027.safetensorsWeights8.0 GB 9e682a16bb2c
model-00028.safetensorsWeights8.0 GB f47ef636b9f2
model-00029.safetensorsWeights8.0 GB 1c1d8e16d0bc
model-00030.safetensorsWeights8.0 GB f28beea8c054
model-00031.safetensorsWeights8.0 GB 7a571aee9765
model-00032.safetensorsWeights8.0 GB 31e6d913aa11
model-00033.safetensorsWeights8.0 GB 3a34973e4e92
model-00034.safetensorsWeights8.0 GB e7167868d0c5
model-00035.safetensorsWeights8.0 GB f4718fd5f925
model-00036.safetensorsWeights8.0 GB 9faccaa2b1af
model-00037.safetensorsWeights8.0 GB a33e1d30518f
model-00038.safetensorsWeights8.0 GB c959c125cde7
model-00039.safetensorsWeights3.4 GB 91ebe29aa432
config.jsonConfiguration2.3 KB —
kept_expert_indices.jsonConfiguration66.3 KB —
model.safetensors.index.jsonConfiguration1.9 MB —
special_tokens_map.jsonConfiguration491 B —
README.mdDocumentation4.9 KB —
chat_template.jinjaOther10.2 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer9.5 MB —
tokenizer_config.jsonTokenizer166.0 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
307.6 GB
Download from Mingyang Song

Released by Mingyang Song through its official repository on Hugging Face. Read the license.

Built From

  • Derived from tencent/Hy3
  • Described by arXiv:2609.30465

Memory Requirements

PrecisionWeights in memory
As published307.6 GB
16-bit307.6 GB
8-bit153.8 GB
4-bit76.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Hy3-Razor-154B-A18B-E96of192

How much GPU memory does Hy3-Razor-154B-A18B-E96of192 need?

About 369.1 GB at 16-bit and 92.3 GB at 4-bit: the weights (153.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Hy3-Razor-154B-A18B-E96of192 on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Hy3-Razor-154B-A18B-E96of192 commercially?

Yes. Hy3-Razor-154B-A18B-E96of192 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Hy3-Razor-154B-A18B-E96of192's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

MiMo-V2.6-Flash-RL

Xiaomi MiMo

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers

Model · Text generation

MiMo-V2.6-Flash-RL

Zachary Howard

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers

Model · Text generation

DeepSeek-V4-Flash-DSpark

DeepSeek

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid…

Open weights mit 165.3B parameters 1,048,576 tokens transformers

Model · Text generation

Vinci-Cyber-123B-1.0

SimpleDirect

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…

Open weights other 125B parameters 262,144 tokens transformers

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Open weights other 124B parameters 262,144 tokens vllm

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers