SAVRN
Search Contact SAVRN

Open-weight model · Text generation

less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50

by XINKAI ZOU jayzou3773/less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50

less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 is an open-weight model for text generation from XINKAI ZOU, released under Apache License 2.0. It has 64.1B parameters and a 262,144-token context. At 16-bit it needs about 153.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method.

Parameters64.1B
Context262,144
Weights128.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 (64.1B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 128.3 GB 153.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 64.1 GB 77.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 32.1 GB 38.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 20, 2026.

less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 on every accelerator the SAVRN Index prices, at every precision

Model Card

By XINKAI ZOU, published under apache-2.0, revision e6669506c798.

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Read XINKAI ZOU's full model card

IntDim-L 50% pruned Qwen/Qwen3.5-122B-A10B

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqa_main configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selection_seed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer max_length, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16.

The source-row selection hash is 790c4c22309def44542965fdde7c5f38f1d8e354602640cfb31518134b8d92e6 and the model-specific token-file hash is 4cecf02da096c0d1c1f8f01bbdf8867186ab3064eccbfbb34cd9c16a89564d62. Full export and zero-mask equivalence metadata are in experiment-export.json. The exact calibration and held-out test rows are in the private dataset jayzou3773/less-is-moe-gpqa-main-calibration-64 revision b9596e85179b3017f77ba1436a5d2e61b6a61a5b, following GPQA's access terms.

Inference requires the Less-is-MoE ragged vLLM plugin from the unified Less-is-MoE GPU image. IntDim-E has one uniform expert width. IntDim-L/G retain the routed MoE topology and store compact per-expert widths in config.json.

Configuration

Architecture
RaggedQwen3_5MoeForCausalLM
Context length (tokens)
262,144
Layers
48
Hidden size
3,072
Attention heads
32
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe_text

Identity and Version

Repository
jayzou3773/less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50
Publisher
XINKAI ZOU
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
64.1B parameters
Languages
moe
Revision
e6669506c798631f571db34b7b7f170239183b63
First published
2026-09-20
Last updated
2026-09-20

Files and Weights

58 files, 128.3 GB in total. The weights are 48 files totalling 128.3 GB in safetensors.

Weights48 files · 128.3 GB
Configuration5 files · 401.4 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 1.5 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights3.3 GB beb6c6104334
model-00002-of-00048.safetensorsWeights3.6 GB d689b782a9e5
model-00003-of-00048.safetensorsWeights2.6 GB f20ccd349294
model-00004-of-00048.safetensorsWeights2.6 GB 2d76689fc52d
model-00005-of-00048.safetensorsWeights2.6 GB 80de59cb83c1
model-00006-of-00048.safetensorsWeights2.6 GB ad446106442f
model-00007-of-00048.safetensorsWeights2.6 GB 42cdb3c06c9c
model-00008-of-00048.safetensorsWeights2.6 GB 3066874b4508
model-00009-of-00048.safetensorsWeights2.6 GB 65671c5ac7a1
model-00010-of-00048.safetensorsWeights2.6 GB 4c80a15be9f5
model-00011-of-00048.safetensorsWeights2.6 GB aace4c262fff
model-00012-of-00048.safetensorsWeights2.6 GB 3f4aa20ca8ee
model-00013-of-00048.safetensorsWeights2.6 GB be57e43620f6
model-00014-of-00048.safetensorsWeights2.6 GB b192eb481834
model-00015-of-00048.safetensorsWeights2.6 GB 235b0c0a5577
model-00016-of-00048.safetensorsWeights2.6 GB e9c56bf25c10
model-00017-of-00048.safetensorsWeights2.6 GB 99b41dda1dcd
model-00018-of-00048.safetensorsWeights2.6 GB 9c8c5bfe65f6
model-00019-of-00048.safetensorsWeights2.6 GB e47ea0c9dbf7
model-00020-of-00048.safetensorsWeights2.6 GB 2f0df1fa1bde
model-00021-of-00048.safetensorsWeights2.6 GB 4dd66cbd0fe3
model-00022-of-00048.safetensorsWeights2.6 GB 0ae4225a12ec
model-00023-of-00048.safetensorsWeights2.6 GB 13705f200735
model-00024-of-00048.safetensorsWeights2.6 GB d5a6675bf7b5
model-00025-of-00048.safetensorsWeights2.6 GB a2304eb0362e
model-00026-of-00048.safetensorsWeights2.6 GB 16c709d457bd
model-00027-of-00048.safetensorsWeights2.6 GB ae6a23b842ce
model-00028-of-00048.safetensorsWeights2.6 GB fd768af8d979
model-00029-of-00048.safetensorsWeights2.6 GB 494176692dff
model-00030-of-00048.safetensorsWeights2.6 GB dab80c556549
model-00031-of-00048.safetensorsWeights2.6 GB 0f4a80dec308
model-00032-of-00048.safetensorsWeights2.6 GB fd7c99e4e1d4
model-00033-of-00048.safetensorsWeights2.6 GB ef81a556736c
model-00034-of-00048.safetensorsWeights2.6 GB 5d664c604a49
model-00035-of-00048.safetensorsWeights2.6 GB 4a189d673bac
model-00036-of-00048.safetensorsWeights2.6 GB 7cec05e2165b
model-00037-of-00048.safetensorsWeights2.6 GB b7980c0d42fd
model-00038-of-00048.safetensorsWeights2.6 GB 4258293cb30c
model-00039-of-00048.safetensorsWeights2.6 GB dbbbd3f9c5a2
model-00040-of-00048.safetensorsWeights2.6 GB 94863f75de79
model-00041-of-00048.safetensorsWeights2.6 GB 42f8bcacf64d
model-00042-of-00048.safetensorsWeights2.6 GB 413c17a11a59
model-00043-of-00048.safetensorsWeights2.6 GB 69c62b526f24
model-00044-of-00048.safetensorsWeights2.6 GB 1f717a2be883
model-00045-of-00048.safetensorsWeights2.6 GB 1b953b360531
model-00046-of-00048.safetensorsWeights2.6 GB 90bffe7ad2ed
model-00047-of-00048.safetensorsWeights2.6 GB f7cb41932cc2
model-00048-of-00048.safetensorsWeights4.0 GB 66133ab6bedb
config.jsonConfiguration160.1 KB
experiment-export.jsonConfiguration159.7 KB
generation_config.jsonConfiguration214 B
model.safetensors.index.jsonConfiguration71.2 KB
publication-metadata.jsonConfiguration10.2 KB
README.mdDocumentation1.5 KB
chat_template.jinjaOther7.8 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
128.3 GB
Download from XINKAI ZOU

Released by XINKAI ZOU through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Qwen/Qwen3.5-122B-A10B

Memory Requirements

PrecisionWeights in memory
As published128.3 GB
16-bit128.3 GB
8-bit64.1 GB
4-bit32.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50

How much GPU memory does less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 need?

About 153.9 GB at 16-bit and 38.5 GB at 4-bit: the weights (64.1B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 commercially?

Yes. less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

Model · Text generation

Qwen3.5-122B-A10B-NVFP4

NVIDIA

The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.5-122B-A10B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems…

Open weights apache-2.0 64.6B parameters 262,144 tokens Model Optimizer

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-NVFP4 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 67.2B parameters 262,144 tokens transformers

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,511 tokens. The source MXFP4 checkpoint was explicitly dequantized to BF16 before scoring and pruning. The source-row selection hash is…

Open weights apache-2.0 59.5B parameters 131,072 tokens

Model · Text generation

Llama-3.3-70B-Instruct

Meta Llama

The Meta Llama 3.3 multilingual large language model (LLM) is an instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.3 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and…

Access requested at publisher llama3.3 70.6B parameters transformers