SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8

by Manish manishkumar2101114/qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8

Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128.

Parameters36B
Context262,144
Weights37.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8 (36B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 71.9 GB 86.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 36.0 GB 43.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 18.0 GB 21.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Manish, published under apache-2.0, revision 34876cee5ef0.

Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128. Chat template: qwen3template333.jinja. Serve with sglang / vllm / transformers.

Read Manish's full model card

Qwen3.6-35B-A3B tool-prompts-ckpt444 FP8

Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128. Chat template: qwen3_template333.jinja. Serve with sglang / vllm / transformers.

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
40
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe
Quantization
fp8

Identity and Version

Repository
manishkumar2101114/qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8
Publisher
Manish
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
36B parameters
Languages
en, hi
Revision
34876cee5ef0dd26bd0d4434e7a3c679ae855bf4
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

54 files, 37.5 GB in total. The weights are 42 files totalling 37.5 GB in safetensors.

Weights42 files · 37.5 GB
Configuration5 files · 6.4 MB
Tokenizer4 files · 30.1 MB
Documentation1 file · 554 B
Other1 file · 7.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
layers-0.safetensorsWeights843.7 MB d1a79d4b0834
layers-1.safetensorsWeights843.7 MB 99794a1ff4ad
layers-10.safetensorsWeights843.7 MB 434fc790e2e0
layers-11.safetensorsWeights837.1 MB bf39611e632f
layers-12.safetensorsWeights843.7 MB 0830c9a1dd3d
layers-13.safetensorsWeights843.7 MB fe3dd4fb1f65
layers-14.safetensorsWeights843.7 MB c830afd9f39b
layers-15.safetensorsWeights837.1 MB 4cc07338763e
layers-16.safetensorsWeights843.7 MB 2dcc86930eb8
layers-17.safetensorsWeights843.7 MB 95c6edbac4fc
layers-18.safetensorsWeights843.7 MB 042fa12718c2
layers-19.safetensorsWeights837.1 MB 2044e10e51b2
layers-2.safetensorsWeights843.7 MB 5dcdc04d7984
layers-20.safetensorsWeights843.7 MB 57f946b9da06
layers-21.safetensorsWeights843.7 MB 1a38a5b3ab24
layers-22.safetensorsWeights843.7 MB 24c242cb97e0
layers-23.safetensorsWeights837.1 MB 97722ede6b8e
layers-24.safetensorsWeights843.7 MB 943cba343e31
layers-25.safetensorsWeights843.7 MB cba498117a7b
layers-26.safetensorsWeights843.7 MB cdf48ebfadba
layers-27.safetensorsWeights837.1 MB 70ae4be3d7e3
layers-28.safetensorsWeights843.7 MB 1b3dad1e99dc
layers-29.safetensorsWeights843.7 MB f5f0d08646aa
layers-3.safetensorsWeights837.1 MB 53515869004a
layers-30.safetensorsWeights843.7 MB f4c1bc824fbb
layers-31.safetensorsWeights837.1 MB b4a347ec0391
layers-32.safetensorsWeights843.7 MB c769571d33e5
layers-33.safetensorsWeights843.7 MB be97a4576226
layers-34.safetensorsWeights843.7 MB ac5ebd496ae3
layers-35.safetensorsWeights837.1 MB 96a5e1fcca74
layers-36.safetensorsWeights843.7 MB f9150a13f352
layers-37.safetensorsWeights843.7 MB 8a307696af3f
layers-38.safetensorsWeights843.7 MB 558c1690581b
layers-39.safetensorsWeights837.1 MB 831de570dcda
layers-4.safetensorsWeights843.7 MB 350e3f4b19bc
layers-5.safetensorsWeights843.7 MB 9da8bae52005
layers-6.safetensorsWeights843.7 MB a0d79aa47f4c
layers-7.safetensorsWeights837.1 MB 2ea7c27ea283
layers-8.safetensorsWeights843.7 MB 0e84769562ef
layers-9.safetensorsWeights843.7 MB 1421379a2c76
mtp.safetensorsWeights853.9 MB 6c4b7365adbd
outside.safetensorsWeights2.9 GB bb9755bc8ebd
config.jsonConfiguration37.0 KB
generation_config.jsonConfiguration202 B
model.safetensors.index.jsonConfiguration6.3 MB
preprocessor_config.jsonConfiguration390 B
video_preprocessor_config.jsonConfiguration385 B
README.mdDocumentation554 B
chat_template.jinjaOther7.5 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
37.5 GB
Download from Manish

Released by Manish through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published37.5 GB
16-bit71.9 GB
8-bit36.0 GB
4-bit18.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8

How much GPU memory does qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8 need?

About 86.3 GB at 16-bit and 21.6 GB at 4-bit: the weights (36B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8 commercially?

Yes. qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is qwen3.6-35b-a3b-tool-prompts-ckpt444-merged-fp8's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.6-35B-A3B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-35B-A3B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-35B-A3B-FP8

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-35B-A3B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers