SAVRN
Search Contact SAVRN

SAVRN Model Hub · Qwen3-235B-A22B-Thinking-2507

Qwen3-235B-A22B-Thinking-2507 GPU Requirements

Qwen3-235B-A22B-Thinking-2507 needs about 564 GB of GPU memory at 16-bit, 282 GB at 8-bit and 141 GB at 4-bit: its 235.1B parameters plus a 20% working margin. The cheapest setup at 16-bit is 2x MI355X from $5.18 an hour, about $3,781 a month around the clock. None of the 9 accelerators the SAVRN Index prices holds it on one card at 16-bit.

Parameters235.1B
Memory at 16-bit564 GB
Memory at 8-bit282 GB
Memory at 4-bit141 GB
Context262,144

Every Accelerator, Every Precision

How many cards of each accelerator the SAVRN Index prices it takes to hold Qwen3-235B-A22B-Thinking-2507, and what that many cards cost an hour at the lowest listed on-demand price. Memory needed: 564 GB at 16-bit, 282 GB at 8-bit, 141 GB at 4-bit.

AcceleratorMemory per cardLowest price per card16-bit8-bit4-bit
H100
Voltage Park
80 GB $1.99 8 cards
$15.92/hr
4 cards
$7.96/hr
2 cards
$3.98/hr
H200
GMI Cloud
141 GB $2.60 5 cards
$13.00/hr
3 cards
$7.80/hr
2 cards
$5.20/hr
B200
Vultr
180 GB $3.50 4 cards
$14.00/hr
2 cards
$7.00/hr
1 card
$3.50/hr
GB200 NVL72
GMI Cloud
186 GB $8.00 4 cards
$32.00/hr
2 cards
$16.00/hr
1 card
$8.00/hr
MI300X
Vultr
192 GB $1.85 3 cards
$5.55/hr
2 cards
$3.70/hr
1 card
$1.85/hr
MI325X
Vultr
256 GB $2.00 3 cards
$6.00/hr
2 cards
$4.00/hr
1 card
$2.00/hr
B300
Massed Compute
268 GB $6.60 3 cards
$19.80/hr
2 cards
$13.20/hr
1 card
$6.60/hr
GB300 NVL72
Verda
279 GB $10.32 3 cards
$30.96/hr
2 cards
$20.64/hr
1 card
$10.32/hr
MI355X
Vultr
288 GB $2.59 2 cards
$5.18/hr
1 card
$2.59/hr
1 card
$2.59/hr

Running It Around the Clock

PrecisionCheapest setupPer hourPer month (730 hours)
16-bit2x MI355X (Vultr) $5.18$3,781
8-bit1x MI355X (Vultr) $2.59$1,891
4-bit1x MI300X (Vultr) $1.85$1,350

One copy of the model on rented cards, busy or idle. Serving more users at once takes more copies or more memory for their contexts.

Memory at Longer Context

Every token in a sequence keeps a key and a value in every layer. From Qwen3-235B-A22B-Thinking-2507's published configuration, that cache adds this much at 16-bit for one sequence:

ContextKey/value cacheTotal with weightsCheapest setup
4,096 tokens0.8 GB565 GB 2x MI355X $5.18/hr
32,768 tokens6.3 GB571 GB 2x MI355X $5.18/hr
131,072 tokens25.2 GB589 GB 3x MI325X $6.00/hr
262,144 tokens (full)50.5 GB615 GB 3x MI325X $6.00/hr

An estimate from layers, key/value heads and head size, assuming full attention in every layer. Runtimes that quantize or page the cache use less.

Questions

How much VRAM does Qwen3-235B-A22B-Thinking-2507 need?

About 564 GB at 16-bit; about 282 GB at 8-bit; about 141 GB at 4-bit: the weights plus 20% for the runtime and a short context. A long context needs more.

What is the cheapest GPU setup to run Qwen3-235B-A22B-Thinking-2507?

At 16-bit, 2x MI355X from $5.18 an hour, at the lowest on-demand price the SAVRN Index lists.

Can Qwen3-235B-A22B-Thinking-2507 run on a single H100?

No. An H100 has 80 GB, and Qwen3-235B-A22B-Thinking-2507 needs 141 GB even at 4-bit, so it takes more than one card.

How much memory does Qwen3-235B-A22B-Thinking-2507 need at its full context length?

About 615 GB at 16-bit for one 262,144-token sequence: 564 GB for the weights and margin plus 50.5 GB of key/value cache, estimated from its published configuration.

Memory is the weights at that precision plus 20% for the runtime and a short context. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026. Setups beyond eight cards, one server, are not listed.