SAVRN
Search Contact SAVRN

SAVRN Model Hub · less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50

less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50 GPU Requirements

less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50 needs about 143 GB of GPU memory at 16-bit, 71.4 GB at 8-bit and 35.7 GB at 4-bit: its 59.5B parameters plus a 20% working margin. The cheapest setup at 16-bit is 1x MI300X from $1.85 an hour, about $1,350 a month around the clock. 7 of the 9 accelerators the SAVRN Index prices hold it on one card at 16-bit: B200, GB200 NVL72, B300, GB300 NVL72, MI300X, MI325X, MI355X.

Parameters59.5B
Memory at 16-bit143 GB
Memory at 8-bit71.4 GB
Memory at 4-bit35.7 GB
Context131,072

Every Accelerator, Every Precision

How many cards of each accelerator the SAVRN Index prices it takes to hold less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50, and what that many cards cost an hour at the lowest listed on-demand price. Memory needed: 143 GB at 16-bit, 71.4 GB at 8-bit, 35.7 GB at 4-bit.

AcceleratorMemory per cardLowest price per card16-bit8-bit4-bit
H100
Voltage Park
80 GB $1.99 2 cards
$3.98/hr
1 card
$1.99/hr
1 card
$1.99/hr
H200
GMI Cloud
141 GB $2.60 2 cards
$5.20/hr
1 card
$2.60/hr
1 card
$2.60/hr
B200
Vultr
180 GB $3.50 1 card
$3.50/hr
1 card
$3.50/hr
1 card
$3.50/hr
GB200 NVL72
GMI Cloud
186 GB $8.00 1 card
$8.00/hr
1 card
$8.00/hr
1 card
$8.00/hr
MI300X
Vultr
192 GB $1.85 1 card
$1.85/hr
1 card
$1.85/hr
1 card
$1.85/hr
MI325X
Vultr
256 GB $2.00 1 card
$2.00/hr
1 card
$2.00/hr
1 card
$2.00/hr
B300
Massed Compute
268 GB $6.60 1 card
$6.60/hr
1 card
$6.60/hr
1 card
$6.60/hr
GB300 NVL72
Verda
279 GB $9.06 1 card
$9.06/hr
1 card
$9.06/hr
1 card
$9.06/hr
MI355X
Vultr
288 GB $2.59 1 card
$2.59/hr
1 card
$2.59/hr
1 card
$2.59/hr

Running It Around the Clock

PrecisionCheapest setupPer hourPer month (730 hours)
16-bit1x MI300X (Vultr) $1.85$1,350
8-bit1x MI300X (Vultr) $1.85$1,350
4-bit1x MI300X (Vultr) $1.85$1,350

One copy of the model on rented cards, busy or idle. Serving more users at once takes more copies or more memory for their contexts.

Questions

How much VRAM does less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50 need?

About 143 GB at 16-bit; about 71.4 GB at 8-bit; about 35.7 GB at 4-bit: the weights plus 20% for the runtime and a short context. A long context needs more.

What is the cheapest GPU setup to run less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50?

At 16-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand price the SAVRN Index lists.

Can less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50 run on a single H100?

Yes at 8-bit, 4-bit: an H100 has 80 GB and less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50 needs 71.4 GB at 8-bit.

Memory is the weights at that precision plus 20% for the runtime and a short context. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 20, 2026. Setups beyond eight cards, one server, are not listed.