# DeepSeek-V4.1-Flash by DeepSeek: GPU Requirements and Cost
Source: https://savrn.com/models/deepseek-v4-1-flash/gpus
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Every Accelerator, Every Precision

How many cards of each accelerator the SAVRN Index prices it takes to hold DeepSeek-V4.1-Flash, and what that many cards cost an hour at the lowest listed on-demand price. Memory needed: 1.8 TB at 16-bit, 916 GB at 8-bit, 458 GB at 4-bit.

| Accelerator | Memory per card | Lowest price per card | 16-bit | 8-bit | 4-bit |
| --- | --- | --- | --- | --- | --- |
| [H100](https://savrn.com/ai-index/pricing/gpus/h100) Voltage Park | 80 GB | $1.99 | More than 8 | More than 8 | 6 cards $11.94/hr |
| [H200](https://savrn.com/ai-index/pricing/gpus/h200) GMI Cloud | 141 GB | $2.60 | More than 8 | 7 cards $18.20/hr | 4 cards $10.40/hr |
| [B200](https://savrn.com/ai-index/pricing/gpus/b200) Vultr | 180 GB | $3.50 | More than 8 | 6 cards $21.00/hr | 3 cards $10.50/hr |
| [GB200 NVL72](https://savrn.com/ai-index/pricing/gpus/gb200) GMI Cloud | 186 GB | $8.00 | More than 8 | 5 cards $40.00/hr | 3 cards $24.00/hr |
| [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) Vultr | 192 GB | $1.85 | More than 8 | 5 cards $9.25/hr | 3 cards $5.55/hr |
| [MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) Vultr | 256 GB | $2.00 | 8 cards $16.00/hr | 4 cards $8.00/hr | 2 cards $4.00/hr |
| [B300](https://savrn.com/ai-index/pricing/gpus/b300) Massed Compute | 268 GB | $6.60 | 7 cards $46.20/hr | 4 cards $26.40/hr | 2 cards $13.20/hr |
| [GB300 NVL72](https://savrn.com/ai-index/pricing/gpus/gb300) Verda | 279 GB | $10.32 | 7 cards $72.24/hr | 4 cards $41.28/hr | 2 cards $20.64/hr |
| [MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) Vultr | 288 GB | $2.59 | 7 cards $18.13/hr | 4 cards $10.36/hr | 2 cards $5.18/hr |

## Running It Around the Clock

| Precision | Cheapest setup | Per hour | Per month (730 hours) |
| --- | --- | --- | --- |
| 16-bit | 8x [MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) (Vultr) | $16.00 | $11,680 |
| 8-bit | 4x [MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) (Vultr) | $8.00 | $5,840 |
| 4-bit | 2x [MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) (Vultr) | $4.00 | $2,920 |

One copy of the model on rented cards, busy or idle. Serving more users at once takes more copies or more memory for their contexts.

## Questions

### How much VRAM does DeepSeek-V4.1-Flash need?

About 1.8 TB at 16-bit; about 916 GB at 8-bit; about 458 GB at 4-bit: the weights plus 20% for the runtime and a short context. A long context needs more.

### What is the cheapest GPU setup to run DeepSeek-V4.1-Flash?

At 16-bit, 8x MI325X from $16.00 an hour, at the lowest on-demand price the SAVRN Index lists.

### Can DeepSeek-V4.1-Flash run on a single H100?

No. An H100 has 80 GB, and DeepSeek-V4.1-Flash needs 458 GB even at 4-bit, so it takes more than one card.

Memory is the weights at that precision plus 20% for the runtime and a short context. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026. Setups beyond eight cards, one server, are not listed.

## DeepSeek-V4.1-Flash

- [Model card, license and download](https://savrn.com/models/deepseek-v4-1-flash)
- [All image and text to text models](https://savrn.com/models/tasks/image-and-text-to-text)
- [GPU prices in the SAVRN Index](https://savrn.com/ai-index/pricing/gpus)
