SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DeepSeek-R1-Distill-Qwen-1.5B

by DeepSeek deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek-R1-Distill-Qwen-1.5B is an open-weight model for text generation from DeepSeek, released under MIT License. It has 1.8B parameters and a 131,072-token context. At 16-bit it needs about 4.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 1M downloads a month.

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1.

Parameters1.8B
Context131,072
Weights3.6 GB
Licensemit
AccessOpen weights
Monthly Downloads1M

Runs On

What it takes to serve DeepSeek-R1-Distill-Qwen-1.5B (1.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3.6 GB 4.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.8 GB 2.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.9 GB 1.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

DeepSeek-R1-Distill-Qwen-1.5B on every accelerator the SAVRN Index prices, at every precision

Model Card

By DeepSeek, published under mit, revision ad9f0ae0864d.

DeepSeek-R1

Paper Link

1. Introduction

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across math, code, and reasoning tasks. To support the research community, we have open-sourced DeepSeek-R1-Zero, DeepSeek-R1, and six dense models distilled from DeepSeek-R1 based on Llama and Qwen. DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI-o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.

NOTE: Before running DeepSeek-R1 series models locally, we kindly recommend reviewing the Usage Recommendation section.

2. Model Summary

Post-Training: Large-Scale Reinforcement Learning on the Base Model

Read the full model card (1,630 words)

Configuration

Architecture
Qwen2ForCausalLM
Context length (tokens)
131,072
Layers
28
Hidden size
1,536
Feed-forward size
8,960
Attention heads
12
Key/value heads
2
Vocabulary size
151,936
Sliding window (tokens)
4,096
RoPE base
10,000
Stored precision
bfloat16
Model type
qwen2

Identity and Version

Repository
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Publisher
DeepSeek
Task
Text generation
Modality
Text
Library
transformers
Parameters
1.8B parameters
Languages
Not stated by the source
Revision
ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562
First published
2025-01-20
Last updated
2025-02-24

Files and Weights

9 files, 3.6 GB in total. The weights are 1 file totalling 3.6 GB in safetensors.

Weights1 file · 3.6 GB
Configuration2 files · 860 B
Tokenizer2 files · 7.0 MB
Documentation2 files · 17.1 KB
Other1 file · 777.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights3.6 GB 58858233513d
config.jsonConfiguration679 B —
generation_config.jsonConfiguration181 B —
LICENSEDocumentation1.1 KB —
README.mdDocumentation16.0 KB —
figures/benchmark.jpgOther777.3 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer7.0 MB —
tokenizer_config.jsonTokenizer3.1 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
3.6 GB
Download from DeepSeek

Released by DeepSeek through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published3.6 GB
16-bit3.6 GB
8-bit1.8 GB
4-bit0.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About DeepSeek-R1-Distill-Qwen-1.5B

How much GPU memory does DeepSeek-R1-Distill-Qwen-1.5B need?

About 4.3 GB at 16-bit and 1.1 GB at 4-bit: the weights (1.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DeepSeek-R1-Distill-Qwen-1.5B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DeepSeek-R1-Distill-Qwen-1.5B commercially?

Yes. DeepSeek-R1-Distill-Qwen-1.5B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is DeepSeek-R1-Distill-Qwen-1.5B's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer - We are working to…

Open weights mit 1.8B parameters 131,072 tokens

Model · Text generation

When2Think-1.5B

Jaejun Shim

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…

Open weights mit 1.8B parameters 131,072 tokens transformers

Model · Text generation

When2Think-ThinkOnly-1.5B

Jaejun Shim

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…

Open weights mit 1.8B parameters 131,072 tokens transformers

Model · Text generation

Bonsai-27B-mlx-1bit

Prism ML

Full 27B-class reasoning in binary transformer weights — the first 27B-class model to run on a phone - ~3.9 GB deployed footprint (down from ~54 GB FP16) — fits within the per-app memory budget of a high-end phone such as the iPhone 17 Pro Max - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 - End-to-end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the…

Open weights apache-2.0 1.7B parameters 262,144 tokens mlx