# DeepSeek-R1-Distill-Qwen-1.5B by DeepSeek: Open-Weight Model
Source: https://savrn.com/models/deepseek-r1-distill-qwen-1-5b
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve DeepSeek-R1-Distill-Qwen-1.5B (1.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 3.6 GB | 4.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 1.8 GB | 2.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.9 GB | 1.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[DeepSeek-R1-Distill-Qwen-1.5B on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b/gpus)

## Model Card

By DeepSeek, published under mit, revision ad9f0ae0864d.

### DeepSeek-R1

[Paper Link](https://github.com/deepseek-ai/DeepSeek-R1/blob/main/DeepSeek_R1.pdf)

### 1. Introduction

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across math, code, and reasoning tasks. To support the research community, we have open-sourced DeepSeek-R1-Zero, DeepSeek-R1, and six dense models distilled from DeepSeek-R1 based on Llama and Qwen. DeepSeek-R1-Distill-Qwen-32B outperforms OpenAI-o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.

NOTE: Before running DeepSeek-R1 series models locally, we kindly recommend reviewing the Usage Recommendation section.

### 2. Model Summary

Post-Training: Large-Scale Reinforcement Learning on the Base Model

[Read the full model card (1,630 words)](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b/card)

## Configuration

Architecture

Qwen2ForCausalLM

Context length (tokens)

131,072

Layers

28

Hidden size

1,536

Feed-forward size

8,960

Attention heads

12

Key/value heads

2

Vocabulary size

151,936

Sliding window (tokens)

4,096

RoPE base

10,000

Stored precision

bfloat16

Model type

qwen2

## Identity and Version

Repository

deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B

Publisher

DeepSeek

Task

Text generation

Modality

Text

Library

transformers

Parameters

1.8B parameters

Languages

Not stated by the source

Revision

ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562

First published

2025-01-20

Last updated

2025-02-24

## Files and Weights

9 files, 3.6 GB in total. The weights are 1 file totalling 3.6 GB in safetensors.

Weights1 file · 3.6 GB

Configuration2 files · 860 B

Tokenizer2 files · 7.0 MB

Documentation2 files · 17.1 KB

Other1 file · 777.3 KB

Repository1 file · 1.5 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 3.6 GB | 58858233513d |
| config.json | Configuration | 679 B | — |
| generation_config.json | Configuration | 181 B | — |
| LICENSE | Documentation | 1.1 KB | — |
| README.md | Documentation | 16.0 KB | — |
| figures/benchmark.jpg | Other | 777.3 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 7.0 MB | — |
| tokenizer_config.json | Tokenizer | 3.1 KB | — |

## License and Download

License

mit

Access

Open weights, no gate

Download size

3.6 GB

[Download from DeepSeek](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B)

Released by DeepSeek through its official repository on Hugging Face. [Read the license](https://opensource.org/license/mit).

## Built From

- Described by [arXiv:2501.12948](https://savrn.com/papers/deepseek-r1-incentivizing-reasoning-capability-in-llms-via-reinforcement-learning)

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 3.6 GB |
| 16-bit | 3.6 GB |
| 8-bit | 1.8 GB |
| 4-bit | 0.9 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Built on This Model

- Derived from[When2Think-1.5B](https://savrn.com/models/when2think-1-5b)
- Derived from[When2Think-ThinkOnly-1.5B](https://savrn.com/models/when2think-thinkonly-1-5b)
- Adapter of[vigyan-1.5b-adapter-mathematics](https://savrn.com/models/vigyan-1-5b-adapter-mathematics)
- Derived from[vigyan-1.5b-adapter-mathematics](https://savrn.com/models/vigyan-1-5b-adapter-mathematics)
- Adapter of[vigyan-1.5b-adapter-engineering](https://savrn.com/models/vigyan-1-5b-adapter-engineering)
- Derived from[vigyan-1.5b-adapter-engineering](https://savrn.com/models/vigyan-1-5b-adapter-engineering)
- Adapter of[vigyan-1.5b-adapter-technology](https://savrn.com/models/vigyan-1-5b-adapter-technology)
- Derived from[vigyan-1.5b-adapter-technology](https://savrn.com/models/vigyan-1-5b-adapter-technology)
- Adapter of[vigyan-1.5b-adapter-science](https://savrn.com/models/vigyan-1-5b-adapter-science)
- Derived from[vigyan-1.5b-adapter-science](https://savrn.com/models/vigyan-1-5b-adapter-science)
- Quantized from[Vigyan-1.5B-4x-MoE](https://savrn.com/models/vigyan-1-5b-4x-moe)
- Derived from[Vigyan-1.5B-4x-MoE](https://savrn.com/models/vigyan-1-5b-4x-moe)
- Derived from[hypernet-sp-distill](https://savrn.com/models/hypernet-sp-distill)
- Quantized from[hypernet-sp-distill](https://savrn.com/models/hypernet-sp-distill)

## Questions About DeepSeek-R1-Distill-Qwen-1.5B

### How much GPU memory does DeepSeek-R1-Distill-Qwen-1.5B need?

About 4.3 GB at 16-bit and 1.1 GB at 4-bit: the weights (1.8B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run DeepSeek-R1-Distill-Qwen-1.5B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use DeepSeek-R1-Distill-Qwen-1.5B commercially?

Yes. DeepSeek-R1-Distill-Qwen-1.5B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

### What is DeepSeek-R1-Distill-Qwen-1.5B's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [JiRackUltra_1b](https://savrn.com/models/jirackultra-1b)

[Center Business Solutions inc](https://savrn.com/model-publishers/cmsmanhattan)

A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer - We are working to…

Open weights mit 1.8B parameters 131,072 tokens

[View model](https://savrn.com/models/jirackultra-1b)

Model · Text generation

### [When2Think-1.5B](https://savrn.com/models/when2think-1-5b)

[Jaejun Shim](https://savrn.com/model-publishers/junshim)

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…

Open weights mit 1.8B parameters 131,072 tokens transformers

[View model](https://savrn.com/models/when2think-1-5b)

Model · Text generation

### [When2Think-ThinkOnly-1.5B](https://savrn.com/models/when2think-thinkonly-1-5b)

[Jaejun Shim](https://savrn.com/model-publishers/junshim)

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…

Open weights mit 1.8B parameters 131,072 tokens transformers

[View model](https://savrn.com/models/when2think-thinkonly-1-5b)

Model · Text generation

### [ovd-math-1-data-full-step500-historical](https://savrn.com/models/ovd-math-1-data-full-step500-historical)

[Jing Xiong](https://savrn.com/model-publishers/menik1126)

Inference weights and tokenizer for the historical 1-data experiment. Part of the OVD collection.

Open weights 1.8B parameters 4,096 tokens transformers

[View model](https://savrn.com/models/ovd-math-1-data-full-step500-historical)

Model · Text generation

### [ovd-math-1-data-full-step300-historical](https://savrn.com/models/ovd-math-1-data-full-step300-historical)

[Jing Xiong](https://savrn.com/model-publishers/menik1126)

Inference weights and tokenizer for the historical 1-data experiment. Part of the OVD collection.

Open weights 1.8B parameters 4,096 tokens transformers

[View model](https://savrn.com/models/ovd-math-1-data-full-step300-historical)

Model · Text generation

### [Bonsai-27B-mlx-1bit](https://savrn.com/models/bonsai-27b-mlx-1bit)

[Prism ML](https://savrn.com/model-publishers/prism-ml)

Full 27B-class reasoning in binary transformer weights — the first 27B-class model to run on a phone - ~3.9 GB deployed footprint (down from ~54 GB FP16) — fits within the per-app memory budget of a high-end phone such as the iPhone 17 Pro Max - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 - End-to-end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the…

Open weights apache-2.0 1.7B parameters 262,144 tokens mlx

[View model](https://savrn.com/models/bonsai-27b-mlx-1bit)

## DeepSeek

[All models and datasets](https://savrn.com/model-publishers/deepseek-ai)

## Versions

- [ad9f0ae0864d](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b/versions/ad9f0ae0864d) · current 2026-10-07

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under mit](https://savrn.com/models/licenses/mit)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-06.
- [DeepSeek website](https://www.deepseek.com/)
- [Hugging Face record](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B)
- [How the hub is built](https://savrn.com/model-hub/methodology)
