# When2Think-1.5B by Jaejun Shim: Open-Weight Model
Source: https://savrn.com/models/when2think-1-5b
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve When2Think-1.5B (1.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 3.6 GB | 4.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 1.8 GB | 2.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.9 GB | 1.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[When2Think-1.5B on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/when2think-1-5b/gpus)

## Model Card

By Jaejun Shim, published under mit, revision 69547b99750d.

### When2Think

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem.

The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource.

### Highlights

- Adaptive Think/NoThink Behavior: Learns when to answer directly and when to invoke explicit multi-step reasoning.
- Accuracy-Efficiency Trade-off: Reduces unnecessary reasoning without uniformly suppressing useful reasoning on difficult problems.
- Standalone Deployment: Requires only the released checkpoint for generation.

### Model Details

#### Model Description

When2Think-1.5B is an RLVR-post-trained version of [deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b).

The checkpoint learns two coupled decisions:

1. Whether to reason - NOTHINK: Answer directly without an extended explicit reasoning trace. - THINK: Generate explicit multi-step reasoning followed by a final answer.
2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target.

The post-training framework combines:

[Read the full model card (1,072 words)](https://savrn.com/models/when2think-1-5b/card)

## Configuration

Architecture

Qwen2ForCausalLM

Context length (tokens)

131,072

Layers

28

Hidden size

1,536

Feed-forward size

8,960

Attention heads

12

Key/value heads

2

Vocabulary size

151,936

Model type

qwen2

## Identity and Version

Repository

junshim/When2Think-1.5B

Publisher

Jaejun Shim

Task

Text generation

Modality

Text

Library

transformers

Parameters

1.8B parameters

Languages

en

Revision

69547b99750d1cb82d75ea82c3afb48454a75627

First published

2026-09-16

Last updated

2026-10-07

## Files and Weights

8 files, 7.1 GB in total. The weights are 1 file totalling 7.1 GB in safetensors.

Weights1 file · 7.1 GB

Configuration2 files · 1.6 KB

Tokenizer2 files · 11.4 MB

Documentation1 file · 12.1 KB

Other1 file · 2.2 KB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 7.1 GB | 43dea75bf87d |
| config.json | Configuration | 1.4 KB | — |
| generation_config.json | Configuration | 207 B | — |
| README.md | Documentation | 12.1 KB | — |
| chat_template.jinja | Other | 2.2 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 11.4 MB | f624f8136bc5 |
| tokenizer_config.json | Tokenizer | 421 B | — |

## License and Download

License

mit

Access

Open weights, no gate

Download size

7.1 GB

[Download from Jaejun Shim](https://huggingface.co/junshim/When2Think-1.5B)

Released by Jaejun Shim through its official repository on Hugging Face. [Read the license](https://opensource.org/license/mit).

## Built From

- Derived from [deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b)
- Described by arXiv:2609.19671
- Trained on (disclosed) agentica-org/DeepScaleR-Preview-Dataset

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 7.1 GB |
| 16-bit | 3.6 GB |
| 8-bit | 1.8 GB |
| 4-bit | 0.9 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About When2Think-1.5B

### How much GPU memory does When2Think-1.5B need?

About 4.3 GB at 16-bit and 1.1 GB at 4-bit: the weights (1.8B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run When2Think-1.5B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use When2Think-1.5B commercially?

Yes. When2Think-1.5B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

### What is When2Think-1.5B's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [JiRackUltra_1b](https://savrn.com/models/jirackultra-1b)

[Center Business Solutions inc](https://savrn.com/model-publishers/cmsmanhattan)

A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer - We are working to…

Open weights mit 1.8B parameters 131,072 tokens

[View model](https://savrn.com/models/jirackultra-1b)

Model · Text generation

### [DeepSeek-R1-Distill-Qwen-1.5B](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b)

[DeepSeek](https://savrn.com/model-publishers/deepseek-ai)

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across…

Open weights mit 1.8B parameters 131,072 tokens transformers

[View model](https://savrn.com/models/deepseek-r1-distill-qwen-1-5b)

Model · Text generation

### [When2Think-ThinkOnly-1.5B](https://savrn.com/models/when2think-thinkonly-1-5b)

[Jaejun Shim](https://savrn.com/model-publishers/junshim)

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…

Open weights mit 1.8B parameters 131,072 tokens transformers

[View model](https://savrn.com/models/when2think-thinkonly-1-5b)

Model · Text generation

### [ovd-math-1-data-full-step500-historical](https://savrn.com/models/ovd-math-1-data-full-step500-historical)

[Jing Xiong](https://savrn.com/model-publishers/menik1126)

Inference weights and tokenizer for the historical 1-data experiment. Part of the OVD collection.

Open weights 1.8B parameters 4,096 tokens transformers

[View model](https://savrn.com/models/ovd-math-1-data-full-step500-historical)

Model · Text generation

### [ovd-math-1-data-full-step300-historical](https://savrn.com/models/ovd-math-1-data-full-step300-historical)

[Jing Xiong](https://savrn.com/model-publishers/menik1126)

Inference weights and tokenizer for the historical 1-data experiment. Part of the OVD collection.

Open weights 1.8B parameters 4,096 tokens transformers

[View model](https://savrn.com/models/ovd-math-1-data-full-step300-historical)

Model · Text generation

### [Bonsai-27B-mlx-1bit](https://savrn.com/models/bonsai-27b-mlx-1bit)

[Prism ML](https://savrn.com/model-publishers/prism-ml)

Full 27B-class reasoning in binary transformer weights — the first 27B-class model to run on a phone - ~3.9 GB deployed footprint (down from ~54 GB FP16) — fits within the per-app memory budget of a high-end phone such as the iPhone 17 Pro Max - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 - End-to-end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the…

Open weights apache-2.0 1.7B parameters 262,144 tokens mlx

[View model](https://savrn.com/models/bonsai-27b-mlx-1bit)

## Jaejun Shim

[All models and datasets](https://savrn.com/model-publishers/junshim)

## Versions

- [69547b99750d](https://savrn.com/models/when2think-1-5b/versions/69547b99750d) · current 2026-10-07

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under mit](https://savrn.com/models/licenses/mit)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-07.
- [Hugging Face record](https://huggingface.co/junshim/When2Think-1.5B)
- [How the hub is built](https://savrn.com/model-hub/methodology)
