SAVRN
Search Contact SAVRN

Open-weight model · Text generation

When2Think-1.5B

by Jaejun Shim junshim/When2Think-1.5B

When2Think-1.5B is an open-weight model for text generation from Jaejun Shim, released under MIT License. It has 1.8B parameters and a 131,072-token context. At 16-bit it needs about 4.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 919 downloads a month.

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem.

Parameters1.8B
Context131,072
Weights7.1 GB
Licensemit
AccessOpen weights
Monthly Downloads919

Runs On

What it takes to serve When2Think-1.5B (1.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3.6 GB 4.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.8 GB 2.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.9 GB 1.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

When2Think-1.5B on every accelerator the SAVRN Index prices, at every precision

Model Card

By Jaejun Shim, published under mit, revision 69547b99750d.

When2Think

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem.

The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource.

Highlights

  • Adaptive Think/NoThink Behavior: Learns when to answer directly and when to invoke explicit multi-step reasoning.
  • Accuracy-Efficiency Trade-off: Reduces unnecessary reasoning without uniformly suppressing useful reasoning on difficult problems.
  • Standalone Deployment: Requires only the released checkpoint for generation.

Model Details

Model Description

When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B.

The checkpoint learns two coupled decisions:

  1. Whether to reason - NOTHINK: Answer directly without an extended explicit reasoning trace. - THINK: Generate explicit multi-step reasoning followed by a final answer.

  2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target.

The post-training framework combines:

Read the full model card (1,072 words)

Configuration

Architecture
Qwen2ForCausalLM
Context length (tokens)
131,072
Layers
28
Hidden size
1,536
Feed-forward size
8,960
Attention heads
12
Key/value heads
2
Vocabulary size
151,936
Model type
qwen2

Identity and Version

Repository
junshim/When2Think-1.5B
Publisher
Jaejun Shim
Task
Text generation
Modality
Text
Library
transformers
Parameters
1.8B parameters
Languages
en
Revision
69547b99750d1cb82d75ea82c3afb48454a75627
First published
2026-09-16
Last updated
2026-10-07

Files and Weights

8 files, 7.1 GB in total. The weights are 1 file totalling 7.1 GB in safetensors.

Weights1 file · 7.1 GB
Configuration2 files · 1.6 KB
Tokenizer2 files · 11.4 MB
Documentation1 file · 12.1 KB
Other1 file · 2.2 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights7.1 GB 43dea75bf87d
config.jsonConfiguration1.4 KB —
generation_config.jsonConfiguration207 B —
README.mdDocumentation12.1 KB —
chat_template.jinjaOther2.2 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer11.4 MB f624f8136bc5
tokenizer_config.jsonTokenizer421 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
7.1 GB
Download from Jaejun Shim

Released by Jaejun Shim through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published7.1 GB
16-bit3.6 GB
8-bit1.8 GB
4-bit0.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About When2Think-1.5B

How much GPU memory does When2Think-1.5B need?

About 4.3 GB at 16-bit and 1.1 GB at 4-bit: the weights (1.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run When2Think-1.5B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use When2Think-1.5B commercially?

Yes. When2Think-1.5B is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is When2Think-1.5B's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer - We are working to…

Open weights mit 1.8B parameters 131,072 tokens

Model · Text generation

DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across…

Open weights mit 1.8B parameters 131,072 tokens transformers

Model · Text generation

When2Think-ThinkOnly-1.5B

Jaejun Shim

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…

Open weights mit 1.8B parameters 131,072 tokens transformers

Model · Text generation

Bonsai-27B-mlx-1bit

Prism ML

Full 27B-class reasoning in binary transformer weights — the first 27B-class model to run on a phone - ~3.9 GB deployed footprint (down from ~54 GB FP16) — fits within the per-app memory budget of a high-end phone such as the iPhone 17 Pro Max - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 - End-to-end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the…

Open weights apache-2.0 1.7B parameters 262,144 tokens mlx