When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…
Open weights
mit
1.8B parameters
131,072 tokens
transformers
When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem. The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource. When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. The checkpoint learns two coupled decisions: 1. Whether to reason 2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target. These…
Open weights
mit
1.8B parameters
131,072 tokens
transformers