SAVRN
Search Contact SAVRN

When2Think-1.5B · Model Card

When2Think-1.5B: Model Card

Written by Jaejun Shim, published under mit, revision 69547b99750d, read 2026-10-07. Shown as written; SAVRN's own facts about this model are on its page.

When2Think

When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem.

The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource.

Highlights

  • Adaptive Think/NoThink Behavior: Learns when to answer directly and when to invoke explicit multi-step reasoning.
  • Accuracy-Efficiency Trade-off: Reduces unnecessary reasoning without uniformly suppressing useful reasoning on difficult problems.
  • Standalone Deployment: Requires only the released checkpoint for generation.

Model Details

Model Description

When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B.

The checkpoint learns two coupled decisions:

  1. Whether to reason - NOTHINK: Answer directly without an extended explicit reasoning trace. - THINK: Generate explicit multi-step reasoning followed by a final answer.

  2. How much to reason - Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target.

The post-training framework combines:

  • Instance-level Difficulty-Aware Control (IDAC): Uses cached reference success and token-cost statistics to modulate a correctness-gated efficiency bonus based on generated token count.
  • Batch-Wise Standardization (BWS): Converts trajectory rewards into standardized advantages for critic-free policy optimization.
  • Importance-sampled THINK/NOTHINK exploration: Adopts balanced mode exploration during post-training. The auxiliary exploration policy is removed at inference time.

These mechanisms are training-time components only. The released checkpoint runs as a standalone causal language model.

Which Model Should I Use?

Model Whether to reason How much to reason Recommended use
When2Think-1.5B Learned THINK/NOTHINK selection IDAC + BWS Adaptive hybrid reasoning in a single checkpoint
When2Think-ThinkOnly-1.5B Always THINK, without hybrid importance sampling IDAC + BWS Explicit reasoning or analysis of within-THINK depth control

The checkpoints represent different accuracy-computation operating points rather than a strict performance ordering.

Model Sources

Uses

Direct Use

When2Think-1.5B is intended for:

  • Mathematical problem solving
  • Adaptive direct-answer and explicit-reasoning generation
  • Research on efficient reasoning models
  • Analysis of Think/NoThink mode selection

The model can be loaded as a standard causal language model using Hugging Face Transformers. No separate router, verifier, critic, difficulty estimator, reward model, or reference policy is required for inference.  

Downstream Use

The checkpoint may be used as a starting point for:

  • Continued post-training on verifiable reasoning tasks
  • Adaptation to mathematical, symbolic, coding, or scientific reasoning
  • Research on adaptive reasoning policies
  • Studies of direct-answer and explicit-reasoning behavior   Additional fine-tuning may alter the learned Think/NoThink balance, response length, and reasoning-depth behavior.

How to Get Started with the Model

Use the code below to get started with the model.

pip install -U torch transformers accelerate
# optional, for fast serving
pip install vllm

by Transformers Pipeline

from transformers import pipeline

model_path = "junshim/When2Think-1.5B"

prompt = "Find the value of $x$ that satisfies the equation $4x+5 = 6x+7$."
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": prompt}
]

generator = pipeline(
    "text-generation",
    model=model_path,
    device_map="auto",
    dtype="auto"
)
outputs = generator(
    messages,
    max_new_tokens=512,
    clean_up_tokenization_spaces=False
)

by Transformers

from accelerate import Accelerator
from transformers import AutoModelForCausalLM, AutoTokenizer

accelerator = Accelerator()
device = accelerator.device

model = AutoModelForCausalLM.from_pretrained(
    model_path,
    torch_dtype="auto",
    device_map=device
)
tokenizer = AutoTokenizer.from_pretrained(model_path)

print(model.generation_config, device)
inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

input_length = inputs["input_ids"].shape[1]
max_new_tokens = model.config.max_position_embeddings - input_length

with torch.inference_mode():
    outputs = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens
    )

output_text = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]

vLLM

vllm serve "junshim/When2Think-1.5B" --reasoning-parser deepseek_r1

Output Parsing

For decoded outputs that use <think>...</think>, the following helper separates the reasoning trace from the final response. Treat this as a convenience parser for that output format, not as a format guarantee across all serving stacks.

Evaluation

Testing Data, Factors & Metrics

See the paper for dataset versions, prompts, answer extraction, and the complete evaluation protocol.

Testing Data

The reported evaluation covers verifiable mathematical reasoning and cross-domain transfer:

  • GSM-Plus
  • OlympiadBench-Math
  • AIME24
  • AIME25
  • Minerva
  • MATH-500
  • Stratified MMLU-Pro transfer
Factors

The analysis considers:

  • Benchmark and task domain
  • Problem difficulty where difficulty labels are available
  • THINK/NOTHINK selection behavior
  • Generated token usage
  • Model scale and training variant
Metrics
  • Pass@3: Sampling-based Pass@k accuracy with k = 3, following the protocol in the paper.
  • Pass@1: Used for difficulty-stratified MATH-500 analysis.
  • Average generated tokens: Mean generated tokens per response, reported independently of the Pass@k sampling budget.
  • THINK ratio: Fraction of responses using explicit reasoning.

Key Results

Model MATH-500 L1 Pass@1 ↑ Tokens ↓ GSM-Plus Pass@3 ↑ Tokens ↓ AIME24 Pass@3 ↑ Tokens ↓ AIME25 Pass@3 ↑ Tokens ↓
R1-Distill-Qwen backbone 92.1 1,199 79.4 590 46.0 14,195 32.0 12,616
When2Think-1.5B 95.8 619 85.7 1,052 56.0 10,236 40.0 9,549
When2Think-ThinkOnly-1.5B N/R N/R 86.8 1,652 57.3 10,046 40.0 9,846

When2Think-ThinkOnly-1.5B corresponds to the paper's w/o IS (IDAC + BWS) variant. Its MATH-500 Level 1 results are not separately reported in the paper and are therefore marked as N/R.

Summary

Key observations:

  • Easy-instance allocation: On MATH-500 Level 1, When2Think improves Pass@1 from 92.1% to 95.8% while reducing average generated tokens from 1,199 to 619, a reduction of 580 tokens, or approximately 48.4%. This demonstrates that the hybrid checkpoint can favor direct answering when extended deliberation provides limited benefit.
  • Task-dependent allocation: On GSM-Plus, When2Think uses more tokens than the backbone while improving Pass@3 from 79.4% to 85.7%. Together with the MATH-500 Level 1 result, this shows that adaptive allocation may reduce, preserve, or increase computation depending on the input.
  • When2Think-1.5B on AIME24: Pass@3 improves by 10.0 percentage points, while average generated tokens decrease by 27.9% relative to the backbone.
  • When2Think-1.5B on AIME25: Pass@3 improves by 8.0 percentage points, while average generated tokens decrease by 24.3% relative to the backbone.
  • When2Think-ThinkOnly-1.5B: Achieves 57.3% Pass@3 on AIME24 and 40.0% Pass@3 on AIME25 while always using explicit reasoning.
  • Within-THINK control: The strong THINK-only results show that difficulty-aware computation control contributes independently of THINK/NOTHINK mode selection.
  • Different operating points: The hybrid and THINK-only checkpoints represent different accuracy-computation operating points rather than a strict performance ordering.

For complete baselines, standard deviations, difficulty-stratified analysis, and ablations, see the paper.

Citation

BibTeX:

@misc{shim2026when2think,
  title         = {When2Think: Learning When and How Much to Reason},
  author        = {Jaejun Shim and HyunJin Kim and Young Jin Kim and JinYeong Bak},
  year          = {2026},
  eprint        = {2609.19671},
  url           = {https://arxiv.org/abs/2609.19671}
}

More Information

  • Begin of Sentence: <|begin▁of▁sentence|>
  • End of Sentence: <|end▁of▁sentence|>
  • Pad: <|end▁of▁sentence|>
  • Begin of Thinking:
  • End of Thinking:
  • User Role: <|User|>
  • Assistant: <|Assistant|>
  • Max position embeddings (length): 131072

Model Card Contact

For questions, open an issue in JJunShim/When2Think.