FP8 version of ThinkLess-2B: 8-bit floating-point weights and activations, 2.5 GB (bf16: 4.3 GB), with near-identical accuracy. Made with llm-compressor (FP8DYNAMIC: per-channel FP8 weights, dynamic per-token FP8 activations, no calibration data). The output head, vision tower and MTP heads stay in 16-bit. The differences are within the 95% confidence intervals, and answers stay just as short (cut-offs ≤ 1%). FP8 compute needs a GPU with FP8 support (NVIDIA Hopper or Ada, e.g. H100, L4, RTX 40-series); vLLM loads the compressed-tensors format directly. Use Qwen3.5's thinking-mode sampling (temperature 1.0, top-p 0.95, top-k 20, presence penalty 1.5). A 4-bit AWQ version of ThinkLess-2B was…
Independent publisher
Sadikh Shaik
Shaik1903
Models
Qwen3.5-2B that thinks less and answers better. ThinkLess-2B uses 57–82% fewer reasoning tokens than the base model on math and science benchmarks while being more accurate on GSM8K, MATH-500 and GPQA-Diamond, with almost no answers cut off mid-thought. Served with vLLM and its built-in MTP speculative decoding, a typical request finishes 3.4× faster than the base model. It was post-trained in two stages: SFT on the model's own shortest correct solutions (self-distillation, with a same-family 9B teacher filling in the hardest problems), then a short GRPO run with an accuracy-and-length reward. All accuracies are at Qwen's recommended 81,920-token thinking budget; see Evaluation. Mean…