Model · Image and text to text
Qwen
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…
Open weights
apache-2.0
9.7B parameters
262,144 tokens
transformers
This repo quantizes the model using data-free quantization technique. As of 2026-02-25, make sure your system has cuda12.8 installed. Then, create a fresh Python environment (e.g. python3.12 venv) and run: Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after…
Open weights
apache-2.0
9.7B parameters
262,144 tokens
transformers
This repository contains the completed training step25 HF export. The repository name retains step24 for compatibility with the requested upload destination. Optimizer and scheduler state are not included.
Open weights
9.7B parameters
262,144 tokens
transformers
Full Qwen3.5-9B model from the September 22, 2026 Harvey notes-only training runs. This revision is epoch 2, step 1550 of a two-epoch run, job 1016164. The final checkpoint is on main and final; epoch 1 is on epoch1. Each saved checkpoint is also available by its checkpoint- reference. - 100,000,076 note labels/epoch before causal shift; 99,993,882 scored labels/epoch after shift. - 223,681 notes, packed into 6,200 rows of 16,384 tokens with seed 731. - 199,987,764 scored supervised labels seen through this checkpoint. - No supervised trajectory examples. These are the notes-only controls, separate from notes + trajectory mixture models. - Eight GH200 GPUs, effective batch 8, learning rate…
Open weights
apache-2.0
9.7B parameters
262,144 tokens
transformers
Full Qwen3.5-9B model from the September 22, 2026 Harvey notes-only training runs. This revision is epoch 2, step 1550 of a two-epoch run, job 1016041. The final checkpoint is on main and final; epoch 1 is on epoch1. Each saved checkpoint is also available by its checkpoint- reference. - 100,000,076 note labels/epoch before causal shift; 99,993,882 scored labels/epoch after shift. - 223,681 notes, packed into 6,200 rows of 16,384 tokens with seed 731. - 199,987,764 scored supervised labels seen through this checkpoint. - No supervised trajectory examples. These are the notes-only controls, separate from notes + trajectory mixture models. - Eight GH200 GPUs, effective batch 8, learning rate…
Open weights
apache-2.0
9.7B parameters
262,144 tokens
transformers
Full Qwen3.5-9B model from the September 22, 2026 Harvey notes-only training runs. This revision is epoch 2, step 466 of a two-epoch run, job 1016040. The final checkpoint is on main and final; epoch 1 is on epoch1. Each saved checkpoint is also available by its checkpoint- reference. - 29,999,869 note labels/epoch before causal shift; 29,998,010 scored labels/epoch after shift. - 68,799 notes, packed into 1,864 rows of 16,384 tokens with seed 731. - 59,996,020 scored supervised labels seen through this checkpoint. - No supervised trajectory examples. These are the notes-only controls, separate from notes + trajectory mixture models. - Eight GH200 GPUs, effective batch 8, learning rate…
Open weights
apache-2.0
9.7B parameters
262,144 tokens
transformers