SAVRN
Search Contact SAVRN

Independent publisher

Will

willamazon1

Models in Library1
Datasets in Library0
Models on Hugging Face36
Followers

Models

Model · Text generation

Qwen3.5-9B-smith-v5-gdpo-exact

Will

Reinforcement-learning checkpoint series from the cposmith... smith-v5-gdpo-exact run: asynchronous multi-turn agentic-environment RL on Qwen/Qwen3.5-9B with an exact (verifiable) reward. The policy was initialized from an internal SFT of Qwen/Qwen3.5-9B (qwen359bsftv3), which also served as the reference model. 92 checkpoints, saved every 4 iterations up to 31, then every 2 iterations, from iter0000003 to iter0000199. Each lives in its own subfolder of this repo so you can compare iter0000003, iter0000007, iter0000011, iter0000015, iter0000019, iter0000023, iter0000027, iter0000031, iter0000033, iter0000035, iter0000037, iter0000039, iter0000041, iter0000043, iter0000045, iter0000047…

Open weights apache-2.0 transformers