Model · Reinforcement learning
rloo_qwen3_1.7b_polaris_hybrid8_b32_len32768_thinking_seed79
Training run started on 2026-09-18. Training is in progress; no quality evaluation is claimed. - RLOO, learning rate 1e-6, KL coefficient 0, sigmoid length reward alpha 0.1, seed 79. - Sampling temperature 1.0, top-p 1.0. No validation run. - Synchronous generation/training on the same eight H100 80GB GPUs; ZeRO-2; bfloat16. Every 20 rollout steps, a resumable DeepSpeed checkpoint (model and optimizer state) is uploaded under checkpoints/actor/globalstepN/, together with its training configuration. It is not a standalone Transformers model folder. Restore that directory and a latest file containing its tag to resume with the same OpenRLHF configuration. Local files are removed only after…