Model · Text generation
exaone-7b-verireason-reproduced-1513-fullft-best-grpo-reproduced-1.0
This model is a fine-tuned version of Jongbin-kr/exaone-7b-verireason-reproduced-1513-fullft-epoch4. It has been trained using TRL. This model was trained with GRPO, a method introduced in DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.