SAVRN
Search Contact SAVRN

Independent publisher

Minjae Oh

Riasok

Models in Library2
Datasets in Library0
Models on Hugging Face52
Followers

Models

This model is a fine-tuned version of cosmos1030/gmp-kd3e-1-s80pct-lr1e-420260916220740 on the trl-lib/ultrafeedbackbinarized dataset. It has been trained using TRL. This model was trained with DPO, a method introduced in Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Open weights 4B parameters 40,960 tokens transformers