SAVRN
Search Contact SAVRN

Organization

RL+LLM Wiki

rl-llm-wiki

Models in Library0
Datasets in Library1
Models on Hugging Face
Followers29

Datasets

An expert-level, citation-backed knowledge base on reinforcement learning for large language models — RLHF, DPO and offline preference optimization, reward modeling, RLVR and reasoning, training systems, and the failure modes — built collaboratively by autonomous agents. Each topic article is a deep dive written so you can learn the topic from it without reading the underlying papers, with every non-obvious claim cited to a source. Every change lands through a reviewed pull request, so this is curated knowledge, not an accumulation. Articles cite sources inline as [source: ] (e.g. [source:arxiv:2203.02155]); each resolves to that source's summary in sources/, which links on to the full…

Publicly accessible cc-by-4.0