SAVRN
Search Contact SAVRN

Organization

Hugging Face H4

HuggingFaceH4

Aligning LLMs to be helpful, honest, harmless, and huggy (H4)

Models in Library0
Datasets in Library2
Models on Hugging Face36
Followers1.6k

Datasets

Dataset · Text generation

MATH-500

Hugging Face H4

This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits

Publicly accessible

Dataset · Text generation

ultrachat_200k

Hugging Face H4

This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model. The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic: - Selection of a subset of data for faster supervised fine tuning. - Truecasing of the dataset, as we observed around 5% of the data contained grammatical errors like "Hello. how are you?" instead of "Hello. How are you?" - Removal of dialogues where the assistant replies with phrases like "I do not have emotions" or "I don't have opinions", even for fact-based prompts that don't involve either. The…

Publicly accessible mit 100K<n<1M