SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Summarization Datasets

2 datasets in the SAVRN Model Hub for summarization, from publishers including LAION eV, JHU Human Language Technology Center of Excellence.

2 datasets.

MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser. The pipeline through which MegaWika was created is complex, and is described in more detail in the paper (linked above), but the…

Publicly accessible cc-by-sa-4.0 10M<n<100M

Dataset · Summarization

Scientific-Summaries

LAION eV

22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has an 18-field structured summary covering methodology, key results, claims, limitations, and more. This public dataset includes full paper text for ~5.3 million papers where open-access status has been confirmed -- either through OpenAlex metadata or because the paper originates from a permissively licensed source such as the arXiv preprint server, bioRxiv, medRxiv, or ChemRxiv. Full text (textsanitized, textraw) is included in this public dataset when either of the following is true: 1. The paper originates from a permissively licensed source: arXiv preprint server…

Publicly accessible cc-by-4.0 10M<n<100M

Who Publishes These Datasets

Other tasks

See all