MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser. The pipeline through which MegaWika was created is complex, and is described in more detail in the paper (linked above), but the…
Publicly accessible
cc-by-sa-4.0
10M<n<100M
22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has an 18-field structured summary covering methodology, key results, claims, limitations, and more. This public dataset includes full paper text for ~5.3 million papers where open-access status has been confirmed -- either through OpenAlex metadata or because the paper originates from a permissively licensed source such as the arXiv preprint server, bioRxiv, medRxiv, or ChemRxiv. Full text (textsanitized, textraw) is included in this public dataset when either of the following is true: 1. The paper originates from a permissively licensed source: arXiv preprint server…
Publicly accessible
cc-by-4.0
10M<n<100M