The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Rows3,708,608
Configurations4
Size643.7 MB
Licensecc-by-sa-3.0
AccessPublicly accessible
Monthly Downloads1.8M
Dataset Card
By Salesforce AI Research, published under cc-by-sa-3.0, revision b08601e04326.
Dataset Card for "wikitext"
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/
- Repository: More Information Needed
- Paper: Pointer Sentinel Mixture Models
- Point of Contact: Stephen Merity
- Size of downloaded dataset files: 391.41 MB
- Size of the generated dataset: 1.12 GB
- Total amount of disk used: 1.52 GB
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.
Structure
wikitext-103-raw-v1 1,809,468 rows
| Split | Rows | Size |
|---|---|---|
| test | 4,358 | 1.4 MB |
| train | 1,801,350 | 525.9 MB |
| validation | 3,760 | 943.9 KB |
textstring
wikitext-103-v1 1,809,468 rows
| Split | Rows | Size |
|---|---|---|
| test | 4,358 | 1.4 MB |
| train | 1,801,350 | 524.6 MB |
| validation | 3,760 | 941.1 KB |
textstring
wikitext-2-raw-v1 44,836 rows
| Split | Rows | Size |
|---|---|---|
| test | 4,358 | 1.4 MB |
| train | 36,718 | 10.7 MB |
| validation | 3,760 | 943.9 KB |
textstring
wikitext-2-v1 44,836 rows
| Split | Rows | Size |
|---|---|---|
| test | 4,358 | 1.3 MB |
| train | 36,718 | 10.6 MB |
| validation | 3,760 | 924.3 KB |
textstring
Details
- Repository
- Salesforce/wikitext
- Publisher
- Salesforce AI Research
- Task category
- Text generation
- Tags
- Not stated by the source
- Size category
- 1M<n<10M
- Languages
- en
- Revision
- b08601e04326c79dfdd32d625aee71d232d685c3
- Last updated
- 2024-01-04
Files
16 files, 643.7 MB in total.
Data14 files · 643.7 MB
Documentation1 file · 10.5 KB
Repository1 file · 1.2 KB
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| wikitext-103-raw-v1/test-00000-of-00001.parquet | Data | 732.6 KB | 5f1bea067869 |
| wikitext-103-raw-v1/train-00000-of-00002.parquet | Data | 157.0 MB | 74da360f2382 |
| wikitext-103-raw-v1/train-00001-of-00002.parquet | Data | 157.1 MB | ba090ac30dbf |
| wikitext-103-raw-v1/validation-00000-of-00001.parquet | Data | 657.2 KB | 204929b7ff9d |
| wikitext-103-v1/test-00000-of-00001.parquet | Data | 721.7 KB | abdfc9f83b11 |
| wikitext-103-v1/train-00000-of-00002.parquet | Data | 155.8 MB | c2ecca8c3250 |
| wikitext-103-v1/train-00001-of-00002.parquet | Data | 155.9 MB | 720f2503551f |
| wikitext-103-v1/validation-00000-of-00001.parquet | Data | 655.1 KB | a586125adab0 |
| wikitext-2-raw-v1/test-00000-of-00001.parquet | Data | 732.6 KB | 5f1bea067869 |
| wikitext-2-raw-v1/train-00000-of-00001.parquet | Data | 6.4 MB | e83889baabc4 |
| wikitext-2-raw-v1/validation-00000-of-00001.parquet | Data | 657.2 KB | 204929b7ff9d |
| wikitext-2-v1/test-00000-of-00001.parquet | Data | 685.4 KB | e6b3913da714 |
| wikitext-2-v1/train-00000-of-00001.parquet | Data | 6.1 MB | dfc27e4360c6 |
| wikitext-2-v1/validation-00000-of-00001.parquet | Data | 617.7 KB | 717de9a0c1c0 |
| README.md | Documentation | 10.5 KB | — |
| .gitattributes | Repository | 1.2 KB | — |
License and Download
- License
- cc-by-sa-3.0
- Access
- No access gate
Download from Salesforce AI Research
Released by Salesforce AI Research through its official repository on Hugging Face.