SAVRN
Search Contact SAVRN

Dataset

minimal-en-corpus-5b

by Fabryka AI SlayerLab/minimal-en-corpus-5b

An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT. The name refers to the approximately 5B-token mixture selected before final BPE tokenization.

Rows
Configurations
Size17.0 GB
License
AccessPublicly accessible
Monthly Downloads254

Dataset Card

An English-language pretraining corpus prepared for controlled experiments with approximately 125M-parameter GPT-2 models based on karpathy/nanoGPT. The name refers to the approximately 5B-token mixture selected before final BPE tokenization. With the included 12,288-token BPE tokenizer, the packaged nanoGPT training split contains 5,396,605,407 tokens. The dataset provides both reusable source text and ready-to-train nanoGPT binaries. - documentid: string - sourceid: string - text: largestring - 12,029 merges - 256 byte tokens - 3 reserved special tokens -: ID 12,285 -: ID 12,286 -: ID 12,287 - One separator appended after every packaged document The tokenizer was trained on a…

Excerpt from the card by Fabryka AI.

Details

Repository
SlayerLab/minimal-en-corpus-5b
Publisher
Fabryka AI
Task category
Not stated by the source
Tags
pretraining, nanogpt, gpt-2
Size category
1M<n<10M
Languages
en
Revision
8624adc77201c4ab022146800fed85128d93b5ed
Last updated
2026-09-22

Files

34 files, 17.0 GB in total.

Data17 files · 6.2 GB
Tokenizer2 files · 831.2 KB
Documentation2 files · 8.3 KB
Other12 files · 10.8 GB
Repository1 file · 2.6 KB
Every file
FileTypeSizeSHA-256
data/manifest.jsonData1.5 KB
manifests/final-splits.jsonData871 B
manifests/mixture.jsonData1.9 KB
manifests/parquet-export.jsonData2.5 KB
manifests/qa.jsonData2.8 KB
runmirror/board_6b.jsonData16.6 KB
runmirror/board_bpe16m_c.jsonData16.6 KB
tokenizer/nanogpt-12k.jsonData559.3 KB
train/shard-00000.parquetData760.6 MBc520f922b7c1
train/shard-00001.parquetData925.0 MB307b3116c943
train/shard-00002.parquetData867.8 MB036866429cde
train/shard-00003.parquetData1.8 GBbbeb75d3388a
train/shard-00004.parquetData383.3 MB601f12697ccd
train/shard-00005.parquetData523.7 MB76df0d3ccc6c
train/shard-00006.parquetData491.6 MB49f6128c1d8e
train/shard-00007.parquetData408.5 MB7ca2a9ca9a93
validation/shard-00000.parquetData6.4 MB5abe9d603867
LICENSEDocumentation728 B
README.mdDocumentation7.6 KB
data/train.binOther10.8 GBcb4c457d2639
data/train.idxOther31.7 MB3c8179598f8e
data/val.binOther10.5 MB8253a10857ca
data/val.idxOther35.8 KB
runmirror/eval_6b.logOther49.3 KB
runmirror/eval_c.logOther68.4 KB
runmirror/train.logOther303.9 KB
scripts/blimp_avg.pyOther624 B
scripts/byte_lm_eval.pyOther18.1 KB
scripts/byte_lm_eval_bpe.pyOther1.9 KB
scripts/read_eval.pyOther694 B
scripts/train_gpt_ref.pyOther17.2 KB
.gitattributesRepository2.6 KB
manifests/tokenizer-training.jsonTokenizer1.1 KB
tokenizer/tokenizer.jsonTokenizer830.1 KB

License and Download

License
Not stated by the source
Access
No access gate
Download from Fabryka AI

Released by Fabryka AI through its official repository on Hugging Face.

Models Trained on This Dataset