SAVRN
Search Contact SAVRN

Dataset

tiny-base-300m-data

by BJY helloboy91/tiny-base-300m-data

This public pilot contains packed uint16 token IDs for pretraining the Tiny Base 300M, 2,048-context language model. It contains 50,000,000 training, 1,000,000 validation, and 1,000,000 test tokens. It does not redistribute raw documents.

Rows
Configurations
Size106.3 MB
Licenseother
AccessPublicly accessible
Monthly Downloads

Dataset Card

This public pilot contains packed uint16 token IDs for pretraining the Tiny Base 300M, 2,048-context language model. It contains 50,000,000 training, 1,000,000 validation, and 1,000,000 test tokens. It does not redistribute raw documents. train.bin, validation.bin, and test.bin are little-endian uint16 streams. Use tokenizer.json with the Hugging Face tokenizers library. manifest.json records exact token counts, split policy, source revisions, mixture, and SHA-256 checksums; settings.json records preparation settings. The deterministic split is SHA-256 of normalized text modulo 1,000: 0–9 test, 10–19 validation, and 20–999 training. The mixture is 70% FineWeb-Edu, 25% English Wikipedia, and…

Excerpt from the card by BJY, licensed other.

Details

Repository
helloboy91/tiny-base-300m-data
Publisher
BJY
Task category
Not stated by the source
Tags
pretraining, tokenized, english
Size category
n<1GB
Languages
en
Revision
385b4eaec80f628d96bc08d77b02109bd76f3527
Last updated
2026-09-18

Files

8 files, 106.3 MB in total.

Data2 files · 2.2 KB
Tokenizer1 file · 2.3 MB
Documentation1 file · 1.8 KB
Other3 files · 104.0 MB
Repository1 file · 2.5 KB
Every file
FileTypeSizeSHA-256
manifest.jsonData1.6 KB
settings.jsonData649 B
README.mdDocumentation1.8 KB
test.binOther2.0 MBe85fb2f775c5
train.binOther100.0 MB8f18a017f3bb
validation.binOther2.0 MB252bf6986e18
.gitattributesRepository2.5 KB
tokenizer.jsonTokenizer2.3 MB

License and Download

License
other
Access
No access gate
Download from BJY

Released by BJY through its official repository on Hugging Face.