SAVRN
Search Contact SAVRN

Dataset

lambada_openai

by EleutherAI EleutherAI/lambada_openai

This dataset is comprised of the LAMBADA test split as pre-processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian.

Rows30,918
Configurations6
Size18.8 MB
Licensemit
AccessPublicly accessible
Monthly Downloads113.9k

Dataset Card

By EleutherAI, published under mit, revision 900124bf3b82.

Dataset Description

Dataset Summary

This dataset is comprised of the LAMBADA test split as pre-processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian.

LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole text, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local context, but must be able to keep track of information in the broader discourse.

Languages

English, German, Spanish, French, and Italian.

Source Data

Read the full dataset card (293 words)

Structure

default 5,153 rows

SplitRowsSize
test5,1531.7 MB
textstring

de 5,153 rows

SplitRowsSize
test5,1531.9 MB
textstring

en 5,153 rows

SplitRowsSize
test5,1531.7 MB
textstring

es 5,153 rows

SplitRowsSize
test5,1531.8 MB
textstring

fr 5,153 rows

SplitRowsSize
test5,1531.9 MB
textstring

it 5,153 rows

SplitRowsSize
test5,1531.8 MB
textstring

Details

Repository
EleutherAI/lambada_openai
Publisher
EleutherAI
Task category
Not stated by the source
Tags
Not stated by the source
Size category
1K<n<10K
Languages
de, en, es, fr, it
Revision
900124bf3b8235c6daf21033af9948b3f07346c4
Last updated
2025-07-10

Files

16 files, 18.8 MB in total.

Data14 files · 18.8 MB
Documentation1 file · 5.5 KB
Repository1 file · 2.7 KB
Every file
FileTypeSizeSHA-256
data/lambada_test.jsonlData1.8 MB4aa8d02cd17c
data/lambada_test_de.jsonlData2.0 MB51c6c1795894
data/lambada_test_en.jsonlData1.8 MB4aa8d02cd17c
data/lambada_test_es.jsonlData1.9 MBffd760026c64
data/lambada_test_fr.jsonlData2.0 MB941ec6a73dba
data/lambada_test_it.jsonlData1.9 MB866542377167
de/test/de.parquetData1.3 MBb229b3828eb8
default/test/default.parquetData1.2 MB99d8a9e54cda
en/test/en.parquetData1.2 MB99d8a9e54cda
es/test/es.parquetData1.2 MB20f22fc14177
fr/test/fr.parquetData1.3 MB1bdbedfcdbc9
it/test/it.parquetData1.2 MBf187d9036e68
lambada_openai.txtData4.8 KB
translation_script.txtData1.7 KB
README.mdDocumentation5.5 KB
.gitattributesRepository2.7 KB

License and Download

License
mit
Access
No access gate
Download from EleutherAI

Released by EleutherAI through its official repository on Hugging Face. Read the license.