SAVRN
Search Contact SAVRN

Dataset · Text classification

snli

by Stanford NLP stanfordnlp/snli

The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known…

Rows570,152
Configurations1
Size20.5 MB
Licensecc-by-sa-4.0
AccessPublicly accessible
Monthly Downloads104.9k

Dataset Card

By Stanford NLP, published under cc-by-sa-4.0, revision cdb5c3d5eed6.

Dataset Card for SNLI

Table of Contents

  • Dataset Description
  • Dataset Summary
  • Supported Tasks and Leaderboards
  • Languages
  • Dataset Structure
  • Data Instances
  • Data Fields
  • Data Splits
  • Dataset Creation
  • Curation Rationale
  • Source Data
  • Annotations
  • Personal and Sensitive Information
  • Considerations for Using the Data
  • Social Impact of Dataset
  • Discussion of Biases
  • Other Known Limitations
  • Additional Information
  • Dataset Curators
  • Licensing Information
  • Citation Information
  • Contributions

Dataset Description

  • Homepage: https://nlp.stanford.edu/projects/snli/
  • Repository: [More Information Needed]
  • Paper: https://aclanthology.org/D15-1075/
  • Paper: https://arxiv.org/abs/1508.05326
  • Leaderboard: https://nlp.stanford.edu/projects/snli/
  • Point of Contact: Samuel Bowman
  • Point of Contact: Gabor Angeli
  • Point of Contact: Chris Manning

Dataset Summary

The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE).

Supported Tasks and Leaderboards

Read the full dataset card (1,826 words)

Structure

plain_text 570,152 rows

SplitRowsSize
test10,0001.3 MB
validation10,0001.3 MB
train550,15268.1 MB
premisestringhypothesisstringlabelClassLabel

Details

Repository
stanfordnlp/snli
Publisher
Stanford NLP
Task category
Text classification
Tags
Not stated by the source
Size category
100K<n<1M
Languages
en
Revision
cdb5c3d5eed6ead6e5a341c8e56e669bb666725b
Last updated
2024-03-06

Files

5 files, 20.5 MB in total.

Data3 files · 20.4 MB
Documentation1 file · 16.0 KB
Repository1 file · 1.2 KB
Every file
FileTypeSizeSHA-256
plain_text/test-00000-of-00001.parquetData411.5 KB4696deda851c
plain_text/train-00000-of-00001.parquetData19.6 MBef9a7b25d973
plain_text/validation-00000-of-00001.parquetData413.2 KB00f5ed8deaed
README.mdDocumentation16.0 KB
.gitattributesRepository1.2 KB

License and Download

License
cc-by-sa-4.0
Access
No access gate
Download from Stanford NLP

Released by Stanford NLP through its official repository on Hugging Face. Read the license.

Models Trained on This Dataset