The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known…
Dataset Card
By Stanford NLP, published under cc-by-sa-4.0, revision cdb5c3d5eed6.
Dataset Card for SNLI
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: https://nlp.stanford.edu/projects/snli/
- Repository: [More Information Needed]
- Paper: https://aclanthology.org/D15-1075/
- Paper: https://arxiv.org/abs/1508.05326
- Leaderboard: https://nlp.stanford.edu/projects/snli/
- Point of Contact: Samuel Bowman
- Point of Contact: Gabor Angeli
- Point of Contact: Chris Manning
Dataset Summary
The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE).
Supported Tasks and Leaderboards
Structure
plain_text 570,152 rows
| Split | Rows | Size |
|---|---|---|
| test | 10,000 | 1.3 MB |
| validation | 10,000 | 1.3 MB |
| train | 550,152 | 68.1 MB |
Details
- Repository
- stanfordnlp/snli
- Publisher
- Stanford NLP
- Task category
- Text classification
- Tags
- Not stated by the source
- Size category
- 100K<n<1M
- Languages
- en
- Revision
- cdb5c3d5eed6ead6e5a341c8e56e669bb666725b
- Last updated
- 2024-03-06
Files
5 files, 20.5 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| plain_text/test-00000-of-00001.parquet | Data | 411.5 KB | 4696deda851c |
| plain_text/train-00000-of-00001.parquet | Data | 19.6 MB | ef9a7b25d973 |
| plain_text/validation-00000-of-00001.parquet | Data | 413.2 KB | 00f5ed8deaed |
| README.md | Documentation | 16.0 KB | — |
| .gitattributes | Repository | 1.2 KB | — |
License and Download
- License
- cc-by-sa-4.0
- Access
- No access gate
Released by Stanford NLP through its official repository on Hugging Face. Read the license.
Models Trained on This Dataset
- Trained on (disclosed)nli-deberta-v3-base
- Trained on (disclosed)nli-deberta-v3-small
- Trained on (disclosed)nli-MiniLM2-L6-H768
- Trained on (disclosed)nli-deberta-v3-xsmall
- Trained on (disclosed)nli-deberta-v3-large
- Trained on (disclosed)nli-distilroberta-base
- Trained on (disclosed)nli-roberta-base
- Trained on (disclosed)nli-deberta-base