GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Dataset Card
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: A manually-curated evaluation dataset for fine-grained analysis of system performance on a broad range of linguistic phenomena. This dataset evaluates sentence understanding through Natural Language Inference (NLI) problems. Use a model trained on MulitNLI to produce predictions for this dataset. The Corpus of Linguistic Acceptability consists of English acceptability judgments drawn from…
Excerpt from the card by NYU Machine Learning for Language, licensed other.
Structure
ax 1,104 rows
| Split | Rows | Size |
|---|---|---|
| test | 1,104 | 243.8 KB |
cola 10,657 rows
| Split | Rows | Size |
|---|---|---|
| train | 8,551 | 522.3 KB |
| validation | 1,043 | 60.8 KB |
| test | 1,063 | 61.0 KB |
mnli 431,992 rows
| Split | Rows | Size |
|---|---|---|
| train | 392,702 | 75.5 MB |
| validation_matched | 9,815 | 1.8 MB |
| validation_mismatched | 9,832 | 1.9 MB |
| test_matched | 9,796 | 1.9 MB |
| test_mismatched | 9,847 | 2.0 MB |
mnli_matched 19,611 rows
| Split | Rows | Size |
|---|---|---|
| validation | 9,815 | 1.8 MB |
| test | 9,796 | 1.9 MB |
mnli_mismatched 19,679 rows
| Split | Rows | Size |
|---|---|---|
| validation | 9,832 | 1.9 MB |
| test | 9,847 | 2.0 MB |
mrpc 5,801 rows
| Split | Rows | Size |
|---|---|---|
| train | 3,668 | 940.4 KB |
| validation | 408 | 106.0 KB |
| test | 1,725 | 442.3 KB |
qnli 115,669 rows
| Split | Rows | Size |
|---|---|---|
| train | 104,743 | 25.4 MB |
| validation | 5,463 | 1.4 MB |
| test | 5,463 | 1.4 MB |
qqp 795,241 rows
| Split | Rows | Size |
|---|---|---|
| train | 363,846 | 51.6 MB |
| validation | 40,430 | 5.7 MB |
| test | 390,965 | 55.1 MB |
rte 5,767 rows
| Split | Rows | Size |
|---|---|---|
| train | 2,490 | 857.1 KB |
| validation | 277 | 90.8 KB |
| test | 3,000 | 724.5 KB |
sst2 70,042 rows
| Split | Rows | Size |
|---|---|---|
| train | 67,349 | 4.7 MB |
| validation | 872 | 106.5 KB |
| test | 1,821 | 218.3 KB |
stsb 8,628 rows
| Split | Rows | Size |
|---|---|---|
| train | 5,749 | 458.1 KB |
| validation | 1,500 | 195.8 KB |
| test | 1,379 | 156.8 KB |
wnli 852 rows
| Split | Rows | Size |
|---|---|---|
| train | 635 | 107.3 KB |
| validation | 71 | 12.2 KB |
| test | 146 | 37.9 KB |
Details
- Repository
- nyu-mll/glue
- Publisher
- NYU Machine Learning for Language
- Task category
- Text classification
- Tags
- qa-nli, coreference-nli, paraphrase-identification
- Size category
- 10K<n<100K
- Languages
- en
- Revision
- bcdcba79d07bc864c1c254ccfcedcce55bcc9a8c
- Last updated
- 2024-01-30
Files
36 files, 162.3 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| ax/test-00000-of-00001.parquet | Data | 80.8 KB | a07b802fe2d4 |
| cola/test-00000-of-00001.parquet | Data | 37.7 KB | 3c4d526b6f49 |
| cola/train-00000-of-00001.parquet | Data | 251.1 KB | 2e7538afa200 |
| cola/validation-00000-of-00001.parquet | Data | 37.6 KB | c14b7219a7d9 |
| mnli/test_matched-00000-of-00001.parquet | Data | 1.2 MB | a330c4f2aeb0 |
| mnli/test_mismatched-00000-of-00001.parquet | Data | 1.3 MB | e5078398d5c8 |
| mnli/train-00000-of-00001.parquet | Data | 52.2 MB | 49a4a5508b89 |
| mnli/validation_matched-00000-of-00001.parquet | Data | 1.2 MB | 7f918c09d9c3 |
| mnli/validation_mismatched-00000-of-00001.parquet | Data | 1.3 MB | 04aba92823a9 |
| mnli_matched/test-00000-of-00001.parquet | Data | 1.2 MB | a330c4f2aeb0 |
| mnli_matched/validation-00000-of-00001.parquet | Data | 1.2 MB | 7f918c09d9c3 |
| mnli_mismatched/test-00000-of-00001.parquet | Data | 1.3 MB | e5078398d5c8 |
| mnli_mismatched/validation-00000-of-00001.parquet | Data | 1.3 MB | 04aba92823a9 |
| mrpc/test-00000-of-00001.parquet | Data | 308.4 KB | a623ed1cbdf4 |
| mrpc/train-00000-of-00001.parquet | Data | 649.3 KB | 61fd41301e0e |
| mrpc/validation-00000-of-00001.parquet | Data | 75.7 KB | 33c007dbf5bf |
| qnli/test-00000-of-00001.parquet | Data | 877.3 KB | f39520cd0792 |
| qnli/train-00000-of-00001.parquet | Data | 17.5 MB | ebc7cb70a5bb |
| qnli/validation-00000-of-00001.parquet | Data | 872.1 KB | e69311b81dc6 |
| qqp/test-00000-of-00001.parquet | Data | 36.7 MB | 95d5d1efcfa3 |
| qqp/train-00000-of-00001.parquet | Data | 33.6 MB | 4d6f02e643f7 |
| qqp/validation-00000-of-00001.parquet | Data | 3.7 MB | efd86a539c41 |
| rte/test-00000-of-00001.parquet | Data | 621.4 KB | 3f44aadbfb8b |
| rte/train-00000-of-00001.parquet | Data | 584.0 KB | a6252ab17015 |
| rte/validation-00000-of-00001.parquet | Data | 69.0 KB | fb2aa2e04f55 |
| sst2/test-00000-of-00001.parquet | Data | 147.8 KB | e9d23cf00672 |
| sst2/train-00000-of-00001.parquet | Data | 3.1 MB | 66a253e67968 |
| sst2/validation-00000-of-00001.parquet | Data | 72.8 KB | a1371f3b3a7b |
| stsb/test-00000-of-00001.parquet | Data | 114.3 KB | 04fa2561f1ff |
| stsb/train-00000-of-00001.parquet | Data | 502.1 KB | bbd93bbb988f |
| stsb/validation-00000-of-00001.parquet | Data | 150.6 KB | 152de7cf1fa3 |
| wnli/test-00000-of-00001.parquet | Data | 13.6 KB | 766d3754c46a |
| wnli/train-00000-of-00001.parquet | Data | 38.8 KB | 40f4c0c60db6 |
| wnli/validation-00000-of-00001.parquet | Data | 11.1 KB | 880037e45e03 |
| README.md | Documentation | 35.3 KB | — |
| .gitattributes | Repository | 1.2 KB | — |
License and Download
- License
- other
- Access
- No access gate
Released by NYU Machine Learning for Language through its official repository on Hugging Face.
Models Trained on This Dataset
- Trained on (disclosed)ModernBERT-base-nli
- Trained on (disclosed)ModernBERT-large-nli
- Trained on (disclosed)deberta-small-long-nli
- Trained on (disclosed)deberta-base-long-nli