This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification.
Dataset Card
This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further details, see the accompanying paper: PAWS: Paraphrase Adversaries from Word Scrambling (https://arxiv.org/abs/1904.01130) PAWS-QQP is not available due to license of QQP. It must be reconstructed by downloading the original data and then running our scripts to produce the data and attach the labels. The text in the dataset is in English. Below are two…
Excerpt from the card by Google Research Datasets, licensed other.
Structure
labeled_final 65,401 rows
| Split | Rows | Size |
|---|---|---|
| train | 49,401 | 12.1 MB |
| test | 8,000 | 2.0 MB |
| validation | 8,000 | 2.0 MB |
labeled_swap 30,397 rows
| Split | Rows | Size |
|---|---|---|
| train | 30,397 | 7.9 MB |
unlabeled_final 655,652 rows
| Split | Rows | Size |
|---|---|---|
| train | 645,652 | 156.3 MB |
| validation | 10,000 | 2.5 MB |
Details
- Repository
- google-research-datasets/paws
- Publisher
- Google Research Datasets
- Task category
- Text classification
- Tags
- paraphrase-identification
- Size category
- 100K<n<1M
- Languages
- en
- Revision
- 161ece9501cf0a11f3e48bd356eaa82de46d6a09
- Last updated
- 2024-01-04
Files
8 files, 129.3 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| labeled_final/test-00000-of-00001.parquet | Data | 1.2 MB | ae342ff12bb8 |
| labeled_final/train-00000-of-00001.parquet | Data | 8.4 MB | 8dc9ad3e5f30 |
| labeled_final/validation-00000-of-00001.parquet | Data | 1.2 MB | 7760d8294537 |
| labeled_swap/train-00000-of-00001.parquet | Data | 5.7 MB | 21b9b68cd939 |
| unlabeled_final/train-00000-of-00001.parquet | Data | 111.0 MB | c9724b30b187 |
| unlabeled_final/validation-00000-of-00001.parquet | Data | 1.7 MB | 438720a7c059 |
| README.md | Documentation | 9.8 KB | — |
| .gitattributes | Repository | 1.2 KB | — |
License and Download
- License
- other
- Access
- No access gate
Released by Google Research Datasets through its official repository on Hugging Face.