SAVRN
Search Contact SAVRN

Dataset · Text classification

paws

by Google Research Datasets google-research-datasets/paws

This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification.

Rows751,450
Configurations3
Size129.3 MB
Licenseother
AccessPublicly accessible
Monthly Downloads213.4k

Dataset Card

This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further details, see the accompanying paper: PAWS: Paraphrase Adversaries from Word Scrambling (https://arxiv.org/abs/1904.01130) PAWS-QQP is not available due to license of QQP. It must be reconstructed by downloading the original data and then running our scripts to produce the data and attach the labels. The text in the dataset is in English. Below are two…

Excerpt from the card by Google Research Datasets, licensed other.

Structure

labeled_final 65,401 rows

SplitRowsSize
train49,40112.1 MB
test8,0002.0 MB
validation8,0002.0 MB
idint32sentence1stringsentence2stringlabelClassLabel

labeled_swap 30,397 rows

SplitRowsSize
train30,3977.9 MB
idint32sentence1stringsentence2stringlabelClassLabel

unlabeled_final 655,652 rows

SplitRowsSize
train645,652156.3 MB
validation10,0002.5 MB
idint32sentence1stringsentence2stringlabelClassLabel

Details

Repository
google-research-datasets/paws
Publisher
Google Research Datasets
Task category
Text classification
Tags
paraphrase-identification
Size category
100K<n<1M
Languages
en
Revision
161ece9501cf0a11f3e48bd356eaa82de46d6a09
Last updated
2024-01-04

Files

8 files, 129.3 MB in total.

Data6 files · 129.3 MB
Documentation1 file · 9.8 KB
Repository1 file · 1.2 KB
Every file
FileTypeSizeSHA-256
labeled_final/test-00000-of-00001.parquetData1.2 MBae342ff12bb8
labeled_final/train-00000-of-00001.parquetData8.4 MB8dc9ad3e5f30
labeled_final/validation-00000-of-00001.parquetData1.2 MB7760d8294537
labeled_swap/train-00000-of-00001.parquetData5.7 MB21b9b68cd939
unlabeled_final/train-00000-of-00001.parquetData111.0 MBc9724b30b187
unlabeled_final/validation-00000-of-00001.parquetData1.7 MB438720a7c059
README.mdDocumentation9.8 KB
.gitattributesRepository1.2 KB

License and Download

License
other
Access
No access gate
Download from Google Research Datasets

Released by Google Research Datasets through its official repository on Hugging Face.