SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically.
Dataset Card
SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original SWE-bench dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the problemstatement (i.e. issue text) and the basecommit which…
Excerpt from the card by Princeton NLP group.
Structure
default 500 rows
| Split | Rows | Size |
|---|---|---|
| test | 500 | 7.8 MB |
Details
- Repository
- princeton-nlp/SWE-bench_Verified
- Publisher
- Princeton NLP group
- Task category
- Not stated by the source
- Tags
- Not stated by the source
- Size category
- Not stated by the source
- Languages
- Not stated by the source
- Revision
- c104f840cc67f8b6eec6f759ebc8b2693d585d4a
- Last updated
- 2025-02-18
Files
3 files, 2.1 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| data/test-00000-of-00001.parquet | Data | 2.1 MB | a45b1fe4e2f0 |
| README.md | Documentation | 3.3 KB | — |
| .gitattributes | Repository | 2.4 KB | — |
License and Download
- License
- Not stated by the source
- Access
- No access gate
Released by Princeton NLP group through its official repository on Hugging Face.