SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically.
Dataset Card
SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process. The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The original SWE-bench dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Want to run inference now? This dataset only contains the problemstatement (i.e. issue text) and the basecommit which…
Excerpt from the card by SWE-bench.
Structure
default 500 rows
| Split | Rows | Size |
|---|---|---|
| test | 500 | 9.1 MB |
Details
- Repository
- SWE-bench/SWE-bench_Verified
- Publisher
- SWE-bench
- Task category
- Not stated by the source
- Tags
- Not stated by the source
- Size category
- Not stated by the source
- Languages
- Not stated by the source
- Revision
- 78f471bf655a3137b2e8a75af1501690ec009ec3
- Last updated
- 2026-08-16
Files
4 files, 6.3 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| data/test-00000-of-00001.parquet | Data | 6.3 MB | 030cfd7f2a70 |
| README.md | Documentation | 3.5 KB | — |
| eval.yaml | Other | 523 B | — |
| .gitattributes | Repository | 2.5 KB | — |
License and Download
- License
- Not stated by the source
- Access
- No access gate
Released by SWE-bench through its official repository on Hugging Face.