For the official data release page, please see microsoft/SWE-bench-Live. SWE-bench-Live is a live benchmark for issue resolving, designed to evaluate an AI system’s ability to complete real-world software engineering tasks.
Dataset Card
By SWE-bench-Live, published under mit, revision b51a86422e10.
A brand-new, continuously updated SWE-bench-like dataset powered by an automated curation pipeline.
For the official data release page, please see microsoft/SWE-bench-Live.
Dataset Summary
SWE-bench-Live is a live benchmark for issue resolving, designed to evaluate an AI system’s ability to complete real-world software engineering tasks. Thanks to our automated dataset curation pipeline, we plan to update SWE-bench-Live on a monthly basis to provide the community with up-to-date task instances and support rigorous and contamination-free evaluation.
News
- 9/17/2025: Dataset updated! We’ve finalized the update process for SWE-bench-Live: Each month, we will add 50 newly verified, high-quality issues to the dataset. The
liteandverifiedsplits will remain frozen, ensuring fair leaderboard comparisons and keeping evaluation costs manageable. To access the latest issues, please refer to thefullsplit! - 07/19/2025: We've employed a LLM filter to automatically filter full dataset to create SWE-bench-Live Verified. The initial Verified subset contains 500 instances from 2024-07 to 2025-04.
- 06/30/2025: We’ve updated the dataset — it now includes a total of 1,565 task instances across 164 repositories!
- 05/21/2025: The initial release of SWE-bench-Live includes 1,319 latest (created after 2024) task instances, each paired with an instance-level Docker image for test execution, covering 93 repositories.
SWE-bench-Live Curation Pipeline
Dataset Structure
| Field | Type | Description |
|---|---|---|
repo |
str |
Repository full name |
pull_number |
str |
Pull request number |
instance_id |
str |
Unique identifier of the task instance |
issue_numbers |
str |
Issues that are resolved by the pull request |
base_commit |
str |
The commit on which the PR is based |
patch |
str |
Gold patch in .diff format |
test_patch |
str |
Test suite modifications |
problem_statement |
str |
Issue description text |
FAIL_TO_PASS |
List[str] |
Tests that should transition from failing to passing |
PASS_TO_PASS |
List[str] |
Tests that should remain passing |
image_key |
str |
Instance-level docker image |
test_cmds |
List[str] |
Commands to run test suite |
log_parser |
str |
Type of log parser, pytest by default |
Evaluation Protocol
[!NOTE] SWE-bench-Live evaluation strictly follows the original
SWE-benchprotocol:
- During a rollout, the agent may access only the
problem_statementfield of the Hugging Face dataset and the docker image of the task instance. It must not access any other fields, such ashint,FAIL_TO_PASS, ortest_patch. Thetest_patchmust not be applied to the repository before or during the rollout. The agent must perform a single rollout based solely on theproblem_statementon the docker container started from the image of the task instance.- Prompts, skills, and workflow instructions provided to the agent must not contain solutions specific to any task instance. They may contain only general instructions for the entire benchmark or, at most, for a specific repository. The agent must not use results from the ground-truth evaluation script to refine its solution.
Compliant prompts should follow the SWE-agent prompt and the OpenHands prompt, which contain only the problem statement and general workflow instructions.
When submitting results to our submissions repository, you must include your agent's raw rollout trajectories so that the maintainers can verify compliance with the SWE-bench protocol. A trajectory consists of the complete sequence of inputs to and outputs from your agent across all rollout rounds for a given task instance, including the initial prompt provided to the agent. Please follow this SWE-agent compliant trajectory example when submitting your result. If your organization's policies prohibit sharing the complete set of trajectories, you must provide at least some representative samples for verification. There is a checklist when submitting a PR to help you check whether you meet the protocol requirements again.
Structure
default 3,688 rows
| Split | Rows | Size |
|---|---|---|
| test | 1,000 | 294.2 MB |
| lite | 300 | 80.9 MB |
| verified | 500 | 140.8 MB |
| full | 1,888 | 539.5 MB |
Details
- Repository
- SWE-bench-Live/SWE-bench-Live
- Publisher
- SWE-bench-Live
- Task category
- Not stated by the source
- Tags
- Not stated by the source
- Size category
- Not stated by the source
- Languages
- Not stated by the source
- Revision
- b51a86422e10cfd403beb4773e5a2947953e36ec
- Last updated
- 2026-09-04
Files
27 files, 241.1 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| data/202401-00000-of-00001.parquet | Data | 2.2 MB | 3f519e0e396b |
| data/202402-00000-of-00001.parquet | Data | 3.4 MB | 45bf826cbafb |
| data/202403-00000-of-00001.parquet | Data | 2.8 MB | 3a885be0a63d |
| data/202404-00000-of-00001.parquet | Data | 2.8 MB | cd97e89db68f |
| data/202405-00000-of-00001.parquet | Data | 4.1 MB | fa773487738b |
| data/202406-00000-of-00001.parquet | Data | 2.9 MB | 056bafc280b4 |
| data/202407-00000-of-00001.parquet | Data | 2.6 MB | 21248cf7dcee |
| data/202408-00000-of-00001.parquet | Data | 2.2 MB | 0399ed86fa33 |
| data/202409-00000-of-00001.parquet | Data | 2.4 MB | 072caf96737f |
| data/202410-00000-of-00001.parquet | Data | 2.4 MB | 0d4a325a310d |
| data/202411-00000-of-00001.parquet | Data | 2.1 MB | 7172a8604e44 |
| data/202412-00000-of-00001.parquet | Data | 3.3 MB | 8e8ea515d0d0 |
| data/202501-00000-of-00001.parquet | Data | 2.7 MB | de82c862ef7a |
| data/202502-00000-of-00001.parquet | Data | 2.0 MB | 5a182ab6f646 |
| data/202503-00000-of-00001.parquet | Data | 2.4 MB | 732b6971430a |
| data/202504-00000-of-00001.parquet | Data | 3.0 MB | f6c5ebc78e0c |
| data/202505-00000-of-00001.parquet | Data | 3.5 MB | 031f9f8b2c6e |
| data/202506-00000-of-00001.parquet | Data | 2.2 MB | 1202acd70b01 |
| data/full-00000-of-00002.parquet | Data | 48.1 MB | 9371afaef65b |
| data/full-00001-of-00002.parquet | Data | 50.1 MB | 12e4956a1761 |
| data/lite-00000-of-00001.parquet | Data | 14.4 MB | 7ee0a75c41bf |
| data/test-00000-of-00001.parquet | Data | 52.9 MB | 7d6e65af7708 |
| data/verified-00000-of-00001.parquet | Data | 24.7 MB | 080e36e46198 |
| README.md | Documentation | 6.5 KB | — |
| assets/banner.png | Other | 822.3 KB | aac999523b6e |
| assets/overview.png | Other | 1.0 MB | fc680abec494 |
| .gitattributes | Repository | 2.5 KB | — |
License and Download
- License
- mit
- Access
- No access gate
Released by SWE-bench-Live through its official repository on Hugging Face. Read the license.