explicit-edit-benchmark · Dataset Card
explicit-edit-benchmark: Dataset Card
Written by Aleksandr Shpuntenko, published under cc-by-4.0, revision 74fe434f8af3, read 2026-09-28. Shown as written; SAVRN's own facts about this dataset are on its page.
226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.
Source code and benchmark runner: GitHub — Explicit Edit Benchmark
Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.
Leaderboard by model route
Score v2 = coverage × quality, where quality is 75% first exact and 25% final exact.
Repeated runs are averaged inside each configuration and task; configurations then have
equal weight inside each task, and tasks have equal weight. The same numbers are in
views.json, and the
Explorer breaks them down by
harness, version and reasoning mode.
| Model | Provider | Score | Coverage | Tasks | Observations | Harnesses |
|---|---|---|---|---|---|---|
deepseek-v4.1-flash |
opencode-go |
99.3% | 100.0% | 226 | 226 | 2 |
deepseek-v4.1-flash |
deepseek |
99.0% | 100.0% | 226 | 2260 | 8 |
muse-spark-1.3-contributor |
opencode-go |
99.0% | 100.0% | 226 | 226 | 1 |
gpt-5.6-sol |
openai-codex |
98.9% | 100.0% | 226 | 226 | 1 |
deepseek-v4-flash |
deepseek |
98.1% | 100.0% | 226 | 2260 | 8 |
deepseek-v4.1-flash |
routerai-deepseek |
97.0% | 100.0% | 226 | 226 | 1 |
gpt-5.6-terra |
openai-codex |
96.5% | 100.0% | 226 | 452 | 2 |
kimi-k2.7-code |
opencode-go |
96.5% | 100.0% | 226 | 226 | 1 |
glm-5.3-flash |
zai |
96.0% | 100.0% | 226 | 1808 | 8 |
grok-4.6 |
opencode-go |
95.8% | 100.0% | 226 | 226 | 1 |
gpt-5.6-luna |
opencode-go |
94.1% | 100.0% | 226 | 226 | 1 |
gpt-5.6-luna |
openai-codex |
94.1% | 100.0% | 226 | 9955 | 34 |
glm-5.3-flash |
routerai-z-ai |
92.8% | 100.0% | 226 | 226 | 1 |
glm-5.3-flash |
opencode-go |
92.6% | 100.0% | 226 | 226 | 1 |
qwen3.7-plus |
opencode-go |
92.4% | 100.0% | 226 | 226 | 1 |
ling-3.0-flash-vl |
routerai-deepinfra |
90.3% | 100.0% | 226 | 226 | 1 |
qwen3.8-flash |
opencode-go |
88.1% | 100.0% | 226 | 226 | 1 |
mimo-v2.5 |
xiaomi |
86.5% | 100.0% | 226 | 1808 | 8 |
qwen3.6-35b-a3b |
unknown |
85.1% | 100.0% | 226 | 226 | 1 |
minimax-m3 |
opencode-go |
80.3% | 100.0% | 226 | 226 | 1 |
mercury-2.5 |
routerai-inception |
76.4% | 100.0% | 226 | 226 | 1 |
The harness list behind each row is in data/models.jsonl.gz, and views.json holds the same
aggregates for the other groupings: by harness, by agent and by reasoning mode.
Tables
| Config | One row per |
|---|---|
profiles |
configuration that was run, with its agent, harness, model and exact versions |
configurations |
recipe behind a configuration, safe to publish |
trials |
task and attempt, with the first and final exact result |
rounds |
attempt, with timing, tokens, cost and timeout state |
tool-calls |
tool the agent used, with its category and outcome |
submissions |
accepted run, with its owner, purpose and definitions |
The Dataset Viewer shows every config. dataset-index.json holds the source hashes, contracts, task sets, completeness and counts, and views.json, leaderboard.json and summary.json hold the aggregated rankings.
Source and contribution
Source code, run instructions and the contribution guide live at alexshpunt/explicit-edit-benchmark. Accepted bundles are kept under source/, the shards and summaries are views rebuilt from them, and each accepted harness family has a README badge under badges/.
Observations hold no prompts, arguments, commands, output, sessions, workspaces or credentials.