SAVRN
Search Contact SAVRN

explicit-edit-benchmark · Dataset Card

explicit-edit-benchmark: Dataset Card

Written by Aleksandr Shpuntenko, published under cc-by-4.0, revision 74fe434f8af3, read 2026-09-28. Shown as written; SAVRN's own facts about this dataset are on its page.

226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte.

Source code and benchmark runner: GitHub — Explicit Edit Benchmark

Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens.

Leaderboard by model route

Score v2 = coverage × quality, where quality is 75% first exact and 25% final exact. Repeated runs are averaged inside each configuration and task; configurations then have equal weight inside each task, and tasks have equal weight. The same numbers are in views.json, and the Explorer breaks them down by harness, version and reasoning mode.

Model Provider Score Coverage Tasks Observations Harnesses
deepseek-v4.1-flash opencode-go 99.3% 100.0% 226 226 2
deepseek-v4.1-flash deepseek 99.0% 100.0% 226 2260 8
muse-spark-1.3-contributor opencode-go 99.0% 100.0% 226 226 1
gpt-5.6-sol openai-codex 98.9% 100.0% 226 226 1
deepseek-v4-flash deepseek 98.1% 100.0% 226 2260 8
deepseek-v4.1-flash routerai-deepseek 97.0% 100.0% 226 226 1
gpt-5.6-terra openai-codex 96.5% 100.0% 226 452 2
kimi-k2.7-code opencode-go 96.5% 100.0% 226 226 1
glm-5.3-flash zai 96.0% 100.0% 226 1808 8
grok-4.6 opencode-go 95.8% 100.0% 226 226 1
gpt-5.6-luna opencode-go 94.1% 100.0% 226 226 1
gpt-5.6-luna openai-codex 94.1% 100.0% 226 9955 34
glm-5.3-flash routerai-z-ai 92.8% 100.0% 226 226 1
glm-5.3-flash opencode-go 92.6% 100.0% 226 226 1
qwen3.7-plus opencode-go 92.4% 100.0% 226 226 1
ling-3.0-flash-vl routerai-deepinfra 90.3% 100.0% 226 226 1
qwen3.8-flash opencode-go 88.1% 100.0% 226 226 1
mimo-v2.5 xiaomi 86.5% 100.0% 226 1808 8
qwen3.6-35b-a3b unknown 85.1% 100.0% 226 226 1
minimax-m3 opencode-go 80.3% 100.0% 226 226 1
mercury-2.5 routerai-inception 76.4% 100.0% 226 226 1

The harness list behind each row is in data/models.jsonl.gz, and views.json holds the same aggregates for the other groupings: by harness, by agent and by reasoning mode.

Tables

Config One row per
profiles configuration that was run, with its agent, harness, model and exact versions
configurations recipe behind a configuration, safe to publish
trials task and attempt, with the first and final exact result
rounds attempt, with timing, tokens, cost and timeout state
tool-calls tool the agent used, with its category and outcome
submissions accepted run, with its owner, purpose and definitions

The Dataset Viewer shows every config. dataset-index.json holds the source hashes, contracts, task sets, completeness and counts, and views.json, leaderboard.json and summary.json hold the aggregated rankings.

Source and contribution

Source code, run instructions and the contribution guide live at alexshpunt/explicit-edit-benchmark. Accepted bundles are kept under source/, the shards and summaries are views rebuilt from them, and each accepted harness family has a README badge under badges/.

Observations hold no prompts, arguments, commands, output, sessions, workspaces or credentials.