226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Score v2 = coverage × quality, where quality is 75% first exact and 25% final exact. Repeated runs are averaged inside each configuration and task; configurations then have equal weight inside each task, and tasks have equal weight. The same numbers are in views.json, and the Explorer breaks them down by harness, version and reasoning mode. The harness…
Independent publisher
Aleksandr Shpuntenko
alexshpunt
Models in Library0
Datasets in Library1
Models on Hugging Face—
Followers—