| Idavidrein/gpqa |
Task diamondMetric diamondComparison conditions not established |
87.8 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| MMMU/MMMU_Pro |
Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established |
75.8 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-05-15 |
| MathArena/aime_2026 |
Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established |
94.1 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| MathArena/hmmt_feb_2026 |
Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established |
84.3 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| SWE-bench/SWE-bench_Multilingual |
Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established |
71.3 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-08-10 |
| SWE-bench/SWE-bench_Verified |
Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established |
77.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| ScaleAI/SWE-bench_Pro |
Task SWE_Bench_ProMetric SWE_Bench_ProComparison conditions not established |
53.5 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| TIGER-Lab/MMLU-Pro |
Task mmlu_proMetric mmlu_proComparison conditions not established |
86.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| benchflow/skillsbench |
Task skillsbench_v1_1Metric skillsbench_v1_1Setup SkillsBench Avg5. Evaluated via OpenCode on 78 tasks (self-contained subset, excluding API-dependent tasks); avg of 5 runs. Mapped to the Hub's benchflow/skillsbench task skillsbench_v1_1.Comparison conditions not established |
48.2 |
Qwen3.6-27B model card — Benchmark Results (Language > Coding Agent) Reported by a third party |
Evaluated revision not stated |
2026-04-21 |
| cais/hle |
Task hleMetric hleComparison conditions not established |
24 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| harborframework/terminal-bench-2.0 |
Task terminalbench_2Metric terminalbench_2Setup Harbor/Terminus-2 harness; 3h timeout, 32 CPU/48 GB RAM; temp=1.0, top_p=0.95, top_k=20, max_tokens=80K, 256K ctx; avg of 5 runs.Comparison conditions not established |
59.3 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-06-23 |
| internlm/WildClawBench |
Task avg_costMetric avg_costComparison conditions not established |
20.91 |
WildClawBench Reported by a third party |
Evaluated revision not stated |
2026-08-11 |
| internlm/WildClawBench |
Task avg_timeMetric avg_timeComparison conditions not established |
421 |
WildClawBench Reported by a third party |
Evaluated revision not stated |
2026-08-11 |
| internlm/WildClawBench |
Task overallMetric overallComparison conditions not established |
43.2 |
WildClawBench Reported by a third party |
Evaluated revision not stated |
2026-08-11 |