| Idavidrein/gpqa |
Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established |
80.8081 |
EvalEval Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| Idavidrein/gpqa |
Task diamondMetric diamondSetup Reasoning: mediumComparison conditions not established |
73.1 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| Idavidrein/gpqa |
Task diamondMetric diamondSetup Reasoning: high, With toolsComparison conditions not established |
80.9 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| Idavidrein/gpqa |
Task diamondMetric diamondSetup Reasoning: highComparison conditions not established |
80.1 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| Idavidrein/gpqa |
Task diamondMetric diamondSetup Reasoning: low, With toolsComparison conditions not established |
68.1 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| Idavidrein/gpqa |
Task diamondMetric diamondSetup Reasoning: medium, With toolsComparison conditions not established |
73.5 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| Idavidrein/gpqa |
Task diamondMetric diamondSetup Reasoning: lowComparison conditions not established |
67.1 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| LEXam-Benchmark/LEXam |
Task mcq_4_choicesMetric mcq_4_choicesComparison conditions not established |
47.71 |
LEXam Leaderboard Reported by a third party |
Evaluated revision not stated |
2026-06-02 |
| LEXam-Benchmark/LEXam |
Task open_questionMetric open_questionComparison conditions not established |
51.74 |
LEXam Leaderboard Reported by a third party |
Evaluated revision not stated |
2026-06-02 |
| SWE-bench/SWE-bench_Verified |
Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: lowComparison conditions not established |
47.9 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| SWE-bench/SWE-bench_Verified |
Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: mediumComparison conditions not established |
52.6 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| SWE-bench/SWE-bench_Verified |
Task swe_bench_%_resolvedMetric swe_bench_%_resolvedSetup Reasoning: highComparison conditions not established |
62.4 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| ScaleAI/SWE-bench_Pro |
Task SWE_Bench_ProMetric SWE_Bench_ProComparison conditions not established |
16.2 |
SWE-Bench Pro official evaluation results Reported by a third party |
Evaluated revision not stated |
2026-02-28 |
| TIGER-Lab/MMLU-Pro |
Task mmlu_proMetric mmlu_proComparison conditions not established |
80.8 |
EvalEval Reported by a third party |
Evaluated revision not stated |
2026-06-30 |
| cais/hle |
Task hleMetric hleSetup Reasoning: lowComparison conditions not established |
5.2 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| cais/hle |
Task hleMetric hleSetup Reasoning: medium, With toolsComparison conditions not established |
11.3 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| cais/hle |
Task hleMetric hleSetup Reasoning: mediumComparison conditions not established |
8.6 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| cais/hle |
Task hleMetric hleSetup Reasoning: highComparison conditions not established |
14.9 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| cais/hle |
Task hleMetric hleSetup Reasoning: high, With toolsComparison conditions not established |
19 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| cais/hle |
Task hleMetric hleSetup Reasoning: low, With toolsComparison conditions not established |
9.1 |
GPT-OSS Model Card Reported by a third party |
Evaluated revision not stated |
2025-08-05 |
| joelniklaus/LEXam-hard |
Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established |
37.55 |
SwissLegalEvals per-sample details (lighteval) Reported by a third party |
Evaluated revision not stated |
2026-06-11 |