| Idavidrein/gpqa |
Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established |
85.3535 |
EvalEval Reported by a third party |
Evaluated revision not stated |
2026-04-20 |
| Idavidrein/gpqa |
Task diamondMetric diamondComparison conditions not established |
84.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-02-25 |
| MMMU/MMMU_Pro |
Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established |
75.1 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-28 |
| MathArena/aime_2026 |
Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established |
93.33 |
Official MathArena Evaluation Reported by a third party |
Evaluated revision not stated |
2026-03-17 |
| MathArena/hmmt_feb_2026 |
Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established |
81.82 |
Official MathArena Evaluation Reported by a third party |
Evaluated revision not stated |
2026-03-17 |
| SWE-bench/SWE-bench_Verified |
Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established |
69.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-02-03 |
| TIGER-Lab/MMLU-Pro |
Task mmlu_proMetric mmlu_proComparison conditions not established |
85.3 |
EvalEval Reported by a third party |
Evaluated revision not stated |
2026-06-30 |
| cais/hle |
Task hleMetric hleSetup chain of thoughtComparison conditions not established |
22.4 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-02-25 |
| harborframework/terminal-bench-2.0 |
Task terminalbench_2Metric terminalbench_2Comparison conditions not established |
40.5 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-02-25 |
| joelniklaus/LEXam-hard |
Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established |
23.93 |
SwissLegalEvals per-sample details (lighteval) Reported by a third party |
Evaluated revision not stated |
2026-08-01 |
| likaixin/ScreenSpot-Pro |
Task overallMetric overallComparison conditions not established |
68.6 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-03-18 |
| llamaindex/ExtractBench |
Task longMetric longSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established |
31.73 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task meanMetric meanSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established |
88 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task mediumMetric mediumSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established |
85.27 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task shortMetric shortSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established |
93.52 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |