| Idavidrein/gpqa |
Task diamondMetric diamondComparison conditions not established |
86 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| MMMU/MMMU_Pro |
Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established |
75.3 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-05-15 |
| MathArena/aime_2026 |
Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established |
92.7 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| MathArena/hmmt_feb_2026 |
Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established |
83.6 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| SWE-bench/SWE-bench_Multilingual |
Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established |
67.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-08-10 |
| SWE-bench/SWE-bench_Verified |
Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established |
73.4 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| ScaleAI/SWE-bench_Pro |
Task SWE_Bench_ProMetric SWE_Bench_ProComparison conditions not established |
49.5 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| TIGER-Lab/MMLU-Pro |
Task mmlu_proMetric mmlu_proComparison conditions not established |
85.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| cais/hle |
Task hleMetric hleComparison conditions not established |
21.4 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| harborframework/terminal-bench-2.0 |
Task terminalbench_2Metric terminalbench_2Setup Harbor/Terminus-2 harness; 3h timeout, 32 CPU/48 GB RAM; temp=1.0, top_p=0.95, top_k=20, max_tokens=80K, 256K ctx; avg of 5 runs.Comparison conditions not established |
51.5 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| llamaindex/ExtractBench |
Task longMetric longSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established |
36.9 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-26 |
| llamaindex/ExtractBench |
Task meanMetric meanSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established |
88.11 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-26 |
| llamaindex/ExtractBench |
Task mediumMetric mediumSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established |
86.37 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-26 |
| llamaindex/ExtractBench |
Task shortMetric shortSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established |
92.86 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-26 |
| llamaindex/ParseBench |
Task chartMetric chartSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established |
5.1 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| llamaindex/ParseBench |
Task layoutMetric layoutSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established |
47.4 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| llamaindex/ParseBench |
Task meanMetric meanSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established |
44.1 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| llamaindex/ParseBench |
Task tableMetric tableSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established |
19.1 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| llamaindex/ParseBench |
Task text_contentMetric text_contentSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established |
90.7 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-16 |
| llamaindex/ParseBench |
Task text_formattingMetric text_formattingSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established |
58.3 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-16 |