| Idavidrein/gpqa |
Task diamondMetric diamondComparison conditions not established |
89.2 |
Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated |
2026-08-14 |
| ScaleAI/SWE-bench_Pro |
Task SWE_Bench_ProMetric SWE_Bench_ProSetup Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, 256K context; baseline models re-evaluated on the same refined task set.Comparison conditions not established |
61.7 |
Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated |
2026-08-14 |
| cais/hle |
Task hleMetric hleSetup Judged by GPT-4o.Comparison conditions not established |
30.8 |
Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated |
2026-08-14 |
| claw-eval/Claw-Eval |
Task multimodalMetric multimodalSetup Reported as ClawEval-MM. Pass@3, the benchmark's own Pass³ methodology; card also reports a secondary 56.9 'Average' metric, not included here.Comparison conditions not established |
57.4 |
Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated |
2026-08-14 |
| datacurve/deep-swe |
Task deep_sweMetric deep_sweComparison conditions not established |
42.2 |
Qwen3.8-27B model card Reported by a third party |
Evaluated revision not stated |
2026-08-14 |
| harborframework/terminal-bench-2.1 |
Task terminalbench_2_1Metric terminalbench_2_1Setup Row labeled "(Terminus)" as the harness; no further hyperparameter footnote given for this row.Comparison conditions not established |
73 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-08-14 |
| internlm/WildClawBench |
Task avg_timeMetric avg_timeComparison conditions not established |
516 |
WildClawBench Reported by a third party |
Evaluated revision not stated |
2026-08-16 |
| internlm/WildClawBench |
Task overallMetric overallComparison conditions not established |
48.0152 |
WildClawBench Reported by a third party |
Evaluated revision not stated |
2026-08-16 |
| llamaindex/ExtractBench |
Task longMetric longSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
38.45 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task meanMetric meanSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
89.75 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task mediumMetric mediumSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
87.54 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task shortMetric shortSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
94.68 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ParseBench |
Task chartMetric chartSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
69.17 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-08-28 |
| llamaindex/ParseBench |
Task layoutMetric layoutSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
69.9 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-08-28 |
| llamaindex/ParseBench |
Task meanMetric meanSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
70.79 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-08-28 |
| llamaindex/ParseBench |
Task tableMetric tableSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
66.82 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-08-28 |
| llamaindex/ParseBench |
Task text_contentMetric text_contentSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
88.28 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-08-28 |
| llamaindex/ParseBench |
Task text_formattingMetric text_formattingSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established |
59.77 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-08-28 |