| Idavidrein/gpqa |
Task diamondMetric diamondComparison conditions not established |
11.9 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-03-02 |
| LiquidAI/ifstruct-v1.0 |
Task ifstruct_v1Metric ifstruct_v1Comparison conditions not established |
15.5 |
Liquid AI — IFStruct v1.0 blog (Qwen3.5-0.8B) Reported by a third party |
Evaluated revision not stated |
2026-06-30 |
| MMMU/MMMU_Pro |
Task mmmu_pro_visionMetric mmmu_pro_visionSetup ThinkingComparison conditions not established |
31.2 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-04-28 |
| TIGER-Lab/MMLU-Pro |
Task mmlu_proMetric mmlu_proComparison conditions not established |
29.7 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-03-02 |
| likaixin/ScreenSpot-Pro |
Task overallMetric overallComparison conditions not established |
46.5 |
Model Card Reported by a third party |
Evaluated revision not stated |
2026-03-18 |
| llamaindex/ExtractBench |
Task longMetric longSetup Pipeline name: qwen3_5_0_8b_vllm_extract_oneshot_structured_output_fileComparison conditions not established |
6.65 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task meanMetric meanSetup Pipeline name: qwen3_5_0_8b_vllm_extract_oneshot_structured_output_fileComparison conditions not established |
46.63 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task mediumMetric mediumSetup Pipeline name: qwen3_5_0_8b_vllm_extract_oneshot_structured_output_fileComparison conditions not established |
26.98 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ExtractBench |
Task shortMetric shortSetup Pipeline name: qwen3_5_0_8b_vllm_extract_oneshot_structured_output_fileComparison conditions not established |
57.44 |
ExtractBench Reported by a third party |
Evaluated revision not stated |
2026-08-24 |
| llamaindex/ParseBench |
Task chartMetric chartSetup Pipeline name: qwen3_5_0_8b_vllm_layoutComparison conditions not established |
0.4 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| llamaindex/ParseBench |
Task layoutMetric layoutSetup Pipeline name: qwen3_5_0_8b_vllm_layoutComparison conditions not established |
15 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| llamaindex/ParseBench |
Task meanMetric meanSetup Pipeline name: qwen3_5_0_8b_vllm_layoutComparison conditions not established |
28.4 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| llamaindex/ParseBench |
Task tableMetric tableSetup Pipeline name: qwen3_5_0_8b_vllm_layoutComparison conditions not established |
1.5 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| llamaindex/ParseBench |
Task text_contentMetric text_contentSetup Pipeline name: qwen3_5_0_8b_vllm_layoutComparison conditions not established |
82 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-22 |
| llamaindex/ParseBench |
Task text_formattingMetric text_formattingSetup Pipeline name: qwen3_5_0_8b_vllm_layoutComparison conditions not established |
43.1 |
ParseBench Reported by a third party |
Evaluated revision not stated |
2026-04-22 |