https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…
SAVRN Model Hub · Models by Task
Visual Question Answering Models
6 open-weight visual question answering models in the SAVRN Model Hub, with Gaoqie publishing the most.
Most Downloaded
| Model | Publisher | Parameters | License | Monthly downloads | Cheapest GPUs at 16-bit |
|---|---|---|---|---|---|
| InternVl2-8B-fire | Gaoqie | 8.1B | apache-2.0 | 78 | 1x MI300X, $1.85/hr |
| DeepSeekVL2-Tiny-fire | Gaoqie | 3.4B | apache-2.0 | 12 | 1x MI300X, $1.85/hr |
| DeepSeekVL-1.3B-Chat-fire | Gaoqie | 2B | apache-2.0 | 12 | 1x MI300X, $1.85/hr |
| Qwen2.5VL-3B-Instruct-fire | Gaoqie | 3.8B | apache-2.0 | 11 | 1x MI300X, $1.85/hr |
| Qwen2VL-2B-Instruct-fire | Gaoqie | 2.2B | apache-2.0 | 9 | 1x MI300X, $1.85/hr |
| DeepSeekVL-7B-Chat-fire | Gaoqie | 7.3B | apache-2.0 | 8 | 1x MI300X, $1.85/hr |
Licenses
| License | Models | Commercial use |
|---|---|---|
| apache-2.0 | 6 | Yes |
Who Publishes Them
| Publisher | Models |
|---|---|
| Gaoqie | 6 |
All 6 Models
https://www.doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…
https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…
https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…
https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…
https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…
Questions
Which Visual question answering models are most downloaded?
By monthly downloads reported by the Hugging Face Hub: InternVl2-8B-fire (78); DeepSeekVL-1.3B-Chat-fire (12); DeepSeekVL2-Tiny-fire (12).