SAVRN
Search Contact SAVRN

SAVRN Model Hub · Models by Task

Visual Question Answering Models

6 open-weight visual question answering models in the SAVRN Model Hub, with Gaoqie publishing the most.

6Models
1Publishers
2B to 8.1BParameter range
1Licenses

Most Downloaded

ModelPublisherParametersLicenseMonthly downloadsCheapest GPUs at 16-bit
InternVl2-8B-fire Gaoqie 8.1B apache-2.0 78 1x MI300X, $1.85/hr
DeepSeekVL2-Tiny-fire Gaoqie 3.4B apache-2.0 12 1x MI300X, $1.85/hr
DeepSeekVL-1.3B-Chat-fire Gaoqie 2B apache-2.0 12 1x MI300X, $1.85/hr
Qwen2.5VL-3B-Instruct-fire Gaoqie 3.8B apache-2.0 11 1x MI300X, $1.85/hr
Qwen2VL-2B-Instruct-fire Gaoqie 2.2B apache-2.0 9 1x MI300X, $1.85/hr
DeepSeekVL-7B-Chat-fire Gaoqie 7.3B apache-2.0 8 1x MI300X, $1.85/hr

Licenses

LicenseModelsCommercial use
apache-2.06Yes

Who Publishes Them

PublisherModels
Gaoqie6

All 6 Models

Model · Visual question answering

InternVl2-8B-fire

Gaoqie

https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…

Open weights apache-2.0 8.1B parameters 32,768 tokens

Model · Visual question answering

DeepSeekVL-1.3B-Chat-fire

Gaoqie

https://www.doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…

Open weights apache-2.0 2B parameters 16,384 tokens

Model · Visual question answering

DeepSeekVL2-Tiny-fire

Gaoqie

https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…

Open weights apache-2.0 3.4B parameters 4,096 tokens

Model · Visual question answering

Qwen2.5VL-3B-Instruct-fire

Gaoqie

https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…

Open weights apache-2.0 3.8B parameters 128,000 tokens

Model · Visual question answering

Qwen2VL-2B-Instruct-fire

Gaoqie

https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…

Open weights apache-2.0 2.2B parameters 32,768 tokens

Model · Visual question answering

DeepSeekVL-7B-Chat-fire

Gaoqie

https://doi.org/10.1007/s10694-026-02000-3 Existing vision-based methods suffer from high false alarm rates in urban flame detection. Applying Multimodal Large Language Models (MLLMs) for secondary filtering shows great potential in reducing false alarms, yet they have high inference latency and are prone to reasoning collapse on negative samples without explicit Chain-of-Thought (CoT) guidance. To overcome these challenges, this study proposed Flash-Cascade, the first sub-second MLLM-based firewall to leverage CoT to efficiently filter false alarms. We deconstructed the flame detection process into four logical stages (planning, observation, analysis, and judgment), which informed the…

Open weights apache-2.0 7.3B parameters 16,384 tokens

Questions

Which Visual question answering models are most downloaded?

By monthly downloads reported by the Hugging Face Hub: InternVl2-8B-fire (78); DeepSeekVL-1.3B-Chat-fire (12); DeepSeekVL2-Tiny-fire (12).

Other Tasks

See all