SAVRN
Search Contact SAVRN

Open-weight model · Any to any

Qwen3-Omni-30B-A3B-Thinking

by Qwen Qwen/Qwen3-Omni-30B-A3B-Thinking

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech.

Parameters31.7B
Context
Weights63.4 GB
Licenseother
AccessOpen weights
Monthly Downloads391.8k

Runs On

What it takes to serve Qwen3-Omni-30B-A3B-Thinking (31.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 63.4 GB 76.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 31.7 GB 38.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 15.9 GB 19.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen3-Omni-30B-A3B-Thinking

The memory figure is the first thing to read here. 63.4 GB of weights at 16-bit becomes 76.1 GB once working memory is counted, which is why the cheapest priced slot is one MI300X with 192 GB at $1.85 an hour. That leaves most of the card for the audio, video, images and text it takes in and the text and speech it streams out. At 8-bit the need falls to 38.1 GB and at 4-bit to 19.0 GB.

The license reads only as other, with no summary on file, so the publisher's own terms have to be pulled and read for commercial deployment before a rack is reserved. No context length is recorded either; confirm the working window against the audio and video durations you plan to feed it. Released 2025-09-15, with 391,783 monthly downloads and no host prices on our Index yet.

Model Card

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: Qwen3-Omni supports a wide range of multimodal application scenarios, covering various domain tasks involving audio, image, video, and audio-visual modalities. Below are several cookbooks demonstrating the usage cases of Qwen3-Omni and these cookbooks include our actual execution logs. You can first follow the QuickStart guide to download the model and install the necessary inference environment…

Excerpt from the card by Qwen, licensed other.

Configuration

Architecture
Qwen3OmniMoeForConditionalGeneration
Model type
qwen3_omni_moe

Identity and Version

Repository
Qwen/Qwen3-Omni-30B-A3B-Thinking
Publisher
Qwen
Task
Any to any
Modality
Multimodal
Library
transformers
Parameters
31.7B parameters
Languages
en
Revision
2f443cfc4c54b14a815c0e2bb9a9d6cbcd9a748b
First published
2025-09-15
Last updated
2025-09-22

Files and Weights

26 files, 63.4 GB in total. The weights are 16 files totalling 63.4 GB in safetensors.

Weights16 files · 63.4 GB
Configuration5 files · 1.9 MB
Tokenizer3 files · 4.5 MB
Documentation1 file · 90.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00016.safetensorsWeights4.0 GB ba3b117f1fe9
model-00002-of-00016.safetensorsWeights4.0 GB 0c419d9fc447
model-00003-of-00016.safetensorsWeights4.0 GB bdd6cc1cb3bf
model-00004-of-00016.safetensorsWeights4.0 GB c8228f28c83b
model-00005-of-00016.safetensorsWeights4.0 GB 4317632b9dd2
model-00006-of-00016.safetensorsWeights4.0 GB c86a2cdcdf7e
model-00007-of-00016.safetensorsWeights4.0 GB e670918aaab9
model-00008-of-00016.safetensorsWeights4.0 GB 1f44c6e93b85
model-00009-of-00016.safetensorsWeights4.0 GB 5439c95a4cbb
model-00010-of-00016.safetensorsWeights4.0 GB c75a2bb738fc
model-00011-of-00016.safetensorsWeights4.0 GB 35439fa67730
model-00012-of-00016.safetensorsWeights4.0 GB 61415e98a78c
model-00013-of-00016.safetensorsWeights4.0 GB 70587b5b7e51
model-00014-of-00016.safetensorsWeights4.0 GB bee3c092d296
model-00015-of-00016.safetensorsWeights4.0 GB 80ea0e6c2cb3
model-00016-of-00016.safetensorsWeights3.4 GB da33b1078d1b
chat_template.jsonConfiguration5.9 KB
config.jsonConfiguration8.5 KB
generation_config.jsonConfiguration113 B
model.safetensors.index.jsonConfiguration1.9 MB
preprocessor_config.jsonConfiguration603 B
README.mdDocumentation90.3 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer1.7 MB
tokenizer_config.jsonTokenizer7.3 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
other
Access
Open weights, no gate
Download size
63.4 GB
Download from Qwen

Released by Qwen through ModelScope.

Memory Requirements

PrecisionWeights in memory
As published63.4 GB
16-bit63.4 GB
8-bit31.7 GB
4-bit15.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3-Omni-30B-A3B-Thinking

How much GPU memory does Qwen3-Omni-30B-A3B-Thinking need?

About 76.1 GB at 16-bit and 19 GB at 4-bit: the weights (31.7B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3-Omni-30B-A3B-Thinking on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Qwen3-Omni-30B-A3B-Thinking released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 33B parameters 262,144 tokens transformers

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 33B parameters 262,144 tokens transformers

Model · Any to any

Qwen3-Omni-30B-A3B-Instruct

Qwen

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: Qwen3-Omni supports a wide range of multimodal application scenarios, covering various domain tasks involving audio, image, video, and audio-visual modalities. Below are several cookbooks demonstrating the usage cases of Qwen3-Omni and these cookbooks include our actual execution logs. You can first follow the QuickStart guide to download the model and install the necessary inference environment…

Open weights other 35.3B parameters transformers

Model · Any to any

Qwen3-Omni-30B-A3B-FP8

Markus

Block-wise FP8 quantization of Qwen/Qwen3-Omni-30B-A3B-Instruct. - Vision encoder (thinker.visual) - Audio tower (thinker.audiotower) - Code2Wav decoder (code2wav) - vLLM >= 0.13.0 with Qwen3-Omni support - 2x 24GB GPUs (e.g., RTX 3090) or equivalent - ~35 GB disk space Block-wise quantization with 128x128 blocks provides better precision than per-tensor quantization while maintaining good compression. Each block has its own scale factor stored as weightscaleinv (inverse scale for efficient multiplication during inference). This is a quantized version of Qwen/Qwen3-Omni-30B-A3B-Instruct. Qwen3-Omni is a natively end-to-end multilingual omni-modal foundation model that processes text…

Open weights other 35.3B parameters transformers

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 18.3B parameters 262,144 tokens transformers

Model · Any to any

gemma-4-12B-it-qat-w4a16-ct

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 13.3B parameters 262,144 tokens transformers