SAVRN
Search Contact SAVRN

Open-weight model · Any to any

Ming-flash-omni-2.0

by inclusionAI inclusionAI/Ming-flash-omni-2.0

Ming-flash-omni-2.0 is an open-weight model for any to any from inclusionAI, released under MIT License. It has 104.2B parameters and a 32,768-token context. At 16-bit it needs about 250.2 GB of GPU memory, which fits on 1x MI325X from $2.00 an hour; at 4-bit, 62.5 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 5.1k downloads a month.

The newly released Ming-flash-omni 2.0 leverages the Ling-2.0 architecture—a Mixture-of-Experts (MoE) framework comprising 100B total and 6B active parameters.

Parameters104.2B
Context32,768
Weights238.0 GB
Licensemit
AccessOpen weights
Monthly Downloads5.1k

Runs On

What it takes to serve Ming-flash-omni-2.0 (104.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 208.5 GB 250.2 GB 1x MI325X (256 GB)
Vultr
$2.00 1x MI355X $2.59 · 2x MI300X $3.70
8-bit 104.2 GB 125.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 52.1 GB 62.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Ming-flash-omni-2.0 on every accelerator the SAVRN Index prices, at every precision

Model Card

By inclusionAI, published under mit, revision c1ed5565e64e.

Technical Report | Hugging Face | ModelScope ## Introduction The newly released Ming-flash-omni 2.0 leverages the [Ling-2.0](https://github.com/inclusionAI/Ling-V2) architecture—a Mixture-of-Experts (MoE) framework comprising 100B total and 6B active parameters. Representing a generational advancement over its predecessor, it establishes new State-of-the-Art (SOTA) benchmarks among open-source omni-MLLMs. Ming-flash-omni 2.0 effectively synergizes foundational abilities with specialized domain expertise. In particular, it exhibits superior performance in visual encyclopedic knowledge, immersive speech synthesis, and high-dynamic image generation and manipulation. ## Updates * [2026.02.11] We release the official version of [Ming-flash-omni 2.0](https://mp.weixin.qq.com/s/hz2fsH1DGpp2zpY-Yngsog), an open-source SOTA omni-MLLM that pushes the boundaries of multimodal understanding and synthesis. * [2025.10.27] We release the preview version of Ming-flash-omni:[Ming-flash-omni Preview](https://github.com/inclusionAI/Ming/tree/main). * [2025.07.15] We release [Ming-lite-omni v1.5](https://github.com/inclusionAI/Ming/tree/v1.5) with significant improvements across all modalities. * [2025.06.12] Our [Technical Report](https://arxiv.org/abs/2506.09344) is in public on arxiv. * [2025.05.28] The official version of [Ming-lite-omni v1](https://github.com/inclusionAI/Ming/tree/v1.0) is released, with better performance and image generation support. * [2025.05.04] We release the test version of Ming-lite-omni:[Ming-lite-omni-Preview](https://github.com/inclusionAI/Ming/tree/Ming-Lite-Omni-Preview). ## Key Features Compared…

Read the full model card (879 words)

Configuration

Architecture
BailingMM2NativeForConditionalGeneration
Context length (tokens)
32,768
Layers
32
Hidden size
4,096
Feed-forward size
9,216
Attention heads
32
Key/value heads
4
Vocabulary size
157,184
Experts
256
Experts active per token
8
RoPE base
2,400,000
Stored precision
bfloat16
Model type
bailingmm_moe_v2_lite

Identity and Version

Repository
inclusionAI/Ming-flash-omni-2.0
Publisher
inclusionAI
Task
Any to any
Modality
Multimodal
Library
diffusers
Parameters
104.2B parameters
Languages
en
Revision
c1ed5565e64e85a76826255c000b66a20d4634c3
First published
2026-02-10
Last updated
2026-09-29

Files and Weights

99 files, 238.0 GB in total. The weights are 60 files totalling 238.0 GB in bin, h5, msgpack, onnx, pt, safetensors.

Weights60 files · 238.0 GB
Configuration29 files · 2.6 MB
Tokenizer5 files · 15.9 MB
Documentation2 files · 16.3 KB
Other1 file · 307.8 KB
Repository2 files · 2.4 KB
Every file
FileTypeSizeSHA-256
byt5/byt5_mapper/byt5_mapper.ptWeights304.6 MB 74a3cba1ce71
byt5/byt5_model/base.ptWeights3.0 GB 5f0816c04e12
byt5/byt5_model/byt5_model.ptWeights877.3 MB ca8c97c89136
byt5/google__byt5-smal/flax_model.msgpackWeights1.2 GB b3aafee96d60
byt5/google__byt5-smal/pytorch_model.binWeights1.2 GB 5c5aaf56299d
byt5/google__byt5-smal/tf_model.h5Weights1.2 GB f97320dd5eb4
connector/model-00001-of-00002.safetensorsWeights5.0 GB c021bcc12289
connector/model-00002-of-00002.safetensorsWeights1.2 GB f49049386ed6
mlp/model.safetensorsWeights45.1 MB 22997acfa6b9
model-00001-of-00042.safetensorsWeights5.0 GB 3edf1c3a549c
model-00002-of-00042.safetensorsWeights5.0 GB 435534ce3b2a
model-00003-of-00042.safetensorsWeights5.0 GB 8c6974eced77
model-00004-of-00042.safetensorsWeights5.0 GB 013ed1e2b926
model-00005-of-00042.safetensorsWeights5.0 GB 22cd341b7925
model-00006-of-00042.safetensorsWeights5.0 GB 9bb1d06e4132
model-00007-of-00042.safetensorsWeights5.0 GB 564d61b480e2
model-00008-of-00042.safetensorsWeights5.0 GB 3b3fecdd1b22
model-00009-of-00042.safetensorsWeights5.0 GB 2f14a4fa5515
model-00010-of-00042.safetensorsWeights5.0 GB 9bf9683ee56e
model-00011-of-00042.safetensorsWeights5.0 GB 4275c8966395
model-00012-of-00042.safetensorsWeights5.0 GB 6970d11a6dfc
model-00013-of-00042.safetensorsWeights5.0 GB e62f403a2a6c
model-00014-of-00042.safetensorsWeights5.0 GB 2e52e582facb
model-00015-of-00042.safetensorsWeights5.0 GB 00da6658de16
model-00016-of-00042.safetensorsWeights5.0 GB 4499f51c9289
model-00017-of-00042.safetensorsWeights5.0 GB 9ca2cb3052b8
model-00018-of-00042.safetensorsWeights5.0 GB 2257ab127438
model-00019-of-00042.safetensorsWeights5.0 GB af19341964d4
model-00020-of-00042.safetensorsWeights5.0 GB 74fe06f8ac5b
model-00021-of-00042.safetensorsWeights5.0 GB 7440a749d924
model-00022-of-00042.safetensorsWeights5.0 GB bcaaa83a9c78
model-00023-of-00042.safetensorsWeights5.0 GB c3a3bd0495a8
model-00024-of-00042.safetensorsWeights5.0 GB b0a35bc3f498
model-00025-of-00042.safetensorsWeights5.0 GB dc97c036c59b
model-00026-of-00042.safetensorsWeights5.0 GB 4489dc484bb6
model-00027-of-00042.safetensorsWeights5.0 GB 5e9be0110cfd
model-00028-of-00042.safetensorsWeights5.0 GB 5dbc75608840
model-00029-of-00042.safetensorsWeights5.0 GB f69f641d1535
model-00030-of-00042.safetensorsWeights5.0 GB 66279b01fb41
model-00031-of-00042.safetensorsWeights5.0 GB a332d936b24e
model-00032-of-00042.safetensorsWeights5.0 GB f07ddcad0bd4
model-00033-of-00042.safetensorsWeights5.0 GB c71cb58bc686
model-00034-of-00042.safetensorsWeights5.0 GB bf71c915534f
model-00035-of-00042.safetensorsWeights5.0 GB af1414dd1120
model-00036-of-00042.safetensorsWeights5.0 GB 6553a3359a0b
model-00037-of-00042.safetensorsWeights5.0 GB 67f7c432429c
model-00038-of-00042.safetensorsWeights5.0 GB ef52ac1e1642
model-00039-of-00042.safetensorsWeights5.0 GB b746d362a232
model-00040-of-00042.safetensorsWeights5.0 GB 3fdb2533d0f9
model-00041-of-00042.safetensorsWeights5.0 GB dcf4415df371
model-00042-of-00042.safetensorsWeights3.6 GB 06386221000f
talker/campplus.onnxWeights28.3 MB a6ac6a639977
talker/model.safetensorsWeights1.4 GB 1403c23a0a91
talker/vae/model.safetensorsWeights1.6 GB 61ffb5e9ca7e
transformer/diffusion_pytorch_model-00001-of-00005.safetensorsWeights3.0 GB ff08b8209825
transformer/diffusion_pytorch_model-00002-of-00005.safetensorsWeights2.9 GB 76abee2c7c2c
transformer/diffusion_pytorch_model-00003-of-00005.safetensorsWeights3.0 GB 9c8f4a400f3e
transformer/diffusion_pytorch_model-00004-of-00005.safetensorsWeights3.0 GB 688c6ef5fd85
transformer/diffusion_pytorch_model-00005-of-00005.safetensorsWeights448.4 MB a45e05dc23c3
vae/diffusion_pytorch_model.safetensorsWeights167.7 MB f5b59a268515
byt5/byt5.jsonConfiguration491 B —
byt5/color_idx.jsonConfiguration2.4 KB —
byt5/font_uni_10-lang_idx.jsonConfiguration24.3 KB —
byt5/google__byt5-smal/config.jsonConfiguration698 B —
byt5/google__byt5-smal/generation_config.jsonConfiguration147 B —
byt5/google__byt5-smal/special_tokens_map.jsonConfiguration2.5 KB —
config.jsonConfiguration4.7 KB —
connector/config.jsonConfiguration1.3 KB —
connector/generation_config.jsonConfiguration242 B —
connector/model.safetensors.index.jsonConfiguration27.7 KB —
mlp/config.jsonConfiguration166 B —
model.safetensors.index.jsonConfiguration2.4 MB —
scheduler/scheduler_config.jsonConfiguration173 B —
talker/__init__.pyConfiguration —
talker/async_vllm_infer.pyConfiguration6.9 KB —
talker/config.jsonConfiguration707 B —
talker/llm/added_tokens.jsonConfiguration1.8 KB —
talker/llm/config.jsonConfiguration659 B —
talker/llm/generation_config.jsonConfiguration242 B —
talker/llm/special_tokens_map.jsonConfiguration2.7 KB —
talker/ming_talker.pyConfiguration8.8 KB —
talker/sync_vllm_infer.pyConfiguration6.1 KB —
talker/talker_vllm_client.pyConfiguration3.2 KB —
talker/talker_vllm_server.pyConfiguration11.4 KB —
talker/vae/config.jsonConfiguration3.1 KB —
talker/vllm_infer.pyConfiguration7.0 KB —
transformer/config.jsonConfiguration584 B —
transformer/diffusion_pytorch_model.safetensors.index.jsonConfiguration49.0 KB —
vae/config.jsonConfiguration805 B —
README.mdDocumentation12.1 KB —
byt5/google__byt5-smal/README.mdDocumentation4.2 KB —
Ming-Omni_TDS Summary 2026.pdfOther307.8 KB 21cf8e761f19
.gitattributesRepository1.7 KB —
byt5/google__byt5-smal/.gitattributesRepository736 B —
byt5/google__byt5-smal/tokenizer_config.jsonTokenizer2.6 KB —
talker/llm/merges.txtTokenizer1.7 MB —
talker/llm/tokenizer.jsonTokenizer11.4 MB 87d54ca3b44e
talker/llm/tokenizer_config.jsonTokenizer16.3 KB —
talker/llm/vocab.jsonTokenizer2.8 MB —

License and Download

License
mit
Access
Open weights, no gate
Download size
238.0 GB
Download from inclusionAI

Released by inclusionAI through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2506.09344
  • Described by arXiv:2510.24821

Memory Requirements

PrecisionWeights in memory
As published238.0 GB
16-bit208.5 GB
8-bit104.2 GB
4-bit52.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Ming-flash-omni-2.0

How much GPU memory does Ming-flash-omni-2.0 need?

About 250.2 GB at 16-bit and 62.5 GB at 4-bit: the weights (104.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Ming-flash-omni-2.0 on?

At 16-bit, 1x MI325X from $2.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Ming-flash-omni-2.0 commercially?

Yes. Ming-flash-omni-2.0 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is Ming-flash-omni-2.0's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Any to any

Qwen3-Omni-30B-A3B-Instruct

Qwen

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: Qwen3-Omni supports a wide range of multimodal application scenarios, covering various domain tasks involving audio, image, video, and audio-visual modalities. Below are several cookbooks demonstrating the usage cases of Qwen3-Omni and these cookbooks include our actual execution logs. You can first follow the QuickStart guide to download the model and install the necessary inference environment…

Open weights other 35.3B parameters transformers

Model · Any to any

Qwen3-Omni-30B-A3B-FP8

Markus

Block-wise FP8 quantization of Qwen/Qwen3-Omni-30B-A3B-Instruct. - Vision encoder (thinker.visual) - Audio tower (thinker.audiotower) - Code2Wav decoder (code2wav) - vLLM >= 0.13.0 with Qwen3-Omni support - 2x 24GB GPUs (e.g., RTX 3090) or equivalent - ~35 GB disk space Block-wise quantization with 128x128 blocks provides better precision than per-tensor quantization while maintaining good compression. Each block has its own scale factor stored as weightscaleinv (inverse scale for efficient multiplication during inference). This is a quantized version of Qwen/Qwen3-Omni-30B-A3B-Instruct. Qwen3-Omni is a natively end-to-end multilingual omni-modal foundation model that processes text…

Open weights other 35.3B parameters transformers

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 33B parameters 262,144 tokens transformers

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 33B parameters 262,144 tokens transformers

Model · Any to any

Qwen3-Omni-30B-A3B-Thinking

Qwen

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: Qwen3-Omni supports a wide range of multimodal application scenarios, covering various domain tasks involving audio, image, video, and audio-visual modalities. Below are several cookbooks demonstrating the usage cases of Qwen3-Omni and these cookbooks include our actual execution logs. You can first follow the QuickStart guide to download the model and install the necessary inference environment…

Open weights other 31.7B parameters transformers

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 18.3B parameters 262,144 tokens transformers