# Ming-flash-omni-2.0 by inclusionAI: Open-Weight Model
Source: https://savrn.com/models/ming-flash-omni-2-0
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Ming-flash-omni-2.0 (104.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 208.5 GB | 250.2 GB | 1x [MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) (256 GB) Vultr | $2.00 | [1x MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) $2.59 · [2x MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) $3.70 |
| 8-bit | 104.2 GB | 125.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 · [1x MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) $2.59 |
| 4-bit | 52.1 GB | 62.5 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[Ming-flash-omni-2.0 on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/ming-flash-omni-2-0/gpus)

## Model Card

By inclusionAI, published under mit, revision c1ed5565e64e.

Technical Report ｜ Hugging Face ｜ ModelScope ## Introduction The newly released Ming-flash-omni 2.0 leverages the [Ling-2.0](https://github.com/inclusionAI/Ling-V2) architecture—a Mixture-of-Experts (MoE) framework comprising 100B total and 6B active parameters. Representing a generational advancement over its predecessor, it establishes new State-of-the-Art (SOTA) benchmarks among open-source omni-MLLMs. Ming-flash-omni 2.0 effectively synergizes foundational abilities with specialized domain expertise. In particular, it exhibits superior performance in visual encyclopedic knowledge, immersive speech synthesis, and high-dynamic image generation and manipulation. ## Updates * [2026.02.11] We release the official version of [Ming-flash-omni 2.0](https://mp.weixin.qq.com/s/hz2fsH1DGpp2zpY-Yngsog), an open-source SOTA omni-MLLM that pushes the boundaries of multimodal understanding and synthesis. * [2025.10.27] We release the preview version of Ming-flash-omni：[Ming-flash-omni Preview](https://github.com/inclusionAI/Ming/tree/main). * [2025.07.15] We release [Ming-lite-omni v1.5](https://github.com/inclusionAI/Ming/tree/v1.5) with significant improvements across all modalities. * [2025.06.12] Our [Technical Report](https://arxiv.org/abs/2506.09344) is in public on arxiv. * [2025.05.28] The official version of [Ming-lite-omni v1](https://github.com/inclusionAI/Ming/tree/v1.0) is released, with better performance and image generation support. * [2025.05.04] We release the test version of Ming-lite-omni：[Ming-lite-omni-Preview](https://github.com/inclusionAI/Ming/tree/Ming-Lite-Omni-Preview). ## Key Features Compared…

[Read the full model card (879 words)](https://savrn.com/models/ming-flash-omni-2-0/card)

## Configuration

Architecture

BailingMM2NativeForConditionalGeneration

Context length (tokens)

32,768

Layers

32

Hidden size

4,096

Feed-forward size

9,216

Attention heads

32

Key/value heads

4

Vocabulary size

157,184

Experts

256

Experts active per token

8

RoPE base

2,400,000

Stored precision

bfloat16

Model type

bailingmm_moe_v2_lite

## Identity and Version

Repository

inclusionAI/Ming-flash-omni-2.0

Publisher

inclusionAI

Task

Any to any

Modality

Multimodal

Library

diffusers

Parameters

104.2B parameters

Languages

en

Revision

c1ed5565e64e85a76826255c000b66a20d4634c3

First published

2026-02-10

Last updated

2026-09-29

## Files and Weights

99 files, 238.0 GB in total. The weights are 60 files totalling 238.0 GB in bin, h5, msgpack, onnx, pt, safetensors.

Weights60 files · 238.0 GB

Configuration29 files · 2.6 MB

Tokenizer5 files · 15.9 MB

Documentation2 files · 16.3 KB

Other1 file · 307.8 KB

Repository2 files · 2.4 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| byt5/byt5_mapper/byt5_mapper.pt | Weights | 304.6 MB | 74a3cba1ce71 |
| byt5/byt5_model/base.pt | Weights | 3.0 GB | 5f0816c04e12 |
| byt5/byt5_model/byt5_model.pt | Weights | 877.3 MB | ca8c97c89136 |
| byt5/google__byt5-smal/flax_model.msgpack | Weights | 1.2 GB | b3aafee96d60 |
| byt5/google__byt5-smal/pytorch_model.bin | Weights | 1.2 GB | 5c5aaf56299d |
| byt5/google__byt5-smal/tf_model.h5 | Weights | 1.2 GB | f97320dd5eb4 |
| connector/model-00001-of-00002.safetensors | Weights | 5.0 GB | c021bcc12289 |
| connector/model-00002-of-00002.safetensors | Weights | 1.2 GB | f49049386ed6 |
| mlp/model.safetensors | Weights | 45.1 MB | 22997acfa6b9 |
| model-00001-of-00042.safetensors | Weights | 5.0 GB | 3edf1c3a549c |
| model-00002-of-00042.safetensors | Weights | 5.0 GB | 435534ce3b2a |
| model-00003-of-00042.safetensors | Weights | 5.0 GB | 8c6974eced77 |
| model-00004-of-00042.safetensors | Weights | 5.0 GB | 013ed1e2b926 |
| model-00005-of-00042.safetensors | Weights | 5.0 GB | 22cd341b7925 |
| model-00006-of-00042.safetensors | Weights | 5.0 GB | 9bb1d06e4132 |
| model-00007-of-00042.safetensors | Weights | 5.0 GB | 564d61b480e2 |
| model-00008-of-00042.safetensors | Weights | 5.0 GB | 3b3fecdd1b22 |
| model-00009-of-00042.safetensors | Weights | 5.0 GB | 2f14a4fa5515 |
| model-00010-of-00042.safetensors | Weights | 5.0 GB | 9bf9683ee56e |
| model-00011-of-00042.safetensors | Weights | 5.0 GB | 4275c8966395 |
| model-00012-of-00042.safetensors | Weights | 5.0 GB | 6970d11a6dfc |
| model-00013-of-00042.safetensors | Weights | 5.0 GB | e62f403a2a6c |
| model-00014-of-00042.safetensors | Weights | 5.0 GB | 2e52e582facb |
| model-00015-of-00042.safetensors | Weights | 5.0 GB | 00da6658de16 |
| model-00016-of-00042.safetensors | Weights | 5.0 GB | 4499f51c9289 |
| model-00017-of-00042.safetensors | Weights | 5.0 GB | 9ca2cb3052b8 |
| model-00018-of-00042.safetensors | Weights | 5.0 GB | 2257ab127438 |
| model-00019-of-00042.safetensors | Weights | 5.0 GB | af19341964d4 |
| model-00020-of-00042.safetensors | Weights | 5.0 GB | 74fe06f8ac5b |
| model-00021-of-00042.safetensors | Weights | 5.0 GB | 7440a749d924 |
| model-00022-of-00042.safetensors | Weights | 5.0 GB | bcaaa83a9c78 |
| model-00023-of-00042.safetensors | Weights | 5.0 GB | c3a3bd0495a8 |
| model-00024-of-00042.safetensors | Weights | 5.0 GB | b0a35bc3f498 |
| model-00025-of-00042.safetensors | Weights | 5.0 GB | dc97c036c59b |
| model-00026-of-00042.safetensors | Weights | 5.0 GB | 4489dc484bb6 |
| model-00027-of-00042.safetensors | Weights | 5.0 GB | 5e9be0110cfd |
| model-00028-of-00042.safetensors | Weights | 5.0 GB | 5dbc75608840 |
| model-00029-of-00042.safetensors | Weights | 5.0 GB | f69f641d1535 |
| model-00030-of-00042.safetensors | Weights | 5.0 GB | 66279b01fb41 |
| model-00031-of-00042.safetensors | Weights | 5.0 GB | a332d936b24e |
| model-00032-of-00042.safetensors | Weights | 5.0 GB | f07ddcad0bd4 |
| model-00033-of-00042.safetensors | Weights | 5.0 GB | c71cb58bc686 |
| model-00034-of-00042.safetensors | Weights | 5.0 GB | bf71c915534f |
| model-00035-of-00042.safetensors | Weights | 5.0 GB | af1414dd1120 |
| model-00036-of-00042.safetensors | Weights | 5.0 GB | 6553a3359a0b |
| model-00037-of-00042.safetensors | Weights | 5.0 GB | 67f7c432429c |
| model-00038-of-00042.safetensors | Weights | 5.0 GB | ef52ac1e1642 |
| model-00039-of-00042.safetensors | Weights | 5.0 GB | b746d362a232 |
| model-00040-of-00042.safetensors | Weights | 5.0 GB | 3fdb2533d0f9 |
| model-00041-of-00042.safetensors | Weights | 5.0 GB | dcf4415df371 |
| model-00042-of-00042.safetensors | Weights | 3.6 GB | 06386221000f |
| talker/campplus.onnx | Weights | 28.3 MB | a6ac6a639977 |
| talker/model.safetensors | Weights | 1.4 GB | 1403c23a0a91 |
| talker/vae/model.safetensors | Weights | 1.6 GB | 61ffb5e9ca7e |
| transformer/diffusion_pytorch_model-00001-of-00005.safetensors | Weights | 3.0 GB | ff08b8209825 |
| transformer/diffusion_pytorch_model-00002-of-00005.safetensors | Weights | 2.9 GB | 76abee2c7c2c |
| transformer/diffusion_pytorch_model-00003-of-00005.safetensors | Weights | 3.0 GB | 9c8f4a400f3e |
| transformer/diffusion_pytorch_model-00004-of-00005.safetensors | Weights | 3.0 GB | 688c6ef5fd85 |
| transformer/diffusion_pytorch_model-00005-of-00005.safetensors | Weights | 448.4 MB | a45e05dc23c3 |
| vae/diffusion_pytorch_model.safetensors | Weights | 167.7 MB | f5b59a268515 |
| byt5/byt5.json | Configuration | 491 B | — |
| byt5/color_idx.json | Configuration | 2.4 KB | — |
| byt5/font_uni_10-lang_idx.json | Configuration | 24.3 KB | — |
| byt5/google__byt5-smal/config.json | Configuration | 698 B | — |
| byt5/google__byt5-smal/generation_config.json | Configuration | 147 B | — |
| byt5/google__byt5-smal/special_tokens_map.json | Configuration | 2.5 KB | — |
| config.json | Configuration | 4.7 KB | — |
| connector/config.json | Configuration | 1.3 KB | — |
| connector/generation_config.json | Configuration | 242 B | — |
| connector/model.safetensors.index.json | Configuration | 27.7 KB | — |
| mlp/config.json | Configuration | 166 B | — |
| model.safetensors.index.json | Configuration | 2.4 MB | — |
| scheduler/scheduler_config.json | Configuration | 173 B | — |
| talker/__init__.py | Configuration |  | — |
| talker/async_vllm_infer.py | Configuration | 6.9 KB | — |
| talker/config.json | Configuration | 707 B | — |
| talker/llm/added_tokens.json | Configuration | 1.8 KB | — |
| talker/llm/config.json | Configuration | 659 B | — |
| talker/llm/generation_config.json | Configuration | 242 B | — |
| talker/llm/special_tokens_map.json | Configuration | 2.7 KB | — |
| talker/ming_talker.py | Configuration | 8.8 KB | — |
| talker/sync_vllm_infer.py | Configuration | 6.1 KB | — |
| talker/talker_vllm_client.py | Configuration | 3.2 KB | — |
| talker/talker_vllm_server.py | Configuration | 11.4 KB | — |
| talker/vae/config.json | Configuration | 3.1 KB | — |
| talker/vllm_infer.py | Configuration | 7.0 KB | — |
| transformer/config.json | Configuration | 584 B | — |
| transformer/diffusion_pytorch_model.safetensors.index.json | Configuration | 49.0 KB | — |
| vae/config.json | Configuration | 805 B | — |
| README.md | Documentation | 12.1 KB | — |
| byt5/google__byt5-smal/README.md | Documentation | 4.2 KB | — |
| Ming-Omni_TDS Summary 2026.pdf | Other | 307.8 KB | 21cf8e761f19 |
| .gitattributes | Repository | 1.7 KB | — |
| byt5/google__byt5-smal/.gitattributes | Repository | 736 B | — |
| byt5/google__byt5-smal/tokenizer_config.json | Tokenizer | 2.6 KB | — |
| talker/llm/merges.txt | Tokenizer | 1.7 MB | — |
| talker/llm/tokenizer.json | Tokenizer | 11.4 MB | 87d54ca3b44e |
| talker/llm/tokenizer_config.json | Tokenizer | 16.3 KB | — |
| talker/llm/vocab.json | Tokenizer | 2.8 MB | — |

## License and Download

License

mit

Access

Open weights, no gate

Download size

238.0 GB

[Download from inclusionAI](https://huggingface.co/inclusionAI/Ming-flash-omni-2.0)

Released by inclusionAI through its official repository on Hugging Face. [Read the license](https://opensource.org/license/mit).

## Built From

- Described by arXiv:2506.09344
- Described by arXiv:2510.24821

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 238.0 GB |
| 16-bit | 208.5 GB |
| 8-bit | 104.2 GB |
| 4-bit | 52.1 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Ming-flash-omni-2.0

### How much GPU memory does Ming-flash-omni-2.0 need?

About 250.2 GB at 16-bit and 62.5 GB at 4-bit: the weights (104.2B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Ming-flash-omni-2.0 on?

At 16-bit, 1x MI325X from $2.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use Ming-flash-omni-2.0 commercially?

Yes. Ming-flash-omni-2.0 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

### What is Ming-flash-omni-2.0's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Any to any

### [Qwen3-Omni-30B-A3B-Instruct](https://savrn.com/models/qwen3-omni-30b-a3b-instruct)

[Qwen](https://savrn.com/model-publishers/qwen)

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: Qwen3-Omni supports a wide range of multimodal application scenarios, covering various domain tasks involving audio, image, video, and audio-visual modalities. Below are several cookbooks demonstrating the usage cases of Qwen3-Omni and these cookbooks include our actual execution logs. You can first follow the QuickStart guide to download the model and install the necessary inference environment…

Open weights other 35.3B parameters transformers

[View model](https://savrn.com/models/qwen3-omni-30b-a3b-instruct)

Model · Any to any

### [Qwen3-Omni-30B-A3B-FP8](https://savrn.com/models/qwen3-omni-30b-a3b-fp8)

[Markus](https://savrn.com/model-publishers/marksverdhei)

Block-wise FP8 quantization of Qwen/Qwen3-Omni-30B-A3B-Instruct. - Vision encoder (thinker.visual) - Audio tower (thinker.audiotower) - Code2Wav decoder (code2wav) - vLLM >= 0.13.0 with Qwen3-Omni support - 2x 24GB GPUs (e.g., RTX 3090) or equivalent - ~35 GB disk space Block-wise quantization with 128x128 blocks provides better precision than per-tensor quantization while maintaining good compression. Each block has its own scale factor stored as weightscaleinv (inverse scale for efficient multiplication during inference). This is a quantized version of Qwen/Qwen3-Omni-30B-A3B-Instruct. Qwen3-Omni is a natively end-to-end multilingual omni-modal foundation model that processes text…

Open weights other 35.3B parameters transformers

[View model](https://savrn.com/models/qwen3-omni-30b-a3b-fp8)

Model · Any to any

### [Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16](https://savrn.com/models/nemotron-3-nano-omni-30b-a3b-reasoning-bf16)

[NVIDIA](https://savrn.com/model-publishers/nvidia)

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 33B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/nemotron-3-nano-omni-30b-a3b-reasoning-bf16)

Model · Any to any

### [Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8](https://savrn.com/models/nemotron-3-nano-omni-30b-a3b-reasoning-fp8)

[NVIDIA](https://savrn.com/model-publishers/nvidia)

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 33B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/nemotron-3-nano-omni-30b-a3b-reasoning-fp8)

Model · Any to any

### [Qwen3-Omni-30B-A3B-Thinking](https://savrn.com/models/qwen3-omni-30b-a3b-thinking)

[Qwen](https://savrn.com/model-publishers/qwen)

Qwen3-Omni is the natively end-to-end multilingual omni-modal foundation models. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: Qwen3-Omni supports a wide range of multimodal application scenarios, covering various domain tasks involving audio, image, video, and audio-visual modalities. Below are several cookbooks demonstrating the usage cases of Qwen3-Omni and these cookbooks include our actual execution logs. You can first follow the QuickStart guide to download the model and install the necessary inference environment…

Open weights other 31.7B parameters transformers

[View model](https://savrn.com/models/qwen3-omni-30b-a3b-thinking)

Model · Any to any

### [Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4](https://savrn.com/models/nemotron-3-nano-omni-30b-a3b-reasoning-nvfp4)

[NVIDIA](https://savrn.com/model-publishers/nvidia)

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family. This model is available for commercial use. This…

Open weights other 18.3B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/nemotron-3-nano-omni-30b-a3b-reasoning-nvfp4)

## inclusionAI

[All models and datasets](https://savrn.com/model-publishers/inclusionai)

## Versions

- [c1ed5565e64e](https://savrn.com/models/ming-flash-omni-2-0/versions/c1ed5565e64e) · current 2026-09-29

## Explore More

- [All any to any models](https://savrn.com/models/tasks/any-to-any)
- [All models under mit](https://savrn.com/models/licenses/mit)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-09-29.
- [Hugging Face record](https://huggingface.co/inclusionAI/Ming-flash-omni-2.0)
- [How the hub is built](https://savrn.com/model-hub/methodology)
