SAVRN
Search Contact SAVRN

Open-weight model · Any to any

MiniCPM-o-4_5

by OpenBMB openbmb/MiniCPM-o-4_5

A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full-Duplex Mulitmodal Live Streaming on | CaseBook(Audio, Omni Full-Duplex) MiniCPM-o 4.5 is the latest and most capable model in the MiniCPM-o series.

Parameters9.4B
Context40,960
Weights20.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads693.8k

Runs On

What it takes to serve MiniCPM-o-4_5 (9.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 18.7 GB 22.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 9.4 GB 11.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 4.7 GB 5.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on MiniCPM-o-4_5

We would look at this one when the job involves a camera and a microphone at the same time. OpenBMB assembled it from SigLip2 for vision, Whisper-medium for hearing, CosyVoice2 for speech and Qwen3-8B for language, and the feature they lead with is full-duplex multimodal live streaming. At 16-bit it needs 22.5 GB; 8-bit brings that to 11.2 GB and 4-bit to 5.6 GB. All three fit on one MI300X with 192 GB at $1.85 per hour, so the card decision is how many live sessions you want on it at once.

Apache License 2.0 applies: commercial use, modification and redistribution are permitted, and you keep the notices and state significant changes. Two checks. The context is 40,960 tokens, short for this hub, so plan session length around it. And the paper on file, arXiv:2408.01800, describes MiniCPM-V, so read the 4.5 card for what changed.

Model Card

By OpenBMB, published under apache-2.0, revision 503e754207c9.

A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full-Duplex Mulitmodal Live Streaming on Your Phone

GitHub | CookBook | Omni-modal Demo | Vision-Language Demo WeChat | Discord | MiniCPM Wiki(Chinese) | CaseBook(Audio, Omni Full-Duplex)

News

[!NOTE] [2026.02.06] We open-sourced a realtime web demo deployable on your own devices like Mac or GPU.Try it now!

MiniCPM-o 4.5

MiniCPM-o 4.5 is the latest and most capable model in the MiniCPM-o series. The model is built in an end-to-end fashion based on SigLip2, Whisper-medium, CosyVoice2, and Qwen3-8B with a total of 9B parameters. It exhibits a significant performance improvement, and introduces new features for full-duplex multimodal live streaming. Notable features of MiniCPM-o 4.5 include:

Read the full model card (6,159 words)

Configuration

Architecture
MiniCPMO
Context length (tokens)
40,960
Layers
36
Hidden size
4,096
Feed-forward size
12,288
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
151,748
RoPE base
1,000,000
Stored precision
bfloat16
Model type
minicpmo

Identity and Version

Repository
openbmb/MiniCPM-o-4_5
Publisher
OpenBMB
Task
Any to any
Modality
Multimodal
Library
transformers
Parameters
9.4B parameters
Languages
Not stated by the source
Revision
503e754207c94da6bb26850b4469f367c9ea3582
First published
2026-02-03
Last updated
2026-08-18

Files and Weights

54 files, 20.0 GB in total. The weights are 8 files totalling 20.0 GB in onnx, pt, safetensors.

Weights8 files · 20.0 GB
Configuration13 files · 573.8 KB
Tokenizer4 files · 15.7 MB
Documentation1 file · 85.8 KB
Other26 files · 57.3 MB
Repository2 files · 3.1 KB
Every file
FileTypeSizeSHA-256
assets/token2wav/campplus.onnxWeights28.3 MB a6ac6a639977
assets/token2wav/flow.ptWeights623.5 MB 15ccff24256f
assets/token2wav/hift.ptWeights83.4 MB 3386cc880324
assets/token2wav/speech_tokenizer_v2_25hz.onnxWeights496.1 MB d43342aa1216
model-00001-of-00004.safetensorsWeights5.3 GB 30c40b9a1038
model-00002-of-00004.safetensorsWeights5.3 GB fe0faef420ac
model-00003-of-00004.safetensorsWeights5.3 GB 5d0b20153f9b
model-00004-of-00004.safetensorsWeights2.9 GB f61addf4747c
added_tokens.jsonConfiguration2.7 KB
assets/token2wav/flow.yamlConfiguration1.1 KB
config.jsonConfiguration6.4 KB
configuration_minicpmo.pyConfiguration10.1 KB
generation_config.jsonConfiguration178 B
model.safetensors.index.jsonConfiguration117.2 KB
modeling_minicpmo.pyConfiguration214.4 KB
modeling_navit_siglip.pyConfiguration42.6 KB
preprocessor_config.jsonConfiguration842 B
processing_minicpmo.pyConfiguration67.3 KB
special_tokens_map.jsonConfiguration12.0 KB
tokenization_minicpmo_fast.pyConfiguration3.2 KB
utils.pyConfiguration95.7 KB
README.mdDocumentation85.8 KB
assets/HT_ref_audio.wavOther192.6 KB cb8f06ba5080
assets/Skiing.mp4Other8.5 MB 479ace116d6a
assets/Trump_WEF_2018_10s.mp3Other161.1 KB 4fb796c2bb95
assets/audio_cases/assistant_ref.mp4Other65.5 KB 7e4a56e44187
assets/audio_cases/assistant_response.mp4Other269.5 KB d46268e3beb7
assets/audio_cases/elon_musk__000_assistant_audio.wavOther1.8 MB ff6dced2e686
assets/audio_cases/elon_musk__system_ref_audio.wavOther539.0 KB 2c4109b2d685
assets/audio_cases/elon_musk_ref.mp4Other165.2 KB 206c48ec4b08
assets/audio_cases/elon_musk_response.mp4Other388.5 KB d83c3a977fd4
assets/audio_cases/hermione__000_assistant_audio.wavOther1.5 MB cfd18a2ee9e2
assets/audio_cases/hermione__system_ref_audio.wavOther197.3 KB 46bd82796ce5
assets/audio_cases/minicpm_assistant__000_assistant_audio.wavOther1.2 MB fe3b793a6436
assets/audio_cases/minicpm_assistant__system_ref_audio.wavOther192.6 KB ad576b50fd2f
assets/audio_cases/paimon__000_assistant_audio.wavOther697.0 KB 9ddd9d835895
assets/audio_cases/paimon__system_ref_audio.wavOther479.3 KB b9625e7115ff
assets/audio_cases/readme.txtOther57 B
assets/bajie.wavOther636.5 KB 16aa8ca3da7d
assets/fossil.pngOther465.9 KB b8b3f1668da6
assets/haimianbaobao.wavOther343.1 KB 27405cf5977f
assets/highway.pngOther840.9 KB 87c32da6ee77
assets/nezha.wavOther457.8 KB fe5d8932013f
assets/omni_duplex1.mp4Other7.3 MB 31622e1efd9a
assets/omni_duplex2.mp4Other29.3 MB c04eaef27a82
assets/sunwukong.wavOther644.6 KB 3bdb6c175bd3
assets/system_ref_audio.wavOther539.0 KB 2c4109b2d685
assets/system_ref_audio_2.wavOther341.2 KB 4a65d1709923
.gitattributesRepository3.1 KB
.gitignoreRepository9 B
merges.txtTokenizer1.4 MB
tokenizer.jsonTokenizer11.4 MB 6d55eb34389b
tokenizer_config.jsonTokenizer101.7 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
20.0 GB
Download from OpenBMB

Released by OpenBMB through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published20.0 GB
16-bit18.7 GB
8-bit9.4 GB
4-bit4.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Compare MiniCPM-o-4_5

Questions About MiniCPM-o-4_5

How much GPU memory does MiniCPM-o-4_5 need?

About 22.5 GB at 16-bit and 5.6 GB at 4-bit: the weights (9.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run MiniCPM-o-4_5 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use MiniCPM-o-4_5 commercially?

Yes. MiniCPM-o-4_5 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is MiniCPM-o-4_5's context length?

40,960 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Any to any

MiniCPM-o-4_5-awq

OpenBMB

A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full-Duplex Mulitmodal Live Streaming on | CaseBook(Audio, Omni Full-Duplex) MiniCPM-o 4.5 is the latest and most capable model in the MiniCPM-o series. The model is built in an end-to-end fashion based on SigLip2, Whisper-medium, CosyVoice2, and Qwen3-8B with a total of 9B parameters. It exhibits a significant performance improvement, and introduces new features for full-duplex multimodal live streaming. Notable features of MiniCPM-o 4.5 include: - Leading Visual Capability. MiniCPM-o 4.5 achieves an average score of 77.6 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 9B parameters, it surpasses widely…

Open weights apache-2.0 9.4B parameters 40,960 tokens transformers

Model · Any to any

gemma-4-E4B-it-qat-w4a16-ct

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 8.7B parameters 131,072 tokens transformers

Model · Any to any

MiniCPM-o-2_6

OpenBMB

[2025.06.20] Our official ollama repository is released. Try our latest models with one click! [2025.03.01] RLAIF-V, which is the alignment technique of MiniCPM-o, is accepted by CVPR 2025!The code, dataset, paper are open-sourced! [2025.01.24] MiniCPM-o 2.6 technical report is released! See Here. [2025.01.19] MiniCPM-o tops GitHub Trending and reaches top-2 on Hugging Face Trending! MiniCPM-o 2.6 is the latest and most capable model in the MiniCPM-o series. The model is built in an end-to-end fashion based on SigLip-400M, Whisper-medium-300M, ChatTTS-200M, and Qwen2.5-7B with a total of 8B parameters. It exhibits a significant performance improvement over MiniCPM-V 2.6, and introduces new…

Open weights apache-2.0 8.7B parameters 32,768 tokens transformers

Model · Any to any

Qwen2.5-Omni-7B

Qwen

Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. We conducted a comprehensive evaluation of Qwen2.5-Omni, which demonstrates strong performance across all modalities when compared to similarly sized single-modality models and closed-source models like Qwen2.5-VL-7B, Qwen2-Audio, and Gemini-1.5-pro. In tasks requiring the integration of multiple modalities, such as OmniBench, Qwen2.5-Omni achieves state-of-the-art performance. Furthermore, in single-modality tasks, it excels in areas including speech recognition (Common…

Open weights other 10.7B parameters transformers

Model · Any to any

Qwen2.5-Omni-7B-AWQ

Qwen

Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. This model card introduces a series of enhancements designed to improve the Qwen2.5-Omni-7B's operability on devices with constrained GPU memory. Key optimizations include: Implemented 4-bit quantization of the Thinker's weights using AWQ, effectively reducing GPU VRAM usage. Enhanced the inference pipeline to load model weights on-demand for each module and offload them to CPU memory once inference is complete, preventing peak VRAM usage from becoming excessive. Converted…

Open weights other 10.7B parameters transformers

Model · Any to any

gemma-4-E4B-it

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 8B parameters 131,072 tokens transformers