SAVRN's Take
The label on this page covers models that accept more than one kind of input. Mostly that means Google's Gemma 4 line, which takes text and images, adds audio on the E2B, E4B and 12B sizes, and answers in text, with a context of up to 256K tokens and more than 140 languages. Qwen3-Omni-30B-A3B-Instruct goes further, reading text, images, audio and video and streaming back both text and speech. Sixty-one models sit here, 20 with Index pricing.
At the bottom, gemma-4-E4B-it-assistant at 79M parameters wants 0.2 GB at 16-bit. The most downloaded, gemma-4-E4B-it at 8B parameters and 4,473,231 pulls a month, needs 19.2 GB at 16-bit and 4.8 GB at 4-bit. The 12B, with a 262,144 token context where the E-series stops at 131,072, climbs to 28.7 GB. Qwen3-Omni tops the range at 35.3B parameters and 84.6 GB at 16-bit, or 21.2 GB at 4-bit. Every indexed entry lands on a single MI300X, priced on the Index at $1.85 an hour, so the hardware question is how much of one card you leave for the KV cache when the context runs long.
Google publishes 16 of the 61 and LM Studio Community 12; the top GGUF and MLX repacks from GGML Org, LM Studio and Unsloth AI each clear a million downloads a month. Forty-eight carry Apache 2.0, publisher weights and repacks alike. Ten are marked other, Qwen3-Omni among them, so read those terms first. Two state no license; we would not deploy those. Decide on audio in, speech out and context length before you commit.
SAVRN Research, 2026-09-18