SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

Qwen3-ASR-1.7B

by Qwen Qwen/Qwen3-ASR-1.7B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects.

Parameters2.3B
Context
Weights4.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads2.4M

Runs On

What it takes to serve Qwen3-ASR-1.7B (2.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 4.7 GB 5.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 2.3 GB 2.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.2 GB 1.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen3-ASR-1.7B

The name says 1.7B, the weight files hold 2.3B parameters, so size by the files. It turns speech into text and identifies which of 52 languages and dialects it hears; a 0.6B sibling shares its Qwen3-Omni foundation, and a separate 0.6B forced aligner handles timestamps. At 16-bit the weights are 4.7 GB and the run needs 5.6 GB, which leaves most of a 192 GB MI300X at $1.85 an hour empty, so scale by instances per GPU, not by GPUs.

Apache 2.0 permits commercial use, modification and redistribution provided the notices stay and changes are stated. Before you commit: no context length is published, no library is named, and the only artifact is safetensors, so 8-bit at 2.8 GB or 4-bit at 1.4 GB is your own quantization work. Released January 28, 2026 and described in arXiv:2601.21337; read the paper before planning around the 52 languages.

Model Card

By Qwen, published under apache-2.0, revision 7278e1e70fe2.

Qwen3-ASR

Overview

Introduction

Read the full model card (3,421 words)

Configuration

Architecture
Qwen3ASRForConditionalGeneration
Model type
qwen3_asr

Identity and Version

Repository
Qwen/Qwen3-ASR-1.7B
Publisher
Qwen
Task
Speech recognition
Modality
Audio
Library
Not stated by the source
Parameters
2.3B parameters
Languages
Not stated by the source
Revision
7278e1e70fe206f11671096ffdd38061171dd6e5
First published
2026-01-28
Last updated
2026-01-30

Files and Weights

12 files, 4.7 GB in total. The weights are 2 files totalling 4.7 GB in safetensors.

Weights2 files · 4.7 GB
Configuration5 files · 72.6 KB
Tokenizer3 files · 4.5 MB
Documentation1 file · 57.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights4.2 GB a4cd1f1a04d9
model-00002-of-00002.safetensorsWeights478.2 MB 6e0b9d9e09e2
chat_template.jsonConfiguration1.2 KB
config.jsonConfiguration6.2 KB
generation_config.jsonConfiguration142 B
model.safetensors.index.jsonConfiguration64.8 KB
preprocessor_config.jsonConfiguration330 B
README.mdDocumentation57.5 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer1.7 MB
tokenizer_config.jsonTokenizer12.5 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.7 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Built From

  • Described by arXiv:2601.21337

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
hf-audio/open-asr-leaderboard Task ami_werMetric ami_werComparison conditions not established 10.56 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task earnings22_werMetric earnings22_werComparison conditions not established 10.25 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task gigaspeech_werMetric gigaspeech_werComparison conditions not established 8.74 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task librispeech_clean_werMetric librispeech_clean_werComparison conditions not established 1.63 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task librispeech_other_werMetric librispeech_other_werComparison conditions not established 3.4 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task mean_werMetric mean_werComparison conditions not established 5.76 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task rtfxMetric rtfxComparison conditions not established 147.93 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task spgispeech_werMetric spgispeech_werComparison conditions not established 2.84 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task tedlium_werMetric tedlium_werComparison conditions not established 2.28 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28
hf-audio/open-asr-leaderboard Task voxpopuli_werMetric voxpopuli_werComparison conditions not established 6.35 open-asr-leaderboard
Reported by a third party
Evaluated revision not stated 2026-01-28

Memory Requirements

PrecisionWeights in memory
As published4.7 GB
16-bit4.7 GB
8-bit2.3 GB
4-bit1.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare Qwen3-ASR-1.7B

Questions About Qwen3-ASR-1.7B

How much GPU memory does Qwen3-ASR-1.7B need?

About 5.6 GB at 16-bit and 1.4 GB at 4-bit: the weights (2.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3-ASR-1.7B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3-ASR-1.7B commercially?

Yes. Qwen3-ASR-1.7B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Speech recognition

seamless-m4t-v2-large

AI at Meta

SeamlessM4T is our foundational all-in-one Massively Multilingual and Multimodal Machine Translation model delivering high-quality translation for speech and text in nearly 100 languages. SeamlessM4T models support the tasks of: - Automatic speech recognition (ASR). - 101 languages for speech input. - 96 Languages for text input/output. - 35 languages for speech output. We are releasing SeamlessM4T v2, an updated version with our novel UnitY2 architecture. This new model improves over SeamlessM4T v1 in quality as well as inference speed in speech generation tasks. The v2 version of SeamlessM4T is a multitask adaptation of our novel UnitY2 architecture. Unity2 with its hierarchical…

Open weights cc-by-nc-4.0 2.3B parameters 4,096 tokens transformers

Model · Speech recognition

cohere-transcribe-03-2026-mlx-4bit

David Larrea

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has been re-validated against the current Swift and Python MLX runtimes. Verified semantic parity on an English fixture: - official CUDA reference path (transformers native Cohere ASR) Fastest and smallest, but introduces a lexical regression on the repo sample (Kaldi → Khaldi). - Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - This repository contains inference artifacts only. Refer to…

Open weights apache-2.0 2.1B parameters mlx

Model · Speech recognition

cohere-transcribe-03-2026-mlx-8bit

David Larrea

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has been re-validated against the current Swift and Python MLX runtimes. Verified semantic parity on an English fixture: - official CUDA reference path (transformers native Cohere ASR) Matches fp16 on the repo sample while reducing memory substantially. - Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - This repository contains inference artifacts only. Refer to the upstream Cohere model…

Open weights apache-2.0 2.1B parameters mlx

Model · Speech recognition

whisper-large-v3

OpenAI

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting. Whisper large-v3 has the same architecture as the previous large and large-v2 models, except for the following minor differences: 1. The spectrogram input uses 128 Mel frequency bins instead of 80 The Whisper large-v3 model was trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio collected using…

Open weights apache-2.0 1.5B parameters transformers

Model · Speech recognition

whisper-ja-1.5B

Efwkjn

For usage instructions follow openai/whisper-large-v3. Large-v3 finetune trained as a baseline with smaller checkpoints in progress. Expecting worse long form and equal short form. Benchmarks. Has occasional repetition issue compared to previous models but achieves competitive/SOTA CER across all tested sets.

Open weights 1.5B parameters

Model · Speech recognition

parakeet-ctc-1.1b

NVIDIA

parakeet-ctc-1.1b is an ASR model that transcribes speech in lower case English alphabet. This model is jointly developed by NVIDIA NeMo and Suno.ai teams. It is an XXL version of FastConformer CTC [1] (around 1.1B parameters) model. See the model architecture section and NeMo documentation for complete architecture details. To train, fine-tune or play with the model you will need to install NVIDIA NeMo. We recommend you install it after you've installed latest PyTorch version. There are several ways to use this model. Choose the one that fits your needs. NeMo-Speech.cpp provides a lightweight native C++ runtime for local inference with this model. After installing the runtime: See the…

Open weights cc-by-4.0 1.1B parameters nemo