SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

cohere-transcribe-03-2026-mlx-4bit

by David Larrea davidalarrea/cohere-transcribe-03-2026-mlx-4bit

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has…

Parameters2.1B
Context
Weights1.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve cohere-transcribe-03-2026-mlx-4bit (2.1B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 4.1 GB 5.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 2.1 GB 2.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.0 GB 1.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By David Larrea, published under apache-2.0, revision fedd1c835713.

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has been re-validated against the current Swift and Python MLX runtimes. Verified semantic parity on an English fixture: - official CUDA reference path (transformers native Cohere ASR) Fastest and smallest, but introduces a lexical regression on the repo sample (Kaldi → Khaldi). - Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - This repository contains inference artifacts only. Refer to…

Read David Larrea's full model card

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16.

Variant

  • Precision: 4-bit
  • Quantization mode: affine
  • Group size: 64

Files

  • model.safetensors
  • config.json
  • tokenizer.model
  • tokenizer_config.json
  • preprocessor_config.json
  • special_tokens_map.json
  • key_map.json
  • conversion_summary.json

Repo-sample benchmark

Sample: Tests/media/conversational_a.wav

  • Generation TPS: 394.6
  • Peak memory: 1.96 GB
  • Output: Coffee's story likely begins in Ethiopia, where legend tells of a goat herder named Khaldi, who noticed his goats became energetic after eating red berries from a particular bush; curious, he tried them himself and felt invigorated.

Parity note

This checkpoint has been re-validated against the current Swift and Python MLX runtimes.

Verified semantic parity on an English fixture:

This is a test recording in English. I am speaking clearly at a normal speed. Please transcribe this sentence exactly as I said.

Matched across:

  • Swift MLX fp16
  • Swift MLX 8-bit
  • Swift MLX 4-bit
  • Python MLX fp16
  • Python MLX 4-bit
  • official CUDA reference path (transformers native Cohere ASR)

Quality note

Fastest and smallest, but introduces a lexical regression on the repo sample (KaldiKhaldi).

Notes

  • Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16.
  • This repository contains inference artifacts only. Refer to the upstream Cohere model card and license for original model details.

Configuration

Architecture
CohereAsrForConditionalGeneration
Vocabulary size
16,384
Model type
cohere_asr

Identity and Version

Repository
davidalarrea/cohere-transcribe-03-2026-mlx-4bit
Publisher
David Larrea
Task
Speech recognition
Modality
Audio
Library
mlx
Parameters
2.1B parameters
Languages
en
Revision
fedd1c83571341374f34525096e5ef67ebcdcfcc
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

10 files, 1.5 GB in total. The weights are 1 file totalling 1.5 GB in safetensors.

Weights1 file · 1.5 GB
Configuration5 files · 193.7 KB
Tokenizer2 files · 541.0 KB
Documentation1 file · 1.9 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.5 GB 5284ab5b678d
config.jsonConfiguration4.3 KB
conversion_summary.jsonConfiguration5.2 KB
key_map.jsonConfiguration179.7 KB
preprocessor_config.jsonConfiguration420 B
special_tokens_map.jsonConfiguration4.1 KB
README.mdDocumentation1.9 KB
.gitattributesRepository1.5 KB
tokenizer.modelTokenizer492.8 KB 6d21e6a83b2d
tokenizer_config.jsonTokenizer48.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.5 GB
Download from David Larrea

Released by David Larrea through its official repository on Hugging Face. Read the license.

Built From

  • Derived from CohereLabs/cohere-transcribe-03-2026
  • Quantized from CohereLabs/cohere-transcribe-03-2026

Memory Requirements

PrecisionWeights in memory
As published1.5 GB
16-bit4.1 GB
8-bit2.1 GB
4-bit1.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About cohere-transcribe-03-2026-mlx-4bit

How much GPU memory does cohere-transcribe-03-2026-mlx-4bit need?

About 5 GB at 16-bit and 1.2 GB at 4-bit: the weights (2.1B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run cohere-transcribe-03-2026-mlx-4bit on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use cohere-transcribe-03-2026-mlx-4bit commercially?

Yes. cohere-transcribe-03-2026-mlx-4bit is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Speech recognition

cohere-transcribe-03-2026-mlx-8bit

David Larrea

Quantized MLX weights for beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - model.safetensors - config.json - tokenizer.model - tokenizerconfig.json - preprocessorconfig.json - specialtokensmap.json - keymap.json - conversionsummary.json This checkpoint has been re-validated against the current Swift and Python MLX runtimes. Verified semantic parity on an English fixture: - official CUDA reference path (transformers native Cohere ASR) Matches fp16 on the repo sample while reducing memory substantially. - Generated from the Swift-compatible fp16 checkpoint beshkenadze/cohere-transcribe-03-2026-mlx-fp16. - This repository contains inference artifacts only. Refer to the upstream Cohere model…

Open weights apache-2.0 2.1B parameters mlx

Model · Speech recognition

seamless-m4t-v2-large

AI at Meta

SeamlessM4T is our foundational all-in-one Massively Multilingual and Multimodal Machine Translation model delivering high-quality translation for speech and text in nearly 100 languages. SeamlessM4T models support the tasks of: - Automatic speech recognition (ASR). - 101 languages for speech input. - 96 Languages for text input/output. - 35 languages for speech output. We are releasing SeamlessM4T v2, an updated version with our novel UnitY2 architecture. This new model improves over SeamlessM4T v1 in quality as well as inference speed in speech generation tasks. The v2 version of SeamlessM4T is a multitask adaptation of our novel UnitY2 architecture. Unity2 with its hierarchical…

Open weights cc-by-nc-4.0 2.3B parameters 4,096 tokens transformers

Model · Speech recognition

Qwen3-ASR-1.7B

Qwen

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features: Novel and strong forced alignment Solution: We introduce Qwen3-ForcedAligner-0.6B, which supports timestamp prediction for arbitrary units within up to 5 minutes of speech in 11 languages. Evaluations show its timestamp…

Open weights apache-2.0 2.3B parameters

Model · Speech recognition

whisper-large-v3

OpenAI

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting. Whisper large-v3 has the same architecture as the previous large and large-v2 models, except for the following minor differences: 1. The spectrogram input uses 128 Mel frequency bins instead of 80 The Whisper large-v3 model was trained on 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio collected using…

Open weights apache-2.0 1.5B parameters transformers

Model · Speech recognition

whisper-ja-1.5B

Efwkjn

For usage instructions follow openai/whisper-large-v3. Large-v3 finetune trained as a baseline with smaller checkpoints in progress. Expecting worse long form and equal short form. Benchmarks. Has occasional repetition issue compared to previous models but achieves competitive/SOTA CER across all tested sets.

Open weights 1.5B parameters

Model · Speech recognition

parakeet-ctc-1.1b

NVIDIA

parakeet-ctc-1.1b is an ASR model that transcribes speech in lower case English alphabet. This model is jointly developed by NVIDIA NeMo and Suno.ai teams. It is an XXL version of FastConformer CTC [1] (around 1.1B parameters) model. See the model architecture section and NeMo documentation for complete architecture details. To train, fine-tune or play with the model you will need to install NVIDIA NeMo. We recommend you install it after you've installed latest PyTorch version. There are several ways to use this model. Choose the one that fits your needs. NeMo-Speech.cpp provides a lightweight native C++ runtime for local inference with this model. After installing the runtime: See the…

Open weights cc-by-4.0 1.1B parameters nemo