SAVRN
Search Contact SAVRN

Open-weight model · Text to speech

Kokoro-82M

by Hexgrad hexgrad/Kokoro-82M

Kokoro is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient.

Parameters
Context
Weights355.5 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads11.7M

SAVRN's Notes on Kokoro-82M

Speech output is the whole job here: text in, audio out, from 82 million parameters that Hexgrad ships as 355 MB of weights in a 363 MB package. We have no runs-on row for it yet, so that weight file is the only sizing number, and a file that small does not decide a hardware purchase; the question is what else shares the card. It also needs the misaki G2P library alongside.

Apache 2.0 permits commercial use, modification and redistribution, provided the license and copyright notices and any NOTICE file stay with it and significant changes are stated; contributors grant patent rights expressly. Read the lineage before committing: the weights derive from yl4579/StyleTTS2-LJSpeech, with arXiv:2306.07691 and arXiv:2203.02395 cited for the method. The last update landed 2025-04-10 on a 2024-12-26 release. Hexgrad publishes a SHA256 hash for the model file; check your copy against it.

Model Card

By Hexgrad, published under apache-2.0, revision f3ff3571791e.

Kokoro is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.

GitHub: https://github.com/hexgrad/kokoro

Demo: https://hf.co/spaces/hexgrad/Kokoro-TTS

Read the full model card (619 words)

Identity and Version

Repository
hexgrad/Kokoro-82M
Publisher
Hexgrad
Task
Text to speech
Modality
Audio
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en
Revision
f3ff3571791e39611d31c381e3a41a3af07b4987
First published
2024-12-26
Last updated
2025-04-10

Files and Weights

72 files, 363.3 MB in total. The weights are 55 files totalling 355.5 MB in pt, pth.

Weights55 files · 355.5 MB
Configuration1 file · 2.4 KB
Documentation5 files · 23.0 KB
Other10 files · 7.8 MB
Repository1 file · 1.9 KB
Every file
FileTypeSizeSHA-256
kokoro-v1_0.pthWeights327.2 MB 496dba118d1a
voices/af_alloy.ptWeights523.4 KB 6d877149dd8b
voices/af_aoede.ptWeights523.4 KB c03bd1a4c371
voices/af_bella.ptWeights523.4 KB 8cb64e02fcc8
voices/af_heart.ptWeights523.4 KB 0ab5709b8ffa
voices/af_jessica.ptWeights523.4 KB cdfdccb8cc97
voices/af_kore.ptWeights523.4 KB 8bfbc512321c
voices/af_nicole.ptWeights523.4 KB c5561808bcf5
voices/af_nova.ptWeights523.4 KB e0233676ddc2
voices/af_river.ptWeights523.4 KB e149459bd9c0
voices/af_sarah.ptWeights523.4 KB 49bd364ea3be
voices/af_sky.ptWeights523.4 KB c799548aed06
voices/am_adam.ptWeights523.4 KB ced7e284aba1
voices/am_echo.ptWeights523.4 KB 8bcfdc852bc9
voices/am_eric.ptWeights523.4 KB ada66f0eefff
voices/am_fenrir.ptWeights523.4 KB 98e507eca1db
voices/am_liam.ptWeights523.4 KB c82550757ddb
voices/am_michael.ptWeights523.4 KB 9a443b79a4b2
voices/am_onyx.ptWeights523.4 KB e8452be16cd0
voices/am_puck.ptWeights523.4 KB dd1d8973f4ce
voices/am_santa.ptWeights523.4 KB 7f2f7582fa2b
voices/bf_alice.ptWeights523.4 KB d292651b6af6
voices/bf_emma.ptWeights523.4 KB d0a423deabf4
voices/bf_isabella.ptWeights523.4 KB cdd4c3700380
voices/bf_lily.ptWeights523.4 KB 6e09c2e481e2
voices/bm_daniel.ptWeights523.4 KB fc3fce4e9c12
voices/bm_fable.ptWeights523.4 KB d44935f31352
voices/bm_george.ptWeights523.4 KB f1bc812213dc
voices/bm_lewis.ptWeights523.4 KB b5204750dcba
voices/ef_dora.ptWeights523.4 KB d9d69b0f8a2b
voices/em_alex.ptWeights523.4 KB 5eac53f767c3
voices/em_santa.ptWeights523.4 KB aa8620cb96ce
voices/ff_siwis.ptWeights523.4 KB 8073bf2d2c4b
voices/hf_alpha.ptWeights523.4 KB 06906fe05746
voices/hf_beta.ptWeights523.4 KB 63c0a1a6272e
voices/hm_omega.ptWeights523.4 KB b55f02a8e848
voices/hm_psi.ptWeights523.4 KB 2f0f055cea4f
voices/if_sara.ptWeights523.4 KB 6c0b253b955f
voices/im_nicola.ptWeights523.3 KB 234ed0664864
voices/jf_alpha.ptWeights523.4 KB 1bf4c9dc69e4
voices/jf_gongitsune.ptWeights523.4 KB 1b171917f18f
voices/jf_nezumi.ptWeights523.4 KB d83f007a7f01
voices/jf_tebukuro.ptWeights523.4 KB 0d6917904438
voices/jm_kumo.ptWeights523.4 KB 98340afd68b1
voices/pf_dora.ptWeights523.4 KB 07e4ff987c5d
voices/pm_alex.ptWeights523.4 KB cf0ba8c573c2
voices/pm_santa.ptWeights523.4 KB d42103169c5c
voices/zf_xiaobei.ptWeights523.4 KB 9b76be63dab4
voices/zf_xiaoni.ptWeights523.4 KB 95b49f169bf1
voices/zf_xiaoxiao.ptWeights523.4 KB cfaf6f2ded1e
voices/zf_xiaoyi.ptWeights523.4 KB b5235dbaeef8
voices/zm_yunjian.ptWeights523.4 KB 76cbf8bad359
voices/zm_yunxi.ptWeights523.4 KB dbe6e1ce7c3d
voices/zm_yunxia.ptWeights523.4 KB bb2b03b08e84
voices/zm_yunyang.ptWeights523.4 KB 5238ac22e0c7
config.jsonConfiguration2.4 KB
DONATE.mdDocumentation2.6 KB
EVAL.mdDocumentation534 B
README.mdDocumentation6.3 KB
SAMPLES.mdDocumentation6.0 KB
VOICES.mdDocumentation7.6 KB
eval/ArtificialAnalysis-2025-02-26.jpegOther939.5 KB a312447a7801
eval/TTS_Arena-2025-02-26.jpegOther560.1 KB f2caf29f9ad1
eval/TTS_Spaces_Arena-2025-02-26.jpegOther515.3 KB b095b1048199
samples/HEARME.wavOther996.0 KB
samples/af_heart_0.wavOther237.6 KB
samples/af_heart_1.wavOther517.2 KB
samples/af_heart_2.wavOther496.8 KB
samples/af_heart_3.wavOther1.4 MB e758efcd852e
samples/af_heart_4.wavOther1.1 MB d50c90f44768
samples/af_heart_5.wavOther1.0 MB bc4515b9479c
.gitattributesRepository1.9 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
355.5 MB
Download from Hexgrad

Released by Hexgrad through its official repository on Hugging Face. Read the license.

Built From

  • Derived from yl4579/StyleTTS2-LJSpeech
  • Described by arXiv:2203.02395
  • Described by arXiv:2306.07691

Memory Requirements

PrecisionWeights in memory
As published355.5 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Kokoro-82M

Can I use Kokoro-82M commercially?

Yes. Kokoro-82M is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to speech

XTTS-v2

Coqui.ai

ⓍTTS is a Voice generation model that lets you clone voices into different languages by using just a quick 6-second audio clip. There is no need for an excessive amount of training data that spans countless hours. This is the same or similar model to what powers Coqui Studio and Coqui API. - Supports 17 languages. - Voice cloning with just a 6-second audio clip. - Emotion and style transfer by cloning. - Cross-language voice cloning. - Multi-lingual speech generation. - 24khz sampling rate. - 2 new languages; Hungarian and Korean - Architectural improvements for speaker conditioning. - Enables the use of multiple speaker references and interpolation between speakers. - Stability…

Open weights other coqui

Model · Text to speech

audio.cpp-gguf

Audio.cpp

This directory contains audio.cpp-native GGUF conversions of multiple speech models. These files are intended for use with audio.cpp. If you enjoy the project, please star audio.cpp on GitHub and this Hugging Face repository. For conversion details, supported layouts, direct-file loading, sidecar embedding, and the latest compatibility notes, see the audio.cpp GGUF guide: - https://github.com/0xShug0/audio.cpp/blob/main/docs/gguf.md!!! Converted and quantized packages are checked with automated metrics, but perceived quality can still differ for human listeners. Please validate the exact package, backend, and route to confirm the output is acceptable for your use case. The table lists the…

Open weights other audio.cpp

Model · Text to speech

chatterbox

Resemble AI

Chatterbox Multilingual V3 is the latest general-purpose multilingual TTS model in the Chatterbox family. It keeps the same 0.5B model size while improving speaker similarity, reducing hallucinations, and producing more natural, conversational speech across languages. V3 is designed for broad language coverage like V2, but with stronger stability and more expressive generation. It is the recommended multilingual model for users who want one voice cloning model that works across many languages. Try it in the Chatterbox Multilingual TTS V3 Space. Alongside V3, we are releasing the Single Language Pack: dedicated finetunes for priority languages where tighter quality control, stronger…

Open weights mit chatterbox

Model · Text to speech

F5-TTS

Yushen CHEN

Download F5-TTS or E2 TTS and place under ckpts/ Paper: F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Open weights cc-by-nc-4.0 f5-tts

Model · Text to speech

Kokoro-82M-v1.0-ONNX

ONNX Community

Kokoro is a frontier TTS model for its size of 82 million parameters (text in/audio out). First, install the kokoro-js library from NPM using: You can then generate speech as follows: Optionally, save the audio to a file: The model is resilient to quantization, enabling efficient high-quality speech synthesis at a fraction of the original model size.

Open weights apache-2.0 transformers.js

Model · Text to speech

Qwen3-TTS-GGUF

Serveurperso

GGUF weights for qwentts.cpp, a C++17/GGML port of Qwen3-TTS 12 Hz (Qwen team, Alibaba). Multilingual zero shot TTS with named speakers and Mandarin dialects, 24 kHz mono. Runs on CPU, CUDA, Metal, Vulkan. qwen-talker-{size}-{mode}-{variant}.gguf Qwen3 LM + code predictor MTP head + optional speaker encoder, text -> 12 Hz codes qwen-tokenizer-12hz-{variant}.gguf SEANet + ConvNeXt + DAC v2 + RVQ, 12 Hz codes 24 kHz audio Three modes are available across two talker sizes: The tokenizer is shared across every talker. Set GGMLBACKEND to force a device, otherwise the runtime picks the best one available. Tokenizer GGUFs are not uniform quants. Three categories get a Conv kernel rows (K=7,3,1)…

Open weights apache-2.0 gguf