SAVRN
Search Contact SAVRN

Open-weight model · Speech recognition

trans-local-speech-models

by Chenggang Chen AladdinChen/trans-local-speech-models

trans-local-speech-models is an open-weight model for speech recognition from Chenggang Chen, released under other. Its published files total 5.6 GB.

Versioned, data-only runtime packages for local speech recognition in Trans. These are conversions of the credited upstream models, not models trained by AladdinChen. The original authors retain their rights. No endorsement is implied.

Parameters
Context
Weights5.6 GB
Licenseother
AccessOpen weights
Monthly Downloads

Model Card

Versioned, data-only runtime packages for local speech recognition in Trans. These are conversions of the credited upstream models, not models trained by AladdinChen. The original authors retain their rights. No endorsement is implied. Each package is downloaded separately. Recognition runs on the user's device; this repository contains no hosted inference service, recordings, training data, app credentials, or application code. See LICENSES.md for source authors, licenses and modification notices. Full license texts and preserved upstream notices are in licenses/ and provenance/. Redistributors must retain the notices and meet the applicable per-model license conditions. This is not a…

Excerpt from the card by Chenggang Chen, licensed other.

Identity and Version

Repository
AladdinChen/trans-local-speech-models
Publisher
Chenggang Chen
Task
Speech recognition
Modality
Audio
Library
Not stated by the source
Parameters
Not stated by the source
Languages
tr, fi, da, sv, he, el, nb, fil
Revision
9fc5ad76c170a9af1b2f5a3abc80129ae3db87dd
First published
2026-09-19
Last updated
2026-09-19

Files and Weights

80 files, 5.6 GB in total. The weights are 29 files totalling 5.6 GB in bin, onnx.

Weights29 files · 5.6 GB
Configuration1 file · 19.5 KB
Tokenizer2 files · 15.0 KB
Documentation34 files · 288.8 KB
Other13 files · 308.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
models/ar/model.int8.onnxWeights131.7 MB 017a46ae61bf
models/da/da-whisper-q8.binWeights264.5 MB 55e73c83b66e
models/el/el-whisper-q8.binWeights874.2 MB f0c778ab88c0
models/fa/decoder.int8.onnxWeights4.0 MB 0adaad326a53
models/fa/encoder.int8.onnxWeights131.3 MB b9d975c1af77
models/fa/joiner.int8.onnxWeights1.4 MB e441ed265c96
models/fi/fi-whisper-q8.binWeights823.4 MB bb803255675c
models/fil/fil-model.int8.onnxWeights174.5 MB 2a3e324be1d9
models/he/he-whisper-q8.binWeights874.2 MB 123a936e686b
models/hi/model.int8.onnxWeights197.6 MB 915c71e04dd7
models/hr/model.int8.onnxWeights131.3 MB 6b84afe742cf
models/id/decoder-iter-100000-avg-15-chunk-32-left-256.int8.onnxWeights540.7 KB 6544848ca80b
models/id/encoder-iter-100000-avg-15-chunk-32-left-256.int8.onnxWeights70.1 MB 3a6f85f5d199
models/id/joiner-iter-100000-avg-15-chunk-32-left-256.int8.onnxWeights259.4 KB 4b89d9629246
models/no/no-whisper-q8.binWeights264.5 MB 3d71ee8d9af1
models/parakeet/decoder.int8.onnxWeights11.8 MB 179e50c43d1a
models/parakeet/encoder.int8.onnxWeights652.2 MB acfc2b445637
models/parakeet/joiner.int8.onnxWeights6.4 MB 3164c13fc282
models/ru/decoder.onnxWeights1.2 MB ea0844367648
models/ru/encoder.int8.onnxWeights224.6 MB aa4e9bd1d8f1
models/ru/joiner.onnxWeights687.8 KB ade116563dbf
models/sv/sv-whisper-q8.binWeights264.5 MB 0073fd16fca4
models/th/decoder-epoch-12-avg-5.onnxWeights5.2 MB 21ebb9302e1c
models/th/encoder-epoch-12-avg-5.int8.onnxWeights154.7 MB 9ceb14519103
models/th/joiner-epoch-12-avg-5.int8.onnxWeights1.0 MB 6ee983924b39
models/tr/tr-whisper-q8.binWeights264.5 MB d7ab7303f2be
models/vi/decoder.int8.onnxWeights1.3 MB e0a156b54547
models/vi/encoder.int8.onnxWeights70.9 MB b528768939c7
models/vi/joiner.int8.onnxWeights1.0 MB 12636559d135
manifest.jsonConfiguration19.5 KB
LICENSES.mdDocumentation11.0 KB
README.mdDocumentation4.5 KB
provenance/ar/README.mdDocumentation3.0 KB
provenance/ar/original/README.mdDocumentation13.9 KB
provenance/da/README.mdDocumentation5.1 KB
provenance/da/THIRD_PARTY_NOTICES.mdDocumentation2.3 KB
provenance/el/README.mdDocumentation6.8 KB
provenance/fa/LICENSEDocumentation11.4 KB
provenance/fa/README.mdDocumentation7.4 KB
provenance/fa/base/README.mdDocumentation4.4 KB
provenance/fa/original/LICENSEDocumentation11.4 KB
provenance/fa/original/README.mdDocumentation1.7 KB
provenance/fi/README.mdDocumentation883 B
provenance/fil/README.mdDocumentation1.5 KB
provenance/fil/original/README.mdDocumentation4.6 KB
provenance/he/README.mdDocumentation215 B
provenance/hi/README.mdDocumentation21.1 KB
provenance/hi/source/LICENSEDocumentation1.1 KB
provenance/hi/source/README.mdDocumentation5.0 KB
provenance/hr/README.mdDocumentation3.0 KB
provenance/hr/original/README.mdDocumentation7.9 KB
provenance/id/README.mdDocumentation902 B
provenance/no/README.mdDocumentation20.3 KB
provenance/parakeet/README.mdDocumentation42.9 KB
provenance/parakeet/original/README.mdDocumentation42.9 KB
provenance/ru/README.mdDocumentation3.8 KB
provenance/ru/original/README.mdDocumentation2.9 KB
provenance/ru/source/LICENSEDocumentation1.1 KB
provenance/ru/source/README.mdDocumentation10.6 KB
provenance/sv/README.mdDocumentation16.4 KB
provenance/th/README.mdDocumentation78 B
provenance/tr/README.mdDocumentation3.1 KB
provenance/vi/README.mdDocumentation7.1 KB
provenance/vi/source/README.mdDocumentation8.8 KB
licenses/Apache-2.0.txtOther11.4 KB
licenses/CC-BY-4.0.txtOther18.7 KB
licenses/CC0-1.0.txtOther7.0 KB
licenses/MIT-terms.txtOther1.0 KB
licenses/Whisper-MIT.txtOther1.1 KB
models/fa/tokens.txtOther12.2 KB
models/fil/fil-tokens.txtOther11.7 KB
models/hi/tokens.txtOther67.6 KB
models/id/tokens.txtOther5.4 KB
models/parakeet/tokens.txtOther93.9 KB
models/ru/tokens.txtOther13.4 KB
models/th/tokens.txtOther39.1 KB
models/vi/tokens.txtOther25.8 KB
.gitattributesRepository1.5 KB
models/ar/vocab.txtTokenizer12.9 KB
models/hr/vocab.txtTokenizer2.0 KB

License and Download

License
other
Access
Open weights, no gate
Download size
5.6 GB
Download from Chenggang Chen

Released by Chenggang Chen through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published5.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About trans-local-speech-models

What license is trans-local-speech-models released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Speech recognition

wav2vec2-large-xlsr-53-japanese

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Japanese using the train and validation splits of Common Voice 6.1, CSS10 and JSUT. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: The model can be evaluated as follows on the Japanese test data of Common Voice. In the table below I report the Word Error Rate (WER) and the Character Error Rate (CER) of the model. I ran the…

Open weights apache-2.0 transformers

Model · Speech recognition

whisperkit-coreml

Argmax

WhisperKit is part of Argmax OSS, an On-device Speech AI SDK for Apple Silicon: https://github.com/argmaxinc/argmax-oss-swift Check out the WhisperKit paper and presentation from ICML 2025: https://icml.cc/virtual/2025/47854 For real-time transcription with speakers and custom vocabulary, check out Argmax Pro SDK: https://www.argmaxinc.com/blog/argmax-sdk-2

Open weights mit whisperkit

Model · Speech recognition

speaker-diarization-3.1

Pyannote

Using this open-source model in production? Consider switching to pyannoteAI for better and faster options. This pipeline is the same as pyannote/speaker-diarization-3.0 except it removes the problematic use of onnxruntime. Both speaker segmentation and embedding now run in pure PyTorch. This should ease deployment and possibly speed up inference. It requires pyannote.audio version 3.1 or higher. It ingests mono audio sampled at 16kHz and outputs speaker diarization as an Annotation instance: - stereo or multi-channel audio files are automatically downmixed to mono by averaging the channels. - audio files sampled at a different rate are resampled to 16kHz automatically upon loading. 1.…

Access requested at publisher mit pyannote-audio

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Portuguese using the train and validation splits of Common Voice 6.1. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata If you want to cite this model you can use this

Open weights apache-2.0 transformers

Model · Speech recognition

speaker-diarization-community-1

Pyannote

This pipeline ingests mono audio sampled at 16kHz and outputs speaker diarization. - stereo or multi-channel audio files are automatically downmixed to mono by averaging the channels. - audio files sampled at a different rate are resampled to 16kHz automatically upon loading. The main improvements brought by Community-1 are: - improved speaker assignment and counting - simpler reconciliation with transcription timestamps with exclusive speaker diarization - easy offline use (i.e. without internet connection) - (optionally) hosted on pyannoteAI cloud 1. pip install pyannote.audio 3. Create access token at hf.co/settings/tokens. Out of the box, Community-1 is much better than…

Access requested at publisher cc-by-4.0 pyannote-audio

Model · Speech recognition

wav2vec2-large-xlsr-53-russian

Jonatas Grosman

Fine-tuned facebook/wav2vec2-large-xlsr-53 on Russian using the train and validation splits of Common Voice 6.1 and CSS10. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud:) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint The model can be used directly (without a language model) as follows... Using the HuggingSound library: 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata If you want to cite this model you can use this

Open weights apache-2.0 transformers