SAVRN
Search Contact SAVRN

Open-weight model · Text classification

turn-detector

by LiveKit livekit/turn-detector

An open-weights language model for contextually-aware end-of-utterance (EOU) detection in voice AI applications.

Parameters135M
Context8,192
Weights703.1 MB
Licenseother
AccessOpen weights
Monthly Downloads951.9k

Runs On

What it takes to serve turn-detector (135M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.3 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on turn-detector

A silence timer cannot tell a pause from a finished sentence, and that gap is the whole job of LiveKit's turn-detector: it reads the transcript and decides whether the speaker is done. At 135M parameters it needs 0.3 GB at 16-bit, so hardware is not a decision here. The cheapest entry on our sheet is one MI300X at $1.85 an hour, and this model uses almost none of that card's 192 GB. It belongs on the same box as the rest of the voice stack, and the ONNX export in the file set gives a second runtime path beyond transformers.

The license reads other, with no summary in our record, so get the actual terms from LiveKit before a commercial deployment. It was derived and quantized from Qwen2.5-0.5B-Instruct, a second set of terms to check. Context is 8,192 tokens, well beyond a single turn.

Model Card

An open-weights language model for contextually-aware end-of-utterance (EOU) detection in voice AI applications. The model predicts whether a user has finished speaking based on the semantic content of their transcribed speech, providing a critical complement to voice activity detection (VAD) systems. Traditional voice agents rely on voice activity detection (VAD) to determine when a user has finished speaking. VAD works by detecting the presence or absence of speech in an audio signal and applying a silence timer. While effective for detecting pauses, VAD lacks language understanding and frequently causes false positives. For example, a user who says "I need to think about that for a…

Excerpt from the card by LiveKit, licensed other.

Configuration

Architecture
LlamaForCausalLM
Context length (tokens)
8,192
Layers
30
Hidden size
576
Feed-forward size
1,536
Attention heads
9
Key/value heads
3
Head dimension
64
Vocabulary size
49,154
RoPE base
100,000
Model type
llama

Identity and Version

Repository
livekit/turn-detector
Publisher
LiveKit
Task
Text classification
Modality
Text
Library
transformers
Parameters
135M parameters
Languages
en, es, fr, de, it, pt, nl, zh
Revision
fba34c38ad5d30a63ebb83a9e6bf271cf4c91d67
First published
2024-12-07
Last updated
2026-02-11

Files and Weights

14 files, 707.9 MB in total. The weights are 2 files totalling 703.1 MB in onnx, safetensors.

Weights2 files · 703.1 MB
Configuration5 files · 2.4 KB
Tokenizer4 files · 4.8 MB
Documentation2 files · 16.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights538.1 MB 2f7b4c93c1cd
model_quantized.onnxWeights165.0 MB 4e685767c364
added_tokens.jsonConfiguration50 B
config.jsonConfiguration793 B
generation_config.jsonConfiguration111 B
ort_config.jsonConfiguration763 B
special_tokens_map.jsonConfiguration699 B
LICENSEDocumentation5.6 KB
README.mdDocumentation10.5 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer466.4 KB
tokenizer.jsonTokenizer3.5 MB
tokenizer_config.jsonTokenizer3.9 KB
vocab.jsonTokenizer800.7 KB

License and Download

License
other
Access
Open weights, no gate
Download size
703.1 MB
Download from LiveKit

Released by LiveKit through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published703.1 MB
16-bit0.3 GB
8-bit0.1 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About turn-detector

How much GPU memory does turn-detector need?

About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (135M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run turn-detector on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is turn-detector released under?

other, as its publisher declares it. Read the license text before commercial use.

What is turn-detector's context length?

8,192 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This model is distilled from the zero-shot classification pipeline on the Multilingual Sentiment dataset using this script. In reality the multilingual-sentiment dataset is annotated of course, but we'll pretend and ignore the annotations for the sake of example. Result can be reproduce using the following commands: If you are training this model on Colab, make the following code changes to avoid Out-of-memory error message: - Transformers 4.28.1 - Pytorch 2.0.0+cu118 - Datasets 2.11.0 - Tokenizers 0.13.3

Open weights apache-2.0 135M parameters 512 tokens transformers

Model · Text classification

roberta-base-go_emotions

Sam Lowe

Model trained from roberta-base on the goemotions dataset for multi-label classification. A version of this model in ONNX format (including an INT8 quantized ONNX version) is now available at https://huggingface.co/SamLowe/roberta-base-goemotions-onnx. These are faster for inference, esp for smaller batch sizes, massively reduce the size of the dependencies required for inference, make inference of the model more multi-platform, and in the case of the quantized version reduce the model file/download size by 75% whilst retaining almost all the accuracy if you only need inference. goemotions is based on Reddit data and has 28 labels. It is a multi-label dataset where one or multiple labels…

Open weights mit 125M parameters 514 tokens transformers

Model · Text classification

cryptobert

Mikolaj Kulakowski

For academic reference, cite the following paper: https://ieeexplore.ieee.org/document/10223689 CryptoBERT is a pre-trained NLP model to analyse the language and sentiments of cryptocurrency-related social media posts and messages. It was built by further training the vinai's bertweet-base language model on the cryptocurrency domain, using a corpus of over 3.2M unique cryptocurrency-related social media posts. (A research paper with more details will follow soon.) The model was trained on the following labels: "Bearish": 0, "Neutral": 1, "Bullish": 2 CryptoBERT's sentiment classification head was fine-tuned on a balanced dataset of 2M labelled StockTwits posts, sampled from…

Open weights mit 125M parameters 514 tokens transformers

Model · Text classification

Firebird-ModernBERT-512-RW

Noumenon, Inc.

Firebird-ModernBERT-512-RW is an experimental post-trained variant of It is a ~149M parameter ModernBERT binary classifier for distinguishing: - 0 — HUMAN The maximum sequence length is 512 tokens. This checkpoint was produced through reward-weighted classifier post-training. The original Firebird checkpoint was kept frozen as a reference model. Training examples were scored by the original classifier, difficult examples received larger loss weights, and the post-trained model was constrained against the frozen reference using a KL penalty. L = weightedcrossentropy + beta KL(reference || policy) Hard human examples received greater weighting than ordinary examples because one goal of the…

Open weights 150M parameters 8,192 tokens transformers

Model · Text classification

inclusively-classification

E-MIMIC

This model is an Italian classification model fine-tuned from the Italian BERT model for the classification of inclusive language in Italian. It has been trained to detect three classes: - inclusive: the sentence is inclusive (e.g. "Il personale docente e non docente") - notinclusive: the sentence is not inclusive (e.g. "I professori") - notpertinent: the sentence is not pertinent to the task (e.g. "La scuola è chiusa") The model has been trained on a dataset containing: - 8580 training sentences - 1073 validation sentences - 1072 test sentences The data collection has been manually annotated by experts in the field of inclusive language (dataset is not publicly available yet). The model…

Open weights cc-by-nc-sa-4.0 111M parameters 512 tokens transformers

Model · Text classification

multi-domain-sentiment-bert

ADITYA GUPTA

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 109M parameters 512 tokens transformers