Nathan Roll 1,2 · Irene Yi 1,2 · Büşra Marşan 1,2 1,3 Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It replaces half of the encoder's temporal depthwise filters with 12,288 fitted, frozen Gabor kernels and trains the remaining parameters on multilingual and multi-accent data. Orukeet outperforms Parakeet on 61 of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). Across all 25 FLEURS languages, pooled WER is 9.85% vs. 11.01%, a 10.6% relative reduction. Final adaptation and checkpoint selection use LibriSpeech test-other. Use Orukeet for recordings, media, batch…
Organization
Oruk
oruk
Speech foundation models for transcription, speaker diarization, vocal emotion, and delivery. We help AI understand what people say, how they say it, and what they mean—directly from the audio.
Models
Submission withdrawn September 19, 2026 (UTC). Open ASR PR #221 is closed. The leaderboard submission has been withdrawn; the checkpoint, code, research results, and exposure disclosures remain available as historical research artifacts. This repository preserves the original R15-0100 Orukeet checkpoint originally proposed as an earlier checkpoint for Open ASR submission #221. The artifact is unchanged from the archived September 6, 2026 model. It has a FastConformer encoder and TDT decoder derived from NVIDIA Parakeet TDT 0.6B v3, with multilingual/accent adaptation and fitted temporal Gabor kernels materialized in the native weights. The NeMo file is orukeet-r15-0100.nemo (2,509,342,720…