Instructions to use kinit/parakeet-tdt-0.6b-v3-sk with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use kinit/parakeet-tdt-0.6b-v3-sk with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("kinit/parakeet-tdt-0.6b-v3-sk") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
parakeet-tdt-0.6b-v3-sk
Slovak fine-tune of nvidia/parakeet-tdt-0.6b-v3 by the KInIT team. Full parameter fine-tuning on a curated Slovak speech corpus, with noise augmentation for robustness to real-world recording conditions. Listed in SLAIH, a catalog of Slovak NLP resources.
Model Details
| Property | Value |
|---|---|
| Base model | nvidia/parakeet-tdt-0.6b-v3 |
| Parameters | ~600M |
| Architecture | FastConformer encoder + TDT (Token-and-Duration Transducer) decoder |
| Fine-tuning method | Full fine-tuning |
| Language | Slovak (sk) |
| Task | Automatic Speech Recognition |
| Framework | NVIDIA NeMo |
| License | CC-BY-4.0 |
Intended Use
This model is intended for Slovak automatic speech recognition across a range of domains and recording conditions. It transcribes with punctuation and capitalization.
Out-of-scope: Non-Slovak audio, safety-critical transcription without human review.
Evaluation
Word Error Rate (WER) and Character Error Rate (CER), lower is better. Measured on two Slovak eval sets:
- CV24 - Common Voice 24.0 test split (5,239 samples, public)
- Internal - held-out KInIT set (9,317 samples, stratified by domain and speaker gender, one third clean and two thirds noise-augmented; not public)
Before vs. after fine-tuning
| Model | CV24 WER โ | CV24 CER โ | Internal WER โ | Internal CER โ |
|---|---|---|---|---|
| kinit/parakeet-tdt-0.6b-v3-sk | 7.92% | 2.06% | 6.62% | 3.12% |
| nvidia/parakeet-tdt-0.6b-v3 | 9.07% | 2.54% | 24.32% | 10.75% |
Slovak fine-tuning reduces WER by 12.8% on Common Voice and 72.8% on the internal eval set, at the cost of multilingual performance (see Limitations).
Parakeet vs. Canary
Both Slovak NeMo models were fine-tuned on the same corpus with the same noise-augmentation recipe and scored with the same pipeline, so this comparison isolates the architecture. Comparison before and after Slovak fine-tuning; WER is measured on our internal benchmark dataset (see above):
| Model | Fine-tuned: WER โ | Base: WER โ |
|---|---|---|
| Parakeet TDT 0.6B v3 | 6.62% | 24.32% |
| Canary 1B v2 | 6.37% | 14.39% |
Parakeet is the faster of the two.
Besides these two NVIDIA NeMo fine-tunes (Canary, Parakeet), the full KInIT ASR collection also includes six Slovak fine-tunes of Whisper, OpenAI's separate speech-recognition architecture, at sizes from tiny to large.
Training Data
Fine-tuned on an internal curated Slovak speech corpus compiled at KInIT. The corpus combines public datasets with internal KInIT recordings. Recordings containing personal data were anonymised prior to use. Samples were quality-filtered using a CER-based threshold validated against multiple ASR models. 75% of the training samples were additionally augmented with synthetic background noise for robustness to real-world recording conditions (see Training Procedure below).
| Data sources |
|---|
| SloPalSpeech |
| Municipal council session recordings |
| Read literature |
| Mozilla Common Voice |
| TEDxSK and JumpSK Lecture Speech Corpus |
| FLEURS read speech |
| Internal KInIT recordings |
Training Procedure
| Hyperparameter | Value |
|---|---|
| Epochs | 3 |
| Learning rate | 4e-4 |
| LR scheduler | Cosine annealing with warmup |
| Optimizer | AdamW |
| Effective batch size | 64 |
| Precision | bf16-mixed |
| Gradient clipping | 1.0 |
| Noise augmentation | Applied to 75% of training samples: phone noise, background speech, background noise, white noise, and packet loss |
| Framework | NVIDIA NeMo with PyTorch Lightning |
Training was performed on the Devana HPC cluster.
Usage
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.from_pretrained("kinit/parakeet-tdt-0.6b-v3-sk")
output = model.transcribe(["audio.wav"])
print(output[0].text)
audio.wav must be mono โ stereo input raises a shape-mismatch error.
License
CC-BY-4.0, inherited from the base model nvidia/parakeet-tdt-0.6b-v3 (ยฉ NVIDIA).
Limitations
- Catastrophic forgetting: Fine-tuning exclusively on Slovak data significantly degrades the base model's performance on the other 24 languages it supports. Use the base nvidia/parakeet-tdt-0.6b-v3 if multilingual transcription is required.
- Performance may degrade on strongly accented, dialectal, or domain-specific speech not represented in the training data.
- The training data consists predominantly of short utterances. The base model's long-form capability is inherited but was not evaluated after fine-tuning.
- Word and segment timestamps are supported by the base architecture but were not evaluated after fine-tuning.
Acknowledgements
Public datasets used in training: SloPalSpeech, Mozilla Common Voice, FLEURS, and the TEDxSK and JumpSK Lecture Speech Corpus (KEMT NLP).
(Part of the) Research results was obtained using the computational resources procured in the national project National competence centre for high performance computing (project code: 311070AKF2) funded by European Regional Development Fund, EU Structural Funds Informatization of society, Operational Program Integrated Infrastructure.
- Downloads last month
- 77
Model tree for kinit/parakeet-tdt-0.6b-v3-sk
Base model
nvidia/parakeet-tdt-0.6b-v3