Instructions to use kristijonas/kmynas-parakeet-lt-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use kristijonas/kmynas-parakeet-lt-v1 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("kristijonas/kmynas-parakeet-lt-v1") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Kmynas 🌾 — Lithuanian ASR with punctuation (Parakeet TDT 0.6B)
A 600M-parameter Lithuanian speech recogniser that emits punctuation and
casing, fine-tuned from NVIDIA's parakeet-tdt-0.6b-v3 on 169.6 h of LIEPA-3.
Named after caraway — lighter than paprika. It is smaller and faster than our Whisper-based Paprika V3, and unlike other public Parakeet fine-tunes for Lithuanian it does not output a bare lowercase stream.
input: a recording of a Seimas speech
output: Ypatingai svarbu diskutuoti šia tema Europos Sąjungos lygiu, nes
ekonominė krizė buvo nelengvas išbandymas socialinės apsaugos sistemų
tvarumui.
Results
Measured with identical chunking (35 s, pause-aligned) and identical normalisation (lowercase, punctuation stripped) for every model, so the numbers compare recognition and not formatting.
| model | FLEURS-lt | VoxPopuli-lt | params | punctuation |
|---|---|---|---|---|
| Kmynas (this model) | 16.60 | 19.37 | 0.6B | yes |
| Noctra parakeet-tdt-0.6b-v3-lt | 16.91 | 19.80 | 0.6B | no |
nvidia/parakeet-tdt-0.6b-v3 (base) |
22.15 | 29.56 | 0.6B | partial |
| Paprika V3 (ours, Whisper) | 11.93 | 17.25 | 0.8B | yes |
FLEURS and VoxPopuli are the honest comparisons — no model here trained on either. On a held-out LIEPA-3 set Kmynas scores 12.09, but that number is not comparable across models: LIEPA-3 ships no test split, so a held-out set is drawn from the same pool everyone trains on. Kmynas excludes those 19 speakers entirely; other LIEPA fine-tunes generally do not, which flatters theirs.
Paprika V3 remains more accurate. Kmynas is the better choice when size and speed matter: measured RTF 0.022 against Paprika's 0.48 — about 22x faster — CoreML-exportable, trained on 169.6 h rather than 3,281 h. 35 minutes of audio transcribes in 46 seconds on one GPU.
Punctuation quality
WER strips punctuation before scoring, so it cannot see the feature this model exists for. Measured directly as marks per 1,000 characters:
| set | Kmynas | human reference |
|---|---|---|
| VoxPopuli (human-punctuated) | 19.65 | 19.02 |
| LIEPA held-out (300 utts) | 28.97 | 29.45 |
100% of hypotheses carry punctuation; 98% start with a capital.
Training
| base | nvidia/parakeet-tdt-0.6b-v3 (FastConformer encoder, TDT decoder) |
| data | LIEPA-3: 94.8 h read + 79.1 h spontaneous, 112,644 utterances |
| clipping | 0.5–15.0 s, 16 kHz mono |
| targets | punctuated and cased (see below) |
| optimiser | AdamW, β=(0.9, 0.98), weight decay 1e-3 |
| LR | 2e-4, cosine annealing, 600 warmup, min 1e-6 |
| batch | micro 16 × grad-accum 4 = effective 64 |
| steps | 6,000 (3.41 epochs) |
| precision | bf16-mixed, single A100 80GB |
| tokenizer | unchanged from base (SentencePiece BPE, 8192) |
How punctuation was added
LIEPA-3 is entirely lowercase and unpunctuated, which is why models trained purely on it emit none. Rather than train on bare targets, every transcript was punctuated by an LLM under a strict word-preservation constraint: only punctuation marks and letter case may change, and a row whose word sequence differs is rejected rather than used.
- 111,469 rows processed; 98.9% passed the word-identity check
- 693 rows (0.6%) flagged as damaged — cut mid-word, unreadable — and dropped
- colloquial forms are preserved (
dvidešimis not "corrected" todvidešimt) - fragments cut mid-sentence get no invented ending
The constraint matters: if a restorer changes words, the audio no longer matches its target and the row silently teaches the model to hallucinate.
Held-out data
read items are excluded by whole speaker (19 speakers, 8,148 rows), spon
items by file hash. LIEPA-3's first 19 read speakers occupy its first ~8,000
rows, so any harvest that starts at shard 0 and does not exclude them trains on
its own evaluation set.
Usage
import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.from_pretrained("kristijonas/kmynas-parakeet-lt-v1")
print(model.transcribe(["audio_lt.wav"])[0].text) # 16 kHz mono
For long recordings, chunk at pauses (~35 s). The TDT decoder handles long audio natively, but chunking at silence gives cleaner sentence boundaries.
Decoding: NeMo's TDT config warns that greedy_batch is inaccurate for TDT
models, but on this checkpoint it made no difference — byte-identical output over
7.3 minutes of audio, 10.2 s vs 10.9 s. Use either.
Load in fp16. NeMo defaults to fp32, which costs 5.2 GB for weights alone. fp16 cut peak memory 6.3 -> 5.3 GB and decode time 76 s -> 46 s, with output 99.18% identical over 790 words.
Word timestamps are reliable: 0.3% zero-width (11 of 3,693), against ~19% for Whisper-family models.
Known issue: ⁇ at quotation marks (fixed in v2)
The inherited SentencePiece vocabulary has no quotation-mark tokens. It
encodes . , ? ! : - ' but not " „ “ ” ( ) ; – —. Punctuation-restored targets
contained Lithuanian quotes and en-dashes, SentencePiece mapped them to <unk>,
and the model faithfully learned to emit <unk> — rendered ⁇ — at those spots.
Rate is roughly 10-18 per 1,000 words depending on register. Every occurrence
sits at a correct quotation or dash position: the model learned placement and
directionality, it simply lacks the characters. Strip ⁇ from output as a
workaround, or treat it as a quote mark.
Affected 4.7% of training rows (67,162 instances: 48,679 quotes, 18,483 dashes).
v2 extends the vocabulary with „ “ – before retraining.
Limitations
- Lithuanian only. The base is multilingual; this fine-tune is not.
- Telephone speech is untested. No 8 kHz narrowband data in training and no telephone evaluation exists yet, so behaviour on call audio is unknown.
- Dialects are untested. Training used standard-register read and broadcast speech.
- Numbers are spelled out, following LIEPA convention (
du tūkstančiai antri, not2002). Benchmarks whose references use digits will score this model worse than it deserves — 18.4% of FLEURS references contain digits, and restricting to digit-free references lowers WER substantially. - 3.41 epochs. Published low-resource recipes typically use 5–50, and the WER-vs-steps curve was still falling, so more training should still help.
License
CC-BY-4.0, matching both the base model and the LIEPA-3 corpus. LIEPA-3 is a Vilnius University / Vytautas Magnus University / Institute of the Lithuanian Language resource — see liepa.lt for corpus terms.
Citation
@misc{kmynas2026,
title = {Kmynas: Lithuanian ASR with punctuation (Parakeet TDT 0.6B)},
author = {Kristijonas Jakubsonas},
year = {2026},
url = {https://huggingface.co/kristijonas/kmynas-parakeet-lt-v1}
}
- Downloads last month
- 10
Model tree for kristijonas/kmynas-parakeet-lt-v1
Base model
nvidia/parakeet-tdt-0.6b-v3Dataset used to train kristijonas/kmynas-parakeet-lt-v1
Evaluation results
- WER on FLEURS Lithuaniantest set self-reported16.600
- WER on VoxPopuli Lithuanian (gold subset)test set self-reported19.370