Kmynas 🌾 — Lithuanian ASR with punctuation (Parakeet TDT 0.6B)

A 600M-parameter Lithuanian speech recogniser that emits punctuation and casing, fine-tuned from NVIDIA's parakeet-tdt-0.6b-v3 on 169.6 h of LIEPA-3.

Named after caraway — lighter than paprika. It is smaller and faster than our Whisper-based Paprika V3, and unlike other public Parakeet fine-tunes for Lithuanian it does not output a bare lowercase stream.

input:  a recording of a Seimas speech
output: Ypatingai svarbu diskutuoti šia tema Europos Sąjungos lygiu, nes
        ekonominė krizė buvo nelengvas išbandymas socialinės apsaugos sistemų
        tvarumui.

Results

Measured with identical chunking (35 s, pause-aligned) and identical normalisation (lowercase, punctuation stripped) for every model, so the numbers compare recognition and not formatting.

model FLEURS-lt VoxPopuli-lt params punctuation
Kmynas (this model) 16.60 19.37 0.6B yes
Noctra parakeet-tdt-0.6b-v3-lt 16.91 19.80 0.6B no
nvidia/parakeet-tdt-0.6b-v3 (base) 22.15 29.56 0.6B partial
Paprika V3 (ours, Whisper) 11.93 17.25 0.8B yes

FLEURS and VoxPopuli are the honest comparisons — no model here trained on either. On a held-out LIEPA-3 set Kmynas scores 12.09, but that number is not comparable across models: LIEPA-3 ships no test split, so a held-out set is drawn from the same pool everyone trains on. Kmynas excludes those 19 speakers entirely; other LIEPA fine-tunes generally do not, which flatters theirs.

Paprika V3 remains more accurate. Kmynas is the better choice when size and speed matter: measured RTF 0.022 against Paprika's 0.48 — about 22x faster — CoreML-exportable, trained on 169.6 h rather than 3,281 h. 35 minutes of audio transcribes in 46 seconds on one GPU.

Punctuation quality

WER strips punctuation before scoring, so it cannot see the feature this model exists for. Measured directly as marks per 1,000 characters:

set Kmynas human reference
VoxPopuli (human-punctuated) 19.65 19.02
LIEPA held-out (300 utts) 28.97 29.45

100% of hypotheses carry punctuation; 98% start with a capital.

Training

base nvidia/parakeet-tdt-0.6b-v3 (FastConformer encoder, TDT decoder)
data LIEPA-3: 94.8 h read + 79.1 h spontaneous, 112,644 utterances
clipping 0.5–15.0 s, 16 kHz mono
targets punctuated and cased (see below)
optimiser AdamW, β=(0.9, 0.98), weight decay 1e-3
LR 2e-4, cosine annealing, 600 warmup, min 1e-6
batch micro 16 × grad-accum 4 = effective 64
steps 6,000 (3.41 epochs)
precision bf16-mixed, single A100 80GB
tokenizer unchanged from base (SentencePiece BPE, 8192)

How punctuation was added

LIEPA-3 is entirely lowercase and unpunctuated, which is why models trained purely on it emit none. Rather than train on bare targets, every transcript was punctuated by an LLM under a strict word-preservation constraint: only punctuation marks and letter case may change, and a row whose word sequence differs is rejected rather than used.

  • 111,469 rows processed; 98.9% passed the word-identity check
  • 693 rows (0.6%) flagged as damaged — cut mid-word, unreadable — and dropped
  • colloquial forms are preserved (dvidešim is not "corrected" to dvidešimt)
  • fragments cut mid-sentence get no invented ending

The constraint matters: if a restorer changes words, the audio no longer matches its target and the row silently teaches the model to hallucinate.

Held-out data

read items are excluded by whole speaker (19 speakers, 8,148 rows), spon items by file hash. LIEPA-3's first 19 read speakers occupy its first ~8,000 rows, so any harvest that starts at shard 0 and does not exclude them trains on its own evaluation set.

Usage

import nemo.collections.asr as nemo_asr
model = nemo_asr.models.ASRModel.from_pretrained("kristijonas/kmynas-parakeet-lt-v1")
print(model.transcribe(["audio_lt.wav"])[0].text)   # 16 kHz mono

For long recordings, chunk at pauses (~35 s). The TDT decoder handles long audio natively, but chunking at silence gives cleaner sentence boundaries.

Decoding: NeMo's TDT config warns that greedy_batch is inaccurate for TDT models, but on this checkpoint it made no difference — byte-identical output over 7.3 minutes of audio, 10.2 s vs 10.9 s. Use either.

Load in fp16. NeMo defaults to fp32, which costs 5.2 GB for weights alone. fp16 cut peak memory 6.3 -> 5.3 GB and decode time 76 s -> 46 s, with output 99.18% identical over 790 words.

Word timestamps are reliable: 0.3% zero-width (11 of 3,693), against ~19% for Whisper-family models.

Known issue: ⁇ at quotation marks (fixed in v2)

The inherited SentencePiece vocabulary has no quotation-mark tokens. It encodes . , ? ! : - ' but not " „ “ ” ( ) ; – —. Punctuation-restored targets contained Lithuanian quotes and en-dashes, SentencePiece mapped them to <unk>, and the model faithfully learned to emit <unk> — rendered ⁇ — at those spots.

Rate is roughly 10-18 per 1,000 words depending on register. Every occurrence sits at a correct quotation or dash position: the model learned placement and directionality, it simply lacks the characters. Strip ⁇ from output as a workaround, or treat it as a quote mark.

Affected 4.7% of training rows (67,162 instances: 48,679 quotes, 18,483 dashes). v2 extends the vocabulary with „ “ – before retraining.

Limitations

  • Lithuanian only. The base is multilingual; this fine-tune is not.
  • Telephone speech is untested. No 8 kHz narrowband data in training and no telephone evaluation exists yet, so behaviour on call audio is unknown.
  • Dialects are untested. Training used standard-register read and broadcast speech.
  • Numbers are spelled out, following LIEPA convention (du tūkstančiai antri, not 2002). Benchmarks whose references use digits will score this model worse than it deserves — 18.4% of FLEURS references contain digits, and restricting to digit-free references lowers WER substantially.
  • 3.41 epochs. Published low-resource recipes typically use 5–50, and the WER-vs-steps curve was still falling, so more training should still help.

License

CC-BY-4.0, matching both the base model and the LIEPA-3 corpus. LIEPA-3 is a Vilnius University / Vytautas Magnus University / Institute of the Lithuanian Language resource — see liepa.lt for corpus terms.

Citation

@misc{kmynas2026,
  title  = {Kmynas: Lithuanian ASR with punctuation (Parakeet TDT 0.6B)},
  author = {Kristijonas Jakubsonas},
  year   = {2026},
  url    = {https://huggingface.co/kristijonas/kmynas-parakeet-lt-v1}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kristijonas/kmynas-parakeet-lt-v1

Finetuned
(108)
this model

Dataset used to train kristijonas/kmynas-parakeet-lt-v1

Evaluation results