Garhwali โ†” Hindi MT โ€” v2 (mT5-small) ยท template baseline

Neural machine translation between Hindi (hi) and Garhwali (gbm), both in Devanagari. Garhwali is a Central Pahari language of Uttarakhand, India.

This is v2, published as a reproducible baseline and a cautionary data point. It is archived and two generations behind production.

BLEU 24.5
chrF++ 68.4
Hindi leakage 22.1%
Training pairs 19,982 (from 27 fixed templates)
Status archived

Read this before citing v2's scores

v2 was trained on ~20k pairs expanded from 27 fixed sentence templates. Its test set shared those templates with training, so its reported scores overstate real translation quality.

On unseen lexical combinations v2 exhibits vocabulary collapse โ€” it learned 27 sentence shapes rather than the language. Treat 24.5 BLEU / 68.4 chrF++ as an in-distribution template score, not open-domain quality.

It is published precisely because that failure is instructive: it is what motivated v3 to cut the dataset to 5,000 more diverse pairs and score better.

Usage

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

model_id = "vineetkukreti/garhwali-hindi-mt-v2"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

inputs = tok("translate Hindi to Garhwali: เคฎเฅˆเค‚ เค เฅ€เค• เคนเฅ‚เคเฅค", return_tensors="pt")
out = model.generate(**inputs, num_beams=4, max_new_tokens=128)
print(tok.decode(out[0], skip_special_tokens=True))

The prefix "translate {Source} to {Target}: " is required.

Training configuration

Recorded in the run manifest, with SHA-256 hashes of config and splits for reproducibility:

Base revision 73fb5dbe4756edadc8fbe8c769b0a109493acf7a (pinned)
Epochs 12
LR / schedule 3e-4, warmup ratio 0.05, weight decay 0.01
Batch 8 ร— grad-accum 4
Optimiser adafactor, label smoothing 0.1, bf16
Max source/target length 64 / 64
Seed 42

Limitations

  • Do not use in production. Archived; v4 supersedes it.
  • 22.1% Hindi leakage.
  • Vocabulary collapse outside its 27 training templates.
  • Reported metrics are inflated by template overlap (see above).

How this model line was built (v1 โ†’ v4)

Every version was an answer to a specific failure of the one before it. The numbers that matter most here are not BLEU but the Hindi leakage rate โ€” how often the model emitted Hindi instead of Garhwali, which is the characteristic failure mode when fine-tuning a multilingual model into a low-resource sibling language.

Ver Name Data Pairs BLEU chrF++ Hindi leakage Status
v1 Classical Linguistic Baseline Chatak (1959) grammar + Grierson LSI seeds 48 11.2 42.1 38.5% deprecated
v2 Synthetic Closed Template 27 fixed templates, generated 19,982 24.5 68.4 22.1% archived
v3 Balanced Multi-Aspect Synthetic 34 topical batches (B-01โ€ฆB-34) 5,000 31.8 78.4 11.2% candidate
v4 Archival + Lexicon + Native Gold Dhyani/Benjwal dicts, archives, native-reviewed gold 8,642 36.4 83.9 6.8% active

v1 โ†’ v2: scale, at the cost of generality

v1 trained on 48 pairs hand-derived from Dr. Govind Chatak's 1959 Garhwali grammar and Grierson's Linguistic Survey of India. It was precise on the exact rules it had seen and collapsed on everything else โ€” 38.5% Hindi leakage, severe OOV behaviour.

The v2 response was volume: 27 fixed templates expanded to ~20k pairs. Leakage nearly halved and BLEU more than doubled.

The learning โ€” and the trap. The gain was partly an illusion. Because the test set shared templates with training, scores flattered the model. On unseen lexical combinations v2 showed vocabulary collapse: it had learned 27 sentence shapes, not the language. This is why v2's reported chrF++ should never be read as open-domain quality.

v2 โ†’ v3: diversity beats volume

v3 deliberately shrank the dataset โ€” 19,982 โ†’ 5,000 pairs โ€” while widening its shape: 34 topical batches (B-01โ€ฆB-34) designed against a matrix schema, with no_repeat_ngram_size=3 to penalise the repetition v2 had rewarded.

The learning: fewer, more varied examples beat many near-duplicates. BLEU rose 24.5 โ†’ 31.8 and leakage halved again (22.1% โ†’ 11.2%) on a quarter of the data. For low-resource MT, corpus entropy mattered more than corpus size.

v3 โ†’ v4: synthetic data has a ceiling

v3 was still synthetic, so its errors were systematic โ€” fluent-looking output with wrong aspect morphology, because no generator had ever seen a native speaker correct it.

v4 changed the kind of data rather than the amount:

  • curated dictionary pairs from the Dhyani and Benjwal Garhwali lexicons
  • University archival texts (real published Garhwali prose)
  • native-speaker-reviewed gold elicitation data, adjudicated rather than auto-accepted

Quarantine and adjudication ledgers (v4_train_quarantine.jsonl, v4_1_adjudication_ledger.jsonl) record what was rejected and why โ€” rejection was a first-class step, not a filter.

The learning: the last stretch of quality came from provenance, not modelling. No hyperparameter change moved progressive-aspect correctness the way native review did. Leakage reached 6.8% โ€” still non-zero, and still the main open problem.

Method notes that stayed constant

  • Base google/mt5-small, pinned to revision 73fb5dbe4756edadc8fbe8c769b0a109493acf7a
  • Task prefix "translate {Source} to {Target}: " โ€” part of the model contract; changing it defines a new experiment
  • Best checkpoint selected by chrF++, not loss (chrF++ is more informative for morphologically rich targets)
  • seed=42, adafactor, label smoothing 0.1, bf16, max source/target length 64
  • Splits grouped by document/speaker/collection, never random by sentence, to prevent leakage between train and test
  • Benchmark data held entirely separate; IndicGenBench never trained on (canary + eval-only licence)

Related

  • vineetkukreti/garhwali-hindi-mt-v4 โ€” production model (private)
  • garhwali-hindi-mt-v3 โ€” synthetic candidate (public)

License

Weights Apache-2.0, inherited from google/mt5-small. The training corpus carries its own provenance terms, not covered by this licence.

Downloads last month
44
Safetensors
Model size
0.3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for vineetkukreti/garhwali-hindi-mt-v2

Base model

google/mt5-small
Finetuned
(764)
this model
Quantizations
1 model