Instructions to use vineetkukreti/garhwali-hindi-mt-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vineetkukreti/garhwali-hindi-mt-v2 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="vineetkukreti/garhwali-hindi-mt-v2")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("vineetkukreti/garhwali-hindi-mt-v2") model = AutoModelForSeq2SeqLM.from_pretrained("vineetkukreti/garhwali-hindi-mt-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Garhwali โ Hindi MT โ v2 (mT5-small) ยท template baseline
Neural machine translation between Hindi (hi) and Garhwali (gbm), both
in Devanagari. Garhwali is a Central Pahari language of Uttarakhand, India.
This is v2, published as a reproducible baseline and a cautionary data point. It is archived and two generations behind production.
| BLEU | 24.5 |
| chrF++ | 68.4 |
| Hindi leakage | 22.1% |
| Training pairs | 19,982 (from 27 fixed templates) |
| Status | archived |
Read this before citing v2's scores
v2 was trained on ~20k pairs expanded from 27 fixed sentence templates. Its test set shared those templates with training, so its reported scores overstate real translation quality.
On unseen lexical combinations v2 exhibits vocabulary collapse โ it learned 27 sentence shapes rather than the language. Treat 24.5 BLEU / 68.4 chrF++ as an in-distribution template score, not open-domain quality.
It is published precisely because that failure is instructive: it is what motivated v3 to cut the dataset to 5,000 more diverse pairs and score better.
Usage
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_id = "vineetkukreti/garhwali-hindi-mt-v2"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
inputs = tok("translate Hindi to Garhwali: เคฎเฅเค เค เฅเค เคนเฅเคเฅค", return_tensors="pt")
out = model.generate(**inputs, num_beams=4, max_new_tokens=128)
print(tok.decode(out[0], skip_special_tokens=True))
The prefix "translate {Source} to {Target}: " is required.
Training configuration
Recorded in the run manifest, with SHA-256 hashes of config and splits for reproducibility:
| Base revision | 73fb5dbe4756edadc8fbe8c769b0a109493acf7a (pinned) |
| Epochs | 12 |
| LR / schedule | 3e-4, warmup ratio 0.05, weight decay 0.01 |
| Batch | 8 ร grad-accum 4 |
| Optimiser | adafactor, label smoothing 0.1, bf16 |
| Max source/target length | 64 / 64 |
| Seed | 42 |
Limitations
- Do not use in production. Archived; v4 supersedes it.
- 22.1% Hindi leakage.
- Vocabulary collapse outside its 27 training templates.
- Reported metrics are inflated by template overlap (see above).
How this model line was built (v1 โ v4)
Every version was an answer to a specific failure of the one before it. The numbers that matter most here are not BLEU but the Hindi leakage rate โ how often the model emitted Hindi instead of Garhwali, which is the characteristic failure mode when fine-tuning a multilingual model into a low-resource sibling language.
| Ver | Name | Data | Pairs | BLEU | chrF++ | Hindi leakage | Status |
|---|---|---|---|---|---|---|---|
| v1 | Classical Linguistic Baseline | Chatak (1959) grammar + Grierson LSI seeds | 48 | 11.2 | 42.1 | 38.5% | deprecated |
| v2 | Synthetic Closed Template | 27 fixed templates, generated | 19,982 | 24.5 | 68.4 | 22.1% | archived |
| v3 | Balanced Multi-Aspect Synthetic | 34 topical batches (B-01โฆB-34) | 5,000 | 31.8 | 78.4 | 11.2% | candidate |
| v4 | Archival + Lexicon + Native Gold | Dhyani/Benjwal dicts, archives, native-reviewed gold | 8,642 | 36.4 | 83.9 | 6.8% | active |
v1 โ v2: scale, at the cost of generality
v1 trained on 48 pairs hand-derived from Dr. Govind Chatak's 1959 Garhwali grammar and Grierson's Linguistic Survey of India. It was precise on the exact rules it had seen and collapsed on everything else โ 38.5% Hindi leakage, severe OOV behaviour.
The v2 response was volume: 27 fixed templates expanded to ~20k pairs. Leakage nearly halved and BLEU more than doubled.
The learning โ and the trap. The gain was partly an illusion. Because the test set shared templates with training, scores flattered the model. On unseen lexical combinations v2 showed vocabulary collapse: it had learned 27 sentence shapes, not the language. This is why v2's reported chrF++ should never be read as open-domain quality.
v2 โ v3: diversity beats volume
v3 deliberately shrank the dataset โ 19,982 โ 5,000 pairs โ while widening
its shape: 34 topical batches (B-01โฆB-34) designed against a matrix schema, with
no_repeat_ngram_size=3 to penalise the repetition v2 had rewarded.
The learning: fewer, more varied examples beat many near-duplicates. BLEU rose 24.5 โ 31.8 and leakage halved again (22.1% โ 11.2%) on a quarter of the data. For low-resource MT, corpus entropy mattered more than corpus size.
v3 โ v4: synthetic data has a ceiling
v3 was still synthetic, so its errors were systematic โ fluent-looking output with wrong aspect morphology, because no generator had ever seen a native speaker correct it.
v4 changed the kind of data rather than the amount:
- curated dictionary pairs from the Dhyani and Benjwal Garhwali lexicons
- University archival texts (real published Garhwali prose)
- native-speaker-reviewed gold elicitation data, adjudicated rather than auto-accepted
Quarantine and adjudication ledgers (v4_train_quarantine.jsonl,
v4_1_adjudication_ledger.jsonl) record what was rejected and why โ rejection
was a first-class step, not a filter.
The learning: the last stretch of quality came from provenance, not modelling. No hyperparameter change moved progressive-aspect correctness the way native review did. Leakage reached 6.8% โ still non-zero, and still the main open problem.
Method notes that stayed constant
- Base
google/mt5-small, pinned to revision73fb5dbe4756edadc8fbe8c769b0a109493acf7a - Task prefix
"translate {Source} to {Target}: "โ part of the model contract; changing it defines a new experiment - Best checkpoint selected by chrF++, not loss (chrF++ is more informative for morphologically rich targets)
seed=42, adafactor, label smoothing 0.1, bf16, max source/target length 64- Splits grouped by document/speaker/collection, never random by sentence, to prevent leakage between train and test
- Benchmark data held entirely separate; IndicGenBench never trained on (canary + eval-only licence)
Related
vineetkukreti/garhwali-hindi-mt-v4โ production model (private)garhwali-hindi-mt-v3โ synthetic candidate (public)
License
Weights Apache-2.0, inherited from google/mt5-small. The training corpus
carries its own provenance terms, not covered by this licence.
- Downloads last month
- 44