Gemma-4-E4B DoRA: Classical Sanskrit Poetry $\to$ Hindi (Half 2)

This repository contains a DoRA (Weight-Decomposed Low-Rank Adaptation) adapter for google/gemma-4-E4B fine-tuned for high-fidelity translation of Classical Sanskrit Verses (Shlokas) into Modern Hindi.

This model represents Half 2 of a dual-half data split experiment designed to study within-domain adaptation stability, magnitude/direction dynamics ($\Delta M$ vs $\Delta D$), and representation divergence against prose translation.


Model Details

  • Base Model: google/gemma-4-E4B
  • Tuning Method: DoRA (Weight-Decomposed Low-Rank Adaptation)
  • Rank ($r$): 16
  • Alpha ($\alpha$): 32 ($\alpha / r = 2.0$)
  • LoRA Dropout: 0.05
  • Target Modules: All 7 linear projection layers in the 42 text decoder layers:
    • q_proj, k_proj, v_proj, o_proj (Self-Attention)
    • gate_proj, up_proj, down_proj (MLP)
  • Trainable Parameters: 37,925,888 / 8,034,082,336 (0.472%)
  • Language Pair: Sanskrit (sa) $\to$ Hindi (hi)
  • Domain: Classical Sanskrit Poetry (Epics, Stotras, Subhashitas, Kavyas)

Evaluation & Training Dynamics

The model was trained for 2 epochs on 29,947 parallel verse pairs and evaluated on 3,778 unseen validation verses.

Validation Loss & Perplexity Trajectory

Milestone Checkpoint Step Epoch Eval Loss (Cross-Entropy) Perplexity ($\text{PPL} = e^{\text{loss}}$)
Milestone 25% 936 0.50 2.617 13.69
Milestone 50% 1,872 1.00 2.435 11.42
Milestone 75% 2,808 1.50 2.354 10.53
Final (100%) 3,744 2.00 2.330 10.28

Within-Domain Stability (Half 1 vs. Half 2)

Comparing both halves of the independent poetry splits confirms high empirical consistency:

Metric Poetry Half 1 Poetry Half 2 Absolute Delta ($\Delta$) Relative Variance
Validation Loss 2.323 2.330 0.007 < 0.30%
Validation Perplexity 10.20 10.28 0.08 < 0.78%
Final Train Loss (Smoothed) 8.60 8.58 0.02 < 0.23%

This tight convergence demonstrates that DoRA preserves structural consistency and semantic alignment across disparate verse selections.


Training Hyperparameters

  • Hardware: 1x NVIDIA H100 80GB SXM5 GPU (Sapphire Rapids host)
  • Runtime: 2 hours 19 minutes (8,340s)
  • Throughput: 7.18 samples/sec (0.45 optimizer steps/sec)
  • Effective Batch Size: 16 (per_device_batch_size=4, gradient_accumulation_steps=4)
  • Optimizer: AdamW (weight_decay=0.01, max_grad_norm=1.0)
  • Learning Rate: 5e-5 with Cosine Annealing and 100 linear warmup steps
  • Sequence Length: 384 tokens (loss computed strictly on target tokens with -100 prompt masking)
  • Precision: bfloat16

How to Use

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

BASE_MODEL = "google/gemma-4-E4B"
ADAPTER_REPO = "NIVED2003/gemma-4-E4B-dora-poetry-half2"

# 1. Load Tokenizer & Base Model
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL)
model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

# 2. Load DoRA Adapter
model = PeftModel.from_pretrained(model, ADAPTER_REPO)
model.eval()

# 3. Format Prompt
shloka = "यदा यदा हि धर्मस्य ग्लानिर्भवति भारत । अभ्युत्थानमधर्मस्य तदात्मानं सृजाम्यहम् ॥"
prompt = f"Instruction: Translate the following Sanskrit classical verse (shloka) to Hindi.\nInput: {shloka}\nOutput: "

inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

# 4. Generate Translation
with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=128,
        temperature=0.3,
        top_p=0.9,
        do_sample=True,
        eos_token_id=tokenizer.eos_token_id
    )

translation = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print("Translation:", translation)

Intermediate Milestones Available

All intermediate milestone checkpoints are versioned in this repository:

  • checkpoints/step_25pct: 25% training milestone (step 936)
  • checkpoints/step_50pct: 50% training milestone (step 1872)
  • checkpoints/step_75pct: 75% training milestone (step 2808)
  • Root directory (.): Final fully-trained 100% adapter (step 3744)

Citation & Architecture

@inproceedings{liu2024dora,
  title={DoRA: Weight-Decomposed Low-Rank Adaptation},
  author={Liu, Shih-Yang and Wang, Chien-Yi and Yin, Hongxu and Khona, Pavlo and Shen, Sheng and Yen, Chen-Yu and Wang, Ting-Kuei and Chen, Kuan-Yu and Darve, Eric and Chen, Kwang-Ting},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2024}
}
Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NIVED2003/gemma-4-E4B-dora-poetry-half2

Adapter
(21)
this model