Sevak-97M Sovereign 🙏

ਪੰਜਾਬੀ-ਫਰਸਟ ਬਾਇਲਿੰਗੁਅਲ ਭਾਸ਼ਾ ਮਾਡਲ — ਲੈਪਟਾਪ 'ਤੇ ਸਿਖਲਾਈ, ਸਿਫ਼ਰ ਤੋਂ। (A Punjabi-first bilingual language model, trained from scratch on a laptop.)

Sevak-97M Sovereign is a 97.2M-parameter Punjabi-first bilingual (ਪੰਜਾਬੀ + English) language model, trained from scratch on a consumer laptop (Apple Silicon, MPS). No cloud, no GPU cluster.

Architecture

Parameters 97.2M
Layers 12 (D=768)
Attention RoPE + QK-Norm (standard causal attention), SwiGLU FFN
Embeddings Tied (std=0.02 init)
Tokenizer 16K BPE (Gurmukhi-optimized, 4.0 chars/token)
Context 256 train / 1024 inference

Training Pipeline (3 phases)

  1. Phase A: Punjabi corpus pretraining (40K chunks, 10K steps): loss 9.8 → 3.66
  2. Phase B: QA SFT (7.6K pairs: Bonsai-27B generated + verified SFT): → 2.10
  3. Phase C (Sovereign): 61.7K unique examples (rejection-sampled self-improvement winners, correction pass, sovereign knowledge, bilingual): 3.93 → 3.08

Note: the Phase B teacher (Bonsai-27B) itself scores only 35.7% on the factual test below and makes confident errors on basic Sikh history, so some of that generated data is likely wrong.

Evaluation

Factual accuracy (does it say the right answer?)

143 Punjabi factual questions (Sikh history, Punjab, general knowledge), decontaminated against all SFT sources (0 of them appear in Phase C data). Greedy decoding, prompt ਸਵਾਲ: …\nਜਵਾਬ:, scored by keyword/substring match against a reference answer. The reference answers were written with Gemini, so the scorer may slightly favour Gemini-style phrasing.

Model Size Correct
Sevak-97M Sovereign 97M 14.7%
Sevak-97M Sovereign + Wikipedia retrieval in prompt 97M 8.4%
sahaj-86m (SFT) 86M 25.9%
Qwen3.5-4B (4-bit) 4B 21.7%
Bonsai-2-27B 27B 35.7%
Qwen2.5-14B-Instruct (4-bit) 14B 39.9%

At this size the model is fluent in Punjabi but hallucinates facts. Example: asked when and where the Khalsa was founded, it answers "1869, Uttarakhand" (correct: 1699, Anandpur Sahib). It also does not yet use retrieved context (it was never trained on context→answer).

Bits-per-byte (language modelling, lower = better)

BPB measures how well a model predicts text, not whether it knows facts.

Multi-domain set of fresh texts written for this benchmark (not in training data). Sample sizes are small: 24 texts per prose bucket, 8 texts per QA bucket, so QA differences of a few hundredths are within noise.

Domain Sevak-97M Qwen3.5-2B Qwen3.5-4B LFM2-350M Gemma-3-1B
ਪੰਜਾਬੀ prose (n=24) 0.625 0.865 0.592 1.846 2.333
ਪੰਜਾਬੀ QA (n=8) 0.547 0.579 0.382 1.383 1.826
English prose (n=24) 1.918 0.864 0.785 1.227 1.212
English QA (n=8) 1.619 0.387 0.354 0.801 1.094

On Punjabi prose, Sevak-97M predicts text better than the 2B Qwen and close to the 4B. English is weak (the training mix was ~90% Punjabi).

An earlier single-domain table (200 "held-out" Phase C items, BPB 0.717) has been withdrawn: 16 of those 200 items are exact copies of training items and 76 share a training question, so it overstated the result.

Usage

import torch, sentencepiece as spm
from sevak_v2_model import SevakV2

model = SevakV2()
model.load_state_dict(torch.load("sevak_97m_sovereign_best.pt", map_location="cpu"))
model.eval()

sp = spm.SentencePieceProcessor(model_file="sehaj_bpe_16k.model")
ids = torch.tensor([sp.encode("ਸਵਾਲ: ਪੰਜਾਬ ਦੀ ਰਾਜਧਾਨੀ ਕੀ ਹੈ?\nਜਵਾਬ:")])
out = model.generate(ids, max_len=60, temp=0.7)
print(sp.decode(out[0].tolist()))

Lineage & Research

Built as the recipient model for the Gurmat knowledge-transplant research program (S1–S20):

  • Quantization Tax falsified (S16: 2-bit donors transplant as well as 8-bit)
  • Recipient Geometry Law (S19: S10's 0.993 cosine replicated; simple recipient geometry maps donors at 0.99, but high cosine ≠ high utility)
  • S20b (2026-10-04), negative: injecting mapped Qwen2.5-14B / Qwen3.5-4B activations at inference did not improve free-generation accuracy on 100 held-out questions (base 13%, injected 12%, shuffled-vector control 10%, random 12%). The mapped vector is question-specific (it beats the shuffled control on answer BPB) but adds no usable knowledge.

Limitations

  • Hallucinates facts; verify before use (14.7% factual accuracy above)
  • English is weak (training mix was ~90% Punjabi)
  • Does not use retrieved context yet
  • Not for medical/legal advice

Author

Gurpreet Singh Dhillon — AMRIT Research / Toon Studio ਉੱਤਮ ਮਤਿ ਰਿਦੈ ਤੁਮ੍ਹ ਆਓ

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support