Sevak-97M Sovereign 🙏
ਪੰਜਾਬੀ-ਫਰਸਟ ਬਾਇਲਿੰਗੁਅਲ ਭਾਸ਼ਾ ਮਾਡਲ — ਲੈਪਟਾਪ 'ਤੇ ਸਿਖਲਾਈ, ਸਿਫ਼ਰ ਤੋਂ। (A Punjabi-first bilingual language model, trained from scratch on a laptop.)
Sevak-97M Sovereign is a 97.2M-parameter Punjabi-first bilingual (ਪੰਜਾਬੀ + English) language model, trained from scratch on a consumer laptop (Apple Silicon, MPS). No cloud, no GPU cluster.
Architecture
| Parameters | 97.2M |
| Layers | 12 (D=768) |
| Attention | RoPE + QK-Norm (standard causal attention), SwiGLU FFN |
| Embeddings | Tied (std=0.02 init) |
| Tokenizer | 16K BPE (Gurmukhi-optimized, 4.0 chars/token) |
| Context | 256 train / 1024 inference |
Training Pipeline (3 phases)
- Phase A: Punjabi corpus pretraining (40K chunks, 10K steps): loss 9.8 → 3.66
- Phase B: QA SFT (7.6K pairs: Bonsai-27B generated + verified SFT): → 2.10
- Phase C (Sovereign): 61.7K unique examples (rejection-sampled self-improvement winners, correction pass, sovereign knowledge, bilingual): 3.93 → 3.08
Note: the Phase B teacher (Bonsai-27B) itself scores only 35.7% on the factual test below and makes confident errors on basic Sikh history, so some of that generated data is likely wrong.
Evaluation
Factual accuracy (does it say the right answer?)
143 Punjabi factual questions (Sikh history, Punjab, general knowledge), decontaminated against all SFT sources (0 of them appear in Phase C data). Greedy decoding, prompt ਸਵਾਲ: …\nਜਵਾਬ:, scored by keyword/substring match against a reference answer. The reference answers were written with Gemini, so the scorer may slightly favour Gemini-style phrasing.
| Model | Size | Correct |
|---|---|---|
| Sevak-97M Sovereign | 97M | 14.7% |
| Sevak-97M Sovereign + Wikipedia retrieval in prompt | 97M | 8.4% |
| sahaj-86m (SFT) | 86M | 25.9% |
| Qwen3.5-4B (4-bit) | 4B | 21.7% |
| Bonsai-2-27B | 27B | 35.7% |
| Qwen2.5-14B-Instruct (4-bit) | 14B | 39.9% |
At this size the model is fluent in Punjabi but hallucinates facts. Example: asked when and where the Khalsa was founded, it answers "1869, Uttarakhand" (correct: 1699, Anandpur Sahib). It also does not yet use retrieved context (it was never trained on context→answer).
Bits-per-byte (language modelling, lower = better)
BPB measures how well a model predicts text, not whether it knows facts.
Multi-domain set of fresh texts written for this benchmark (not in training data). Sample sizes are small: 24 texts per prose bucket, 8 texts per QA bucket, so QA differences of a few hundredths are within noise.
| Domain | Sevak-97M | Qwen3.5-2B | Qwen3.5-4B | LFM2-350M | Gemma-3-1B |
|---|---|---|---|---|---|
| ਪੰਜਾਬੀ prose (n=24) | 0.625 | 0.865 | 0.592 | 1.846 | 2.333 |
| ਪੰਜਾਬੀ QA (n=8) | 0.547 | 0.579 | 0.382 | 1.383 | 1.826 |
| English prose (n=24) | 1.918 | 0.864 | 0.785 | 1.227 | 1.212 |
| English QA (n=8) | 1.619 | 0.387 | 0.354 | 0.801 | 1.094 |
On Punjabi prose, Sevak-97M predicts text better than the 2B Qwen and close to the 4B. English is weak (the training mix was ~90% Punjabi).
An earlier single-domain table (200 "held-out" Phase C items, BPB 0.717) has been withdrawn: 16 of those 200 items are exact copies of training items and 76 share a training question, so it overstated the result.
Usage
import torch, sentencepiece as spm
from sevak_v2_model import SevakV2
model = SevakV2()
model.load_state_dict(torch.load("sevak_97m_sovereign_best.pt", map_location="cpu"))
model.eval()
sp = spm.SentencePieceProcessor(model_file="sehaj_bpe_16k.model")
ids = torch.tensor([sp.encode("ਸਵਾਲ: ਪੰਜਾਬ ਦੀ ਰਾਜਧਾਨੀ ਕੀ ਹੈ?\nਜਵਾਬ:")])
out = model.generate(ids, max_len=60, temp=0.7)
print(sp.decode(out[0].tolist()))
Lineage & Research
Built as the recipient model for the Gurmat knowledge-transplant research program (S1–S20):
- Quantization Tax falsified (S16: 2-bit donors transplant as well as 8-bit)
- Recipient Geometry Law (S19: S10's 0.993 cosine replicated; simple recipient geometry maps donors at 0.99, but high cosine ≠ high utility)
- S20b (2026-10-04), negative: injecting mapped Qwen2.5-14B / Qwen3.5-4B activations at inference did not improve free-generation accuracy on 100 held-out questions (base 13%, injected 12%, shuffled-vector control 10%, random 12%). The mapped vector is question-specific (it beats the shuffled control on answer BPB) but adds no usable knowledge.
Limitations
- Hallucinates facts; verify before use (14.7% factual accuracy above)
- English is weak (training mix was ~90% Punjabi)
- Does not use retrieved context yet
- Not for medical/legal advice
Author
Gurpreet Singh Dhillon — AMRIT Research / Toon Studio ਉੱਤਮ ਮਤਿ ਰਿਦੈ ਤੁਮ੍ਹ ਆਓ