Llama 1B with LongCat n-gram embeddings (50%), 6B tokens

A ~1B-parameter Llama 3-style decoder that spends about 48% of its parameter budget on hashed n-gram embedding tables, following the LongCat n-gram embedding module. To keep the total near 1B, it has 16 transformer layers instead of the baseline's 32. Everything else matches the baseline, including the 6B FineWeb training tokens.

This is a base model. It is not instruction-tuned or safety-tuned.

Results

Training

Metric Value vs. baseline
Final train loss (step 3,053) 2.6823 +0.1124
Final eval loss (step 3,000) 2.6950 +0.1043
Final grad norm (step 3,053) 0.0520 +0.0066
Peak grad norm after step 200 0.648 +0.109
Tokens / steps 6B / 3,053 same

Zero-shot benchmarks

Scores from lm-eval on each task's full split. Shared-9 is the unweighted mean of the nine tasks.

Benchmark Metric Score vs. baseline
HellaSwag acc_norm 35.22 βˆ’3.76
WinoGrande acc 49.25 βˆ’2.37
ARC-Easy acc_norm 38.47 βˆ’1.68
ARC-Challenge acc_norm 22.87 βˆ’0.76
PIQA acc_norm 65.07 βˆ’1.52
OpenBookQA acc_norm 28.20 βˆ’0.80
CommonsenseQA acc 19.90 +0.08
SciQ acc_norm 63.30 βˆ’0.20
LAMBADA acc 33.57 βˆ’4.18
Shared-9 average 39.54 βˆ’1.69
Shared-9 average, 4-bit NF4 38.44 βˆ’2.38

Few-NERD (LoRA fine-tuned)

Metric Score vs. baseline
Micro F1 0.598 βˆ’0.057
Macro F1 0.535 βˆ’0.060
Sentence accuracy 37.5% βˆ’4.7 pts

Usage

The architecture class ships with the checkpoint, so load it with trust_remote_code=True. Tested with transformers==5.8.0.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "Mercity/pretrain-longcat-ngram-50pct"

tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

inputs = tokenizer("The capital of France is", return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Cached generation works: the model rebuilds the n-gram context from the full token history at each decoding step.

To score text instead of generating it, pass labels and read the loss:

batch = tokenizer("FineWeb is a large web-text dataset.", return_tensors="pt").to(model.device)
with torch.no_grad():
    loss = model(**batch, labels=batch["input_ids"]).loss
print(f"loss={loss.item():.3f}  ppl={loss.exp().item():.1f}")

When quantizing, keep the n-gram tables in higher precision. They are embedding lookups, not ordinary weight matrices.

Model details

Setting Value
Architecture LlamaLongCatNgram (Llama 3-style dense decoder with n-gram input embeddings)
Total parameters 1.037B
Share of parameters in n-gram tables ~48%
Layers 16
N-gram orders 2, 3, 4
Hash tables 6 (2 per order)
Hidden size 1,536
Intermediate size (SwiGLU) 5,120
Attention heads / KV heads 12 / 6 (GQA)
QK normalization On
Max sequence length 8,192
Tokenizer Llama 2, 32,000 tokens
Embeddings Tied input and output

Training

Setting Value
Data FineWeb sample-10BT, packed 8,192-token sequences
Tokens / steps 6B / 3,053
Batch 12 per device Γ— 20 gradient accumulation (~1.97M tokens per step)
Optimizer Muon (LR 0.02, momentum 0.95, 5 Newton-Schulz steps, WD 0.1) + AdamW (LR 3e-4, Ξ² 0.9/0.95, WD 0.1)
Schedule Cosine, 150 warmup steps
Hardware 1 Γ— NVIDIA B200, ~10.5 hours
Stack TorchTitan, FlashAttention 4, Liger kernels

Related checkpoints

Model Change from the baseline Shared-9
Baseline (QK-norm) Reference model 41.23
KDA 8 of 32 attention layers replaced with Kimi Delta Attention 41.37
N-gram 25% ~25% of parameters in n-gram tables, 23 layers 40.56

Limitations

Trained on 6B English web tokens only, a small budget for a 1B model. Benchmark scores are single-seed. The model will repeat or make up facts and has had no alignment training.

Downloads last month
163
Safetensors
Model size
1B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train Mercity/pretrain-longcat-ngram-50pct

Collection including Mercity/pretrain-longcat-ngram-50pct

Paper for Mercity/pretrain-longcat-ngram-50pct