Instructions to use Mercity/pretrain-longcat-ngram-50pct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mercity/pretrain-longcat-ngram-50pct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Mercity/pretrain-longcat-ngram-50pct", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Mercity/pretrain-longcat-ngram-50pct", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Mercity/pretrain-longcat-ngram-50pct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mercity/pretrain-longcat-ngram-50pct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mercity/pretrain-longcat-ngram-50pct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Mercity/pretrain-longcat-ngram-50pct
- SGLang
How to use Mercity/pretrain-longcat-ngram-50pct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Mercity/pretrain-longcat-ngram-50pct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mercity/pretrain-longcat-ngram-50pct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Mercity/pretrain-longcat-ngram-50pct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mercity/pretrain-longcat-ngram-50pct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Mercity/pretrain-longcat-ngram-50pct with Docker Model Runner:
docker model run hf.co/Mercity/pretrain-longcat-ngram-50pct
Llama 1B with LongCat n-gram embeddings (50%), 6B tokens
A ~1B-parameter Llama 3-style decoder that spends about 48% of its parameter budget on hashed n-gram embedding tables, following the LongCat n-gram embedding module. To keep the total near 1B, it has 16 transformer layers instead of the baseline's 32. Everything else matches the baseline, including the 6B FineWeb training tokens.
This is a base model. It is not instruction-tuned or safety-tuned.
Results
Training
| Metric | Value | vs. baseline |
|---|---|---|
| Final train loss (step 3,053) | 2.6823 | +0.1124 |
| Final eval loss (step 3,000) | 2.6950 | +0.1043 |
| Final grad norm (step 3,053) | 0.0520 | +0.0066 |
| Peak grad norm after step 200 | 0.648 | +0.109 |
| Tokens / steps | 6B / 3,053 | same |
Zero-shot benchmarks
Scores from lm-eval on each task's full split. Shared-9 is the unweighted mean of the nine tasks.
| Benchmark | Metric | Score | vs. baseline |
|---|---|---|---|
| HellaSwag | acc_norm | 35.22 | β3.76 |
| WinoGrande | acc | 49.25 | β2.37 |
| ARC-Easy | acc_norm | 38.47 | β1.68 |
| ARC-Challenge | acc_norm | 22.87 | β0.76 |
| PIQA | acc_norm | 65.07 | β1.52 |
| OpenBookQA | acc_norm | 28.20 | β0.80 |
| CommonsenseQA | acc | 19.90 | +0.08 |
| SciQ | acc_norm | 63.30 | β0.20 |
| LAMBADA | acc | 33.57 | β4.18 |
| Shared-9 average | 39.54 | β1.69 | |
| Shared-9 average, 4-bit NF4 | 38.44 | β2.38 |
Few-NERD (LoRA fine-tuned)
| Metric | Score | vs. baseline |
|---|---|---|
| Micro F1 | 0.598 | β0.057 |
| Macro F1 | 0.535 | β0.060 |
| Sentence accuracy | 37.5% | β4.7 pts |
Usage
The architecture class ships with the checkpoint, so load it with trust_remote_code=True. Tested with transformers==5.8.0.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Mercity/pretrain-longcat-ngram-50pct"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
device_map="auto",
)
inputs = tokenizer("The capital of France is", return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Cached generation works: the model rebuilds the n-gram context from the full token history at each decoding step.
To score text instead of generating it, pass labels and read the loss:
batch = tokenizer("FineWeb is a large web-text dataset.", return_tensors="pt").to(model.device)
with torch.no_grad():
loss = model(**batch, labels=batch["input_ids"]).loss
print(f"loss={loss.item():.3f} ppl={loss.exp().item():.1f}")
When quantizing, keep the n-gram tables in higher precision. They are embedding lookups, not ordinary weight matrices.
Model details
| Setting | Value |
|---|---|
| Architecture | LlamaLongCatNgram (Llama 3-style dense decoder with n-gram input embeddings) |
| Total parameters | 1.037B |
| Share of parameters in n-gram tables | ~48% |
| Layers | 16 |
| N-gram orders | 2, 3, 4 |
| Hash tables | 6 (2 per order) |
| Hidden size | 1,536 |
| Intermediate size (SwiGLU) | 5,120 |
| Attention heads / KV heads | 12 / 6 (GQA) |
| QK normalization | On |
| Max sequence length | 8,192 |
| Tokenizer | Llama 2, 32,000 tokens |
| Embeddings | Tied input and output |
Training
| Setting | Value |
|---|---|
| Data | FineWeb sample-10BT, packed 8,192-token sequences |
| Tokens / steps | 6B / 3,053 |
| Batch | 12 per device Γ 20 gradient accumulation (~1.97M tokens per step) |
| Optimizer | Muon (LR 0.02, momentum 0.95, 5 Newton-Schulz steps, WD 0.1) + AdamW (LR 3e-4, Ξ² 0.9/0.95, WD 0.1) |
| Schedule | Cosine, 150 warmup steps |
| Hardware | 1 Γ NVIDIA B200, ~10.5 hours |
| Stack | TorchTitan, FlashAttention 4, Liger kernels |
Related checkpoints
| Model | Change from the baseline | Shared-9 |
|---|---|---|
| Baseline (QK-norm) | Reference model | 41.23 |
| KDA | 8 of 32 attention layers replaced with Kimi Delta Attention | 41.37 |
| N-gram 25% | ~25% of parameters in n-gram tables, 23 layers | 40.56 |
Limitations
Trained on 6B English web tokens only, a small budget for a 1B model. Benchmark scores are single-seed. The model will repeat or make up facts and has had no alignment training.
- Downloads last month
- 163