Instructions to use EryriLabs/Caelum-G4-38B-A12.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use EryriLabs/Caelum-G4-38B-A12.5B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="EryriLabs/Caelum-G4-38B-A12.5B", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForMultimodalLM model = AutoModelForMultimodalLM.from_pretrained("EryriLabs/Caelum-G4-38B-A12.5B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use EryriLabs/Caelum-G4-38B-A12.5B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EryriLabs/Caelum-G4-38B-A12.5B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EryriLabs/Caelum-G4-38B-A12.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/EryriLabs/Caelum-G4-38B-A12.5B
- SGLang
How to use EryriLabs/Caelum-G4-38B-A12.5B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "EryriLabs/Caelum-G4-38B-A12.5B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EryriLabs/Caelum-G4-38B-A12.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "EryriLabs/Caelum-G4-38B-A12.5B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EryriLabs/Caelum-G4-38B-A12.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use EryriLabs/Caelum-G4-38B-A12.5B with Docker Model Runner:
docker model run hf.co/EryriLabs/Caelum-G4-38B-A12.5B
Caelum-G4-38B-A12.5B
Four Gemma 4 12B specialists. One expert active per token. Built for useful local agentic work.
Caelum-G4-38B-A12.5B is an experimental, top-1 sparse Mixture-of-Experts model assembled from four architecture-compatible Gemma 4 Unified 12B checkpoints. It contains 38.01B total parameters while activating approximately 12.53B parameters per token.
The model is intended for coding, structured tool use, agent planning, systems work, general reasoning, and careful scientific analysis. It preserves the official Gemma 4 Unified shared attention and multimodal trunk, but only text generation is currently release-tested.
This is an experimental community model, not an official Google model. Its positive result is from a small, frozen internal selection suite—not an established external leaderboard.
Why this model is interesting
- Four full-width routed MLP experts: general, agentic/coding, systems/DevOps, and science/medical.
- Top-1 routing at every one of the 48 language layers.
- Calibrated router rather than prompt initialization alone.
- Stock llama.cpp-compatible Gemma 4 MoE tensor layout.
- Q4_K_M beat every dense source checkpoint on the frozen AirGapBench v1 suite.
- The complete BF16 Transformers checkpoint loaded with zero missing, unexpected, or mismatched keys.
Sparse activation reduces compute relative to activating all 38B parameters, but it does not make the model a 12B download. All expert weights still need storage and normally need to be mapped or resident during inference.
Architecture
| Property | Value |
|---|---|
| Architecture | Gemma 4 Unified sparse MoE |
| Total parameters | 38,007,832,816 |
| Active parameters per token | 12,527,436,016 |
| Language layers | 48 |
| Routed experts | 4 |
| Experts active per token | 1 |
| Routed expert intermediate width | 15,360 |
| Shared MLP intermediate width | 1,024 |
| Context configured by architecture | 262,144 tokens |
| BF16 weight size | 76,015,665,632 bytes |
The 1,024-wide shared MLP remains deliberately zero-initialized in this development checkpoint. The calibrated router contains approximately 737K trainable routing weights. The 262K architectural context limit is not a claim that very long contexts are practical or validated on ordinary local hardware.
Source experts
| Expert | Source | Intended specialization |
|---|---|---|
| 0 | google/gemma-4-12B-it | General instruction following and multimodal fallback |
| 1 | Fable v2 | Coding, terminal work, tool use, multi-step agents |
| 2 | Esper4 | Systems, DevOps, infrastructure, incident response |
| 3 | Guardpoint | Scientific and medical reasoning |
Only each source's dense MLP tensors become routed expert weights. Embeddings, attention, normalization, LM head, tokenizer, chat template, and multimodal components come from the pinned official base. Exact source revisions and build receipts should be retained in this repository.
Evaluation
AirGapBench v1
AirGapBench v1 is a frozen 28-case internal selection suite with four cases in each category. It was authored and frozen before any candidate output was observed.
All models below used Q4_K_M GGUF, CPU-only stock llama.cpp, a 2,048-token context, temperature 0, a 256-token output limit, and thinking disabled. All 140 requests completed without request errors.
| Model | Overall | Agent | Coding | DevOps | Medical | Science | State | Tools |
|---|---|---|---|---|---|---|---|---|
| Caelum-G4-38B-A12.5B | 20/28 (71.43%) | 3/4 | 3/4 | 2/4 | 4/4 | 2/4 | 3/4 | 3/4 |
| Fable v2 | 18/28 (64.29%) | 3/4 | 3/4 | 2/4 | 4/4 | 0/4 | 3/4 | 3/4 |
| Gemma 4 12B IT | 17/28 (60.71%) | 3/4 | 2/4 | 1/4 | 4/4 | 2/4 | 2/4 | 3/4 |
| Esper4 | 8/28 (28.57%) | 1/4 | 0/4 | 1/4 | 3/4 | 0/4 | 1/4 | 2/4 |
| Guardpoint | 8/28 (28.57%) | 1/4 | 1/4 | 0/4 | 2/4 | 0/4 | 3/4 | 1/4 |
On this exact suite, Caelum passed three cases that the official base failed and did not lose a case that the base passed. It led Fable by two cases. The suite SHA-256 is:
5bab1a4a7c27967fddb36af97b809b680b01191134a2b81f32032c24a3d821e4
These results are promising engineering evidence. With only 28 locally authored cases, they are not sufficient for a broad benchmark, SOTA, safety, or general-superiority claim. AirGapBench v0 was saturated: Caelum, Fable, and the official base each scored 10/12.
Transformers usage
The repository contains custom modeling code and requires a compatible Transformers release. Review the code before enabling trust_remote_code.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = "EryriLabs/Caelum-G4-38B-A12.5B"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
low_cpu_mem_usage=True,
)
messages = [
{
"role": "user",
"content": "Inspect this deployment plan, identify the riskiest assumption, and propose a verification step.",
}
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
).to(next(model.parameters()).device)
with torch.inference_mode():
output = model.generate(
inputs,
max_new_tokens=256,
do_sample=False,
)
print(tokenizer.decode(output[0, inputs.shape[-1]:], skip_special_tokens=True))
The validated development environment used Transformers 5.14.1. BF16 weights are approximately 76 GB, so most local users should prefer the GGUF release.
GGUF
The companion repository is intended for EryriLabs/Caelum-G4-38B-A12.5B-GGUF. Recommended release variants are:
Q4_K_M— default balance; already functionally and internally benchmarked.Q5_K_M— higher-quality local option; functional smoke passed.Q6_K— near-high-precision option after regression testing.Q8_0— large near-lossless reference quantization.Q3_K_M— low-RAM option after regression testing.IQ2_M— smallest experimental option; expect the largest quality risk.
Every quantization must be produced directly from F16/BF16, never from another quantization.
Known limitations
- No established external evaluation has been completed.
- AirGapBench v1 is small and internally authored.
- The shared MLP is zero-initialized rather than trained or distilled.
- Q4_K_M is the only quantization with the complete current AirGapBench v1 matrix.
- A physical 32 GB RAM host has not yet been validated.
- Long-context quality, memory use, and throughput have not been characterized.
- Multimodal weights are present, but the current llama.cpp combined vision/audio projector fails during audio-frontend initialization before image inference. Do not describe this release as validated for image, video, or audio use yet.
- A routed specialist can still be wrong, fabricate evidence, emit malformed tool calls, or choose an unsuitable expert.
- The medical-safety result is four synthetic cases. This model is not a medical device and must not replace qualified professional judgment.
- Outputs can reflect biases, inaccuracies, unsafe suggestions, or limitations inherited from every source model and dataset.
Responsible use
Use human review for consequential decisions, destructive shell actions, security changes, health information, and any output that affects other people. Sandboxed tools, least-privilege credentials, explicit action confirmation, and post-action verification are strongly recommended for agentic deployment.
Reproducibility files
Publish these small files alongside the weights:
caelum_mergekit.yaml— exact merge configuration.caelum_manifest.json— measured architecture and parameter counts.caelum_router.jsonandcaelum_router.safetensors— calibrated router receipt and weights.caelum_build_receipt.json— immutable build receipt.- The AirGapBench v1 suite, protocol, raw reports, and machine-readable matrix.
GGUF artifacts were created with stock llama.cpp commit 69bf6437914596fbbc4caf09a7ac16f2acdd1a94.
License and acknowledgements
Released under Apache-2.0, subject to the licenses and notices of all source checkpoints. Caelum is an unofficial derivative and is not endorsed by Google, the Gemma team, or the donor authors.
Thanks to Google DeepMind and the Gemma team, the Fable/Composer model authors, Valiant Labs, MergeKit, Hugging Face, and llama.cpp.
- Downloads last month
- 147