Caelum-G4-38B-A12.5B

Four Gemma 4 12B specialists. One expert active per token. Built for useful local agentic work.

Caelum-G4-38B-A12.5B is an experimental, top-1 sparse Mixture-of-Experts model assembled from four architecture-compatible Gemma 4 Unified 12B checkpoints. It contains 38.01B total parameters while activating approximately 12.53B parameters per token.

The model is intended for coding, structured tool use, agent planning, systems work, general reasoning, and careful scientific analysis. It preserves the official Gemma 4 Unified shared attention and multimodal trunk, but only text generation is currently release-tested.

This is an experimental community model, not an official Google model. Its positive result is from a small, frozen internal selection suite—not an established external leaderboard.

Why this model is interesting

  • Four full-width routed MLP experts: general, agentic/coding, systems/DevOps, and science/medical.
  • Top-1 routing at every one of the 48 language layers.
  • Calibrated router rather than prompt initialization alone.
  • Stock llama.cpp-compatible Gemma 4 MoE tensor layout.
  • Q4_K_M beat every dense source checkpoint on the frozen AirGapBench v1 suite.
  • The complete BF16 Transformers checkpoint loaded with zero missing, unexpected, or mismatched keys.

Sparse activation reduces compute relative to activating all 38B parameters, but it does not make the model a 12B download. All expert weights still need storage and normally need to be mapped or resident during inference.

Architecture

Property Value
Architecture Gemma 4 Unified sparse MoE
Total parameters 38,007,832,816
Active parameters per token 12,527,436,016
Language layers 48
Routed experts 4
Experts active per token 1
Routed expert intermediate width 15,360
Shared MLP intermediate width 1,024
Context configured by architecture 262,144 tokens
BF16 weight size 76,015,665,632 bytes

The 1,024-wide shared MLP remains deliberately zero-initialized in this development checkpoint. The calibrated router contains approximately 737K trainable routing weights. The 262K architectural context limit is not a claim that very long contexts are practical or validated on ordinary local hardware.

Source experts

Expert Source Intended specialization
0 google/gemma-4-12B-it General instruction following and multimodal fallback
1 Fable v2 Coding, terminal work, tool use, multi-step agents
2 Esper4 Systems, DevOps, infrastructure, incident response
3 Guardpoint Scientific and medical reasoning

Only each source's dense MLP tensors become routed expert weights. Embeddings, attention, normalization, LM head, tokenizer, chat template, and multimodal components come from the pinned official base. Exact source revisions and build receipts should be retained in this repository.

Evaluation

AirGapBench v1

AirGapBench v1 is a frozen 28-case internal selection suite with four cases in each category. It was authored and frozen before any candidate output was observed.

All models below used Q4_K_M GGUF, CPU-only stock llama.cpp, a 2,048-token context, temperature 0, a 256-token output limit, and thinking disabled. All 140 requests completed without request errors.

Model Overall Agent Coding DevOps Medical Science State Tools
Caelum-G4-38B-A12.5B 20/28 (71.43%) 3/4 3/4 2/4 4/4 2/4 3/4 3/4
Fable v2 18/28 (64.29%) 3/4 3/4 2/4 4/4 0/4 3/4 3/4
Gemma 4 12B IT 17/28 (60.71%) 3/4 2/4 1/4 4/4 2/4 2/4 3/4
Esper4 8/28 (28.57%) 1/4 0/4 1/4 3/4 0/4 1/4 2/4
Guardpoint 8/28 (28.57%) 1/4 1/4 0/4 2/4 0/4 3/4 1/4

On this exact suite, Caelum passed three cases that the official base failed and did not lose a case that the base passed. It led Fable by two cases. The suite SHA-256 is:

5bab1a4a7c27967fddb36af97b809b680b01191134a2b81f32032c24a3d821e4

These results are promising engineering evidence. With only 28 locally authored cases, they are not sufficient for a broad benchmark, SOTA, safety, or general-superiority claim. AirGapBench v0 was saturated: Caelum, Fable, and the official base each scored 10/12.

Transformers usage

The repository contains custom modeling code and requires a compatible Transformers release. Review the code before enabling trust_remote_code.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "EryriLabs/Caelum-G4-38B-A12.5B"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
    low_cpu_mem_usage=True,
)

messages = [
    {
        "role": "user",
        "content": "Inspect this deployment plan, identify the riskiest assumption, and propose a verification step.",
    }
]
inputs = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    enable_thinking=False,
    return_tensors="pt",
).to(next(model.parameters()).device)

with torch.inference_mode():
    output = model.generate(
        inputs,
        max_new_tokens=256,
        do_sample=False,
    )

print(tokenizer.decode(output[0, inputs.shape[-1]:], skip_special_tokens=True))

The validated development environment used Transformers 5.14.1. BF16 weights are approximately 76 GB, so most local users should prefer the GGUF release.

GGUF

The companion repository is intended for EryriLabs/Caelum-G4-38B-A12.5B-GGUF. Recommended release variants are:

  • Q4_K_M — default balance; already functionally and internally benchmarked.
  • Q5_K_M — higher-quality local option; functional smoke passed.
  • Q6_K — near-high-precision option after regression testing.
  • Q8_0 — large near-lossless reference quantization.
  • Q3_K_M — low-RAM option after regression testing.
  • IQ2_M — smallest experimental option; expect the largest quality risk.

Every quantization must be produced directly from F16/BF16, never from another quantization.

Known limitations

  • No established external evaluation has been completed.
  • AirGapBench v1 is small and internally authored.
  • The shared MLP is zero-initialized rather than trained or distilled.
  • Q4_K_M is the only quantization with the complete current AirGapBench v1 matrix.
  • A physical 32 GB RAM host has not yet been validated.
  • Long-context quality, memory use, and throughput have not been characterized.
  • Multimodal weights are present, but the current llama.cpp combined vision/audio projector fails during audio-frontend initialization before image inference. Do not describe this release as validated for image, video, or audio use yet.
  • A routed specialist can still be wrong, fabricate evidence, emit malformed tool calls, or choose an unsuitable expert.
  • The medical-safety result is four synthetic cases. This model is not a medical device and must not replace qualified professional judgment.
  • Outputs can reflect biases, inaccuracies, unsafe suggestions, or limitations inherited from every source model and dataset.

Responsible use

Use human review for consequential decisions, destructive shell actions, security changes, health information, and any output that affects other people. Sandboxed tools, least-privilege credentials, explicit action confirmation, and post-action verification are strongly recommended for agentic deployment.

Reproducibility files

Publish these small files alongside the weights:

  • caelum_mergekit.yaml — exact merge configuration.
  • caelum_manifest.json — measured architecture and parameter counts.
  • caelum_router.json and caelum_router.safetensors — calibrated router receipt and weights.
  • caelum_build_receipt.json — immutable build receipt.
  • The AirGapBench v1 suite, protocol, raw reports, and machine-readable matrix.

GGUF artifacts were created with stock llama.cpp commit 69bf6437914596fbbc4caf09a7ac16f2acdd1a94.

License and acknowledgements

Released under Apache-2.0, subject to the licenses and notices of all source checkpoints. Caelum is an unofficial derivative and is not endorsed by Google, the Gemma team, or the donor authors.

Thanks to Google DeepMind and the Gemma team, the Fable/Composer model authors, Valiant Labs, MergeKit, Hugging Face, and llama.cpp.

Downloads last month
147
Safetensors
Model size
38B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EryriLabs/Caelum-G4-38B-A12.5B

Collection including EryriLabs/Caelum-G4-38B-A12.5B