Access request

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This repository is public metadata with a manual download gate. Submit an accurate intended-use statement. Approval is discretionary and may be revoked for misuse or redistribution without authorization.

Log in or Sign Up to review the conditions and access this model content.

GLM-5.3-DERISKED-UD-Q3_K_XL

Client drop · full GLM-5.3 MoE · Unsloth UD-Q3_K_XL GGUF · Blackfrost DWM a=3

Built by Blackfrost · Las Vegas, Nevada

Format Access Edit Status

Public metadata, manual download gate. This is a client delivery of the GLM-5.3 de-risked Q3 GGUF. Request access on the Hub. Do not redistribute the weights without authorization.

What this is

Nine-shard Unsloth UD-Q3_K_XL GGUF of GLM-5.3 (full ~753B MoE, not Flash) after a Blackfrost direction-weight modification (DWM) pass.

The intended behavior is in the weights. It does not depend on a system prompt, LoRA, or a runtime filter. Production DWM details are proprietary and are not disclosed beyond the locked recipe below.

This exact Q3 artifact has not been through the R1-HARMFUL-BENCH-450 judged suite. Do not copy NVFP4/BF16 refusal percentages onto this GGUF.

Specifications

Architecture GlmMoeDsaForCausalLM (GLM-5.3 full, DSA MoE)
Parameters Full ~753B MoE topology · no expert pruning
Quant Unsloth UD-Q3_K_XL GGUF (mixed K-quants; expert down-proj mostly IQ4_XS / Q6_K)
Artifact size 342,965,977,029 bytes (319.4 GiB)
Shards 9 (00001-of-00009 … 00009-of-00009)
Layers 78 main layers + MTP head at blk.78
Hidden size 6,144
Experts 256 routed · top-8 / token · shared expert
Context Native GLM-5.3 long context (set --ctx-size from RAM, not the architectural ceiling)
Runtime llama.cpp / compatible GGUF loaders
Languages English and Chinese

Lineage

zai-org/GLM-5.3-BF16
  └─ unsloth/GLM-5.3-GGUF  tree 346b3591c7f28d1a23716f97a065ecf12ec14771
       └─ UD-Q3_K_XL (9 shards, 342,965,977,029 B)
            └─ Blackfrost DWM a=3 skip-early 2  ← this repository
Upstream BF16 zai-org/GLM-5.3-BF16
Quant parent unsloth/GLM-5.3-GGUF UD-Q3_K_XL/
Quant tree 346b3591c7f28d1a23716f97a065ecf12ec14771
Blackfrost change Weight-level DWM on an independent copy of the Q3 GGUF
Not applied Extra SFT · DPO · RLHF · expert pruning · GGUF requant

The Unsloth source GGUF was left untouched. This repo is the independent candidate-a3-skip2 writer output.

DWM recipe (locked)

Alpha / scale 3.0
Passes 1
Skip-early 2 (layers 0–1 unchanged)
Norm restore off (--no-norm-restore)
Direction shape 78 × 6,144 f32
Writer gguf-dwm-writer · 8 threads · 2 expert workers
Edited surface 227 packed GGUF tensors · layers 2–77
MTP (blk.78) not in the 227-target set

Finish markers on the build host: DWM_COMPLETE processed=227 targets=227, SOURCE_UNTOUCHED_CANDIDATE_SIZES_OK, APPLY_OK.

Shard 00001-of-00009 is byte-identical to the Unsloth source (sha256 56d6d59fc554a84c503c2f786e6978d55681e42b8c01370afd923a6286d05e0a) because skip-early leaves the first-shard tensors alone. Later shards match size of the source and are the DWM-edited payload.

Files

file bytes
GLM-5.3-UD-Q3_K_XL-00001-of-00009.gguf 9,428,677
GLM-5.3-UD-Q3_K_XL-00002-of-00009.gguf 48,804,973,120
GLM-5.3-UD-Q3_K_XL-00003-of-00009.gguf 48,508,432,544
GLM-5.3-UD-Q3_K_XL-00004-of-00009.gguf 48,508,432,544
GLM-5.3-UD-Q3_K_XL-00005-of-00009.gguf 48,508,432,544
GLM-5.3-UD-Q3_K_XL-00006-of-00009.gguf 48,508,432,544
GLM-5.3-UD-Q3_K_XL-00007-of-00009.gguf 48,508,432,544
GLM-5.3-UD-Q3_K_XL-00008-of-00009.gguf 48,717,290,336
GLM-5.3-UD-Q3_K_XL-00009-of-00009.gguf 2,892,122,176
Total 342,965,977,029

Point llama.cpp at shard 00001. It will pull the rest from the same directory.

Load

Recent llama.cpp with GLM-5.3 / GLM-MoE DSA support. Example shape (tune context and GPU layers to the box):

llama-server \
  --model GLM-5.3-UD-Q3_K_XL-00001-of-00009.gguf \
  --ctx-size 32768 \
  --n-gpu-layers 99 \
  --jinja \
  --host 0.0.0.0 --port 8080

Sampling baseline used on the GLM-5.3 line: temperature 1.0, top-p 0.95, thinking on unless you explicitly disable it in the template.

This drop was not load-qualified on Spark after apply. Qualify locally before production: /health, a known-positive completion, and a known-negative control.

What this is not

  • Not GLM-5.3-Flash (different architecture and GGUF).
  • Not the BF16 or NVFP4 DERISKED safetensors products.
  • Not a 450-prompt judged refusal result. Those numbers belong to other checkpoints and must not be cited for this GGUF.
  • Not a safety-stock model. Output is untrusted. You own access control, logging, and policy.

License and access

Upstream GLM-5.3 remains under Z.AI's GLM-5.3 terms. This derivative is a gated Blackfrost client delivery. Recipients may not republish the weights. Export controls, local law, and the operator's authorization boundary still apply.

Responsible use

Intended for authorized security testing, evaluation, and controlled local inference. Provided as is, without warranty.


Blackfrost · Las Vegas · x.com/Blackfrost_AI

Downloads last month
-
GGUF
Model size
754B params
Architecture
glm-dsa
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for defiobi00/GLM-5.3-DERISKED-UD-Q3_K_XL

Quantized
(28)
this model