Instructions to use SixVolts/GLM-5.3-ewaste-edition-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL # Run inference directly in the terminal: llama cli -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL # Run inference directly in the terminal: llama cli -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Use Docker
docker model run hf.co/SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SixVolts/GLM-5.3-ewaste-edition-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SixVolts/GLM-5.3-ewaste-edition-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
- Ollama
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with Ollama:
ollama run hf.co/SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
- Unsloth Desktop
- Pi
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with Docker Model Runner:
docker model run hf.co/SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
- Lemonade
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Run and chat with the model
lemonade run user.GLM-5.3-ewaste-edition-GGUF-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use SixVolts/GLM-5.3-ewaste-edition-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SixVolts/GLM-5.3-ewaste-edition-GGUF:Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- GLM-5.3 — "e-waste edition" GGUF
GLM-5.3 — "e-waste edition" GGUF
This is a quantization of GLM-5.3 that is optimized to work better on older hardware. This quant was built using i-matrix GGUF quantizations with the same principle as the GLM-5.2 e-waste edition: K-quants for the routed experts, not codebook i-quants, because on pre-AVX-512 CPUs and previous-gen datacenter GPUs (MI100 / gfx908) expert dequantization, not bandwidth, is the decode bottleneck. GLM-5.3 is GLM-5.2 with further post-training (identical architecture and config), so the 5.2 recipes carry over tensor for tensor.
Available quantizations
| Variant | Size | Routed experts | Non-experts | Target |
|---|---|---|---|---|
GLM-5.3-Q4_K_XL |
403.8 GiB | Q4_K | Q8_0 | Highest quality; big-RAM CPU-expert boxes (≈512 GB) |
GLM-5.3-Q3_K_XL |
314.1 GiB | Q3_K | Q8_0 | Max quality on GPU; experts spill to CPU on <11×32 GB |
GLM-5.3-Q3_K_M ⭐ recommended |
297.5 GiB | Q3_K + Q2_K (cold layers) | Q6_K | Fits 0-spill on 10×32 GB (320 GB HBM) |
GLM-5.3-Q2_K_XL |
228.7 GiB | Q2_K + IQ2_XXS (13 coldest layers) | Q4_K | Fits 0-spill on 8×32 GB; fastest decode, lowest quality |
GLM-5.3-MTP-Q3_K_M-sc |
7.2 GB | — | — | self-contained NextN/MTP draft head (see Speculative decoding) |
All are imatrix-quantized with unsloth's GLM-5.3 importance matrix (209 chunks × 3584 tokens). The per-tensor types are the 5.2 files' types read from their GGUF headers, so each variant is the exact 5.3 counterpart of its 5.2 namesake, with one deliberate change (the MTP head, below).
What changed from the 5.2 build
- F32 intermediate, not BF16. GLM-5.3 ships as FP8 e4m3 with 128×128 block scales, and the scales are arbitrary
floats (194 of 95,040 in the first shard are powers of two). A BF16 intermediate re-rounds 88% of the dequantized
weights (typically 0.1%, up to 2⁻⁸ relative). The conversion here writes
fp8 × scalein F32 (3.0 TB) andllama-quantizeworks in f32 anyway, so the quantizer sees exactly the dequantized checkpoint. - Fewer tensors (1,524 vs 1,809). The checkpoint only carries DSA indexer weights on the 22
indexer_types: fulllayers; the other layers reuse the previous full layer's selection (what the reference implementation does). The 5.2 Q3_K_M / Q2_K_XL files were made by an older converter that duplicated the indexer into every layer; these files recordindexer_typesinstead. The 5.2_XLfiles already had the 1,524-tensor layout. - The MTP head (blk.78) is usable. The 5.2 files kept the NextN/MTP block's experts at Q2_K because mainline did
not use it. llama.cpp now drafts with it (
--spec-type draft-mtp), so here its experts are Q4_K and itsnextn.eh_projis Q6_K/Q8_0 in every variant (+2.1 GiB vs the 5.2 recipe). Measured on the head's own weights: the MoE block's time on an MI100 is within 30 µs between every 4-bit type (and q3_K is the slowest at 413 µs vs q4_K 253), while relative RMS error is q2_K 0.30, q3_K 0.15, q4_K 0.072, q5_K 0.036 — so the head costs nothing in speed and Q4_K is the best of the equal-cost options.
GLM-5.3-Q3_K_M — composition
| Tensor group | Quant | Notes |
|---|---|---|
| Routed experts gate/up, blk.35–77 (43 hottest MoE layers) | Q3_K | |
| Routed experts down, blk.35–77 | Q4_K | every hot ffn_down_exps |
| Routed experts, blk.3–34 (32 coldest MoE layers) | Q2_K | the cold cut, see below |
| Routed experts, blk.78 (MTP head) | Q4_K | changed from the 5.2 recipe's Q2_K |
Attention (q_a, q_b, kv_a_mqa, v_b, output), indexer attn_q_b |
Q6_K | |
attn_k_b |
Q4_0 | ncols=192 not divisible by 256 → automatic fallback |
Indexer attn_k, proj |
Q3_K | |
Shared expert (ffn_{gate,up,down}_shexp) |
Q6_K | fires every token |
| Dense FFN blk.0–2 | gate/up Q3_K, down Q5_K | |
token_embd, output |
Q6_K | |
norms, router (ffn_gate_inp, exp_probs_b) |
F32 |
Total: 297.45 GiB, 3.39 bpw, 1,524 tensors.
The cold cut is the same in 5.3
"Cold" is read from the imatrix (in_sum2, the sum of squared input activations per channel). The mean
ffn_down_exps importance is monotonic in depth and spans six orders of magnitude in 5.3 exactly as in 5.2:
blk.3 ≈ 0.003 ← coldest
blk.8 ≈ 2.3 (the one outlier, as in 5.2)
blk.20 ≈ 2.0
blk.34 ≈ 70
blk.35 ≈ 84
blk.77 ≈ 14860 ← hottest
so the 32 coldest MoE layers (blk.3–34) take Q2_K experts and the hot tail (blk.35–77) stays Q3_K/Q4_K.
GLM-5.3-Q2_K_XL, Q3_K_XL, Q4_K_XL — composition
Tensor for tensor the 5.2 files' types (Q2_K_XL: Q2_K experts with IQ2_XXS on blk.3–7 and 9–16, Q4_K
everything else, 6 shards; Q3_K_XL / Q4_K_XL: Q3_K / Q4_K experts, Q8_0 everything else, 8 / 11 shards), plus
the MTP head change above. All four were quantized from the same F32 intermediate with the same imatrix.
Quality (perplexity)
wikitext-2 test, 100 chunks × 512 tokens, f16 KV cache, -fa off (the DSA path), identical settings for all:
| Quant | Size | PPL ↓ | 5.2 counterpart (q8_0 KV) |
|---|---|---|---|
GLM-5.3-Q4_K_XL |
403.8 GiB | 2.6537 ± 0.032 | 2.6733 |
GLM-5.3-Q3_K_XL |
314.1 GiB | 2.7910 ± 0.034 | 2.8176 |
GLM-5.3-Q3_K_M |
297.5 GiB | 2.8092 ± 0.034 | 2.8348 |
GLM-5.3-Q2_K_XL |
228.7 GiB | 3.6057 ± 0.047 | 3.6129 |
The 5.2 column is the previous README's measurement (q8_0 KV cache); the two models and the KV setting differ, so compare within the 5.3 column.
Performance — 10 × MI100 (gfx908), 0-spill, Q3_K_M
Mainline-equivalent (dense DSA attention, no draft head): 11.5 tok/s decode, single stream, 16k context.
With this build's llama.cpp branch (fused sparse top-k attention for glm-dsa, the MTP head as a self-contained
drafter on its own GPU, per-tensor placement across the ten cards, pre-scaled int8 MMQ tiles for Q2_K/Q3_K on
gfx908): 19.6 tok/s decode at temperature 1.0 with 82–84% of drafted tokens accepted, ~110–140 tok/s prefill
at ubatch 512–1024, and 112k of context with the drafter still resident. The branch is
sixvolts/llama.cpp glm53/bringup; the pieces that matter for this model
are the sparse attention path (LLAMA_DSA_SPARSE=1), the Q2_K/Q3_K MMQ kernels (2–2.2× faster per MoE block on
gfx908), and a batch-size gate on HIP graphs that fixes a prefill corruption on this GPU class.
Speculative decoding with the MTP head
GLM-5.3-MTP-Q3_K_M-sc.gguf is blk.78 plus token_embd / output_norm / output, copied byte for byte from the
Q3_K_M file, so it can run as a separate draft model on a device the main model does not use:
llama-server -m GLM-5.3-Q3_K_M-00001-of-00008.gguf \
-md GLM-5.3-MTP-Q3_K_M-sc.gguf --spec-type draft-mtp --spec-draft-n-max 2 \
-sm layer -fa off -c 16384
Two drafted tokens per step is the sweet spot here (three loses on prose: acceptance falls to 0.56).
Usage
# Q3_K_M — 0-spill on ten 32 GB cards (first shard loads the set)
llama-server -m GLM-5.3-Q3_K_M-00001-of-00008.gguf -sm layer -fit on -fa off -c 16384
# Q4_K_XL — big-RAM CPU-expert box (≈512 GB): experts in system RAM, the rest on one GPU
llama-server -m GLM-5.3-Q4_K_XL-00001-of-00010.gguf -ngl 999 -ot 'ffn_.*_exps=CPU' -fa off -c 16384
Use -sm layer, not -sm tensor. Sampling per Z.ai: temperature 1.0, top-p 0.95, min-p 0.01.
Credits / provenance
- Base model: GLM-5.3 by zai-org (FP8 checkpoint, revision
aca966e4). - Importance matrix: unsloth/GLM-5.3-GGUF
imatrix_unsloth.gguf_file(not redistributed here). Thanks to the unsloth team for the calibration data. - Quantized with llama.cpp
convert_hf_to_gguf.py --outtype f32andllama-quantize --imatrix --tensor-type-file.
- Downloads last month
- 367
2-bit
3-bit
4-bit
Model tree for SixVolts/GLM-5.3-ewaste-edition-GGUF
Base model
zai-org/GLM-5.3