|
Download README.md from festr2/glm51-nvfp4-mxfp8-l42-62-bf16src-20260526: direct link, hf CLI and curl.
- Browser
- Download file 2.12 kB
-
https://huggingface.co/festr2/glm51-nvfp4-mxfp8-l42-62-bf16src-20260526/resolve/f3a52ec785346624d2416877bfcdb0485adbf636/README.md
- Command line
-
hf download hf://festr2/glm51-nvfp4-mxfp8-l42-62-bf16src-20260526@f3a52ec785346624d2416877bfcdb0485adbf636/README.md
-
curl -L -o README.md https://huggingface.co/festr2/glm51-nvfp4-mxfp8-l42-62-bf16src-20260526/resolve/f3a52ec785346624d2416877bfcdb0485adbf636/README.md
2.12 kB
metadata
license: other
base_model:
- lukealonso/GLM-5.1-NVFP4-MTP
- zai-org/GLM-5.1
tags:
- glm
- glm-5.1
- nvfp4
- mxfp8
- modelopt
- vllm
- b12x
GLM-5.1 NVFP4 + MXFP8 L42-62
Mixed-precision GLM-5.1 checkpoint for local vLLM/B12X experiments.
Sparse expert layers 42-62 were rebuilt as ModelOpt MXFP8 from the BF16
zai-org/GLM-5.1 checkpoint. The remaining sparse expert layers stay NVFP4.
Important: GLM-5.1 stores its MTP / next-token-prediction layer as
model.layers.78.*. The FP8-PB-WO mixed base checkpoint had that layer in BF16,
so this checkpoint transplants model.layers.78.* from
lukealonso/GLM-5.1-NVFP4-MTP to keep the production NVFP4 MTP layer.
Build Sources
- Base mixed checkpoint:
/root/.cache/huggingface/hub/models--festr2--glm51-nvfp4-w4a16-fp8pbwo-l42-62-20260517/snapshots/e37f1787435d2b2c111a5f5eac924a556a06e257 - BF16 MXFP8 source:
/root/.cache/huggingface/hub/models--zai-org--GLM-5.1/snapshots/26e1bd6e011feb778d25ae34b09b07074139d92d - NVFP4 MTP source:
/root/.cache/huggingface/hub/models--lukealonso--GLM-5.1-NVFP4-MTP/snapshots/78b7fe365f3905b4e0261a85182fefdbd5137989
Reproduction
The builder script is here:
https://github.com/local-inference-lab/quant-toolkit/blob/codex/glm51-mxfp8-converter-20260526/tools/patch_glm51_experts_mxfp8_from_bf16.py
Commit:
https://github.com/local-inference-lab/quant-toolkit/commit/6f406718271f45b1365abf388a00483b66067fd8
Exact reproduction command is documented here:
https://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm5.1_mxfp8_checkpoint.md
Static Validation
Expected checkpoint structure:
model.layers.42-62.mlp.experts.* -> MXFP8
model.layers.78.* -> NVFP4 MTP
other sparse expert layers -> inherited NVFP4/mixed base state
Sample tensors:
model.layers.42.mlp.experts.0.gate_proj.weight
shape: [2048, 6144]
dtype: F8_E4M3
model.layers.42.mlp.experts.0.gate_proj.weight_scale
shape: [2048, 192]
dtype: U8
model.layers.78.mlp.experts.0.gate_proj.weight
shape: [2048, 3072]
dtype: U8