Text Generation
MLX
Safetensors
xing4_0
quantization
apple-silicon
Mixture of Experts
mla
hyper-connections
xing
telechat
base_model_size:10B to 100B
conversational
custom_code
4-bit precision
Instructions to use TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 9,762 Bytes
2d29ae2 7d9990c 2d29ae2 7d9990c 2d29ae2 7d9990c 2d29ae2 7d9990c 2d29ae2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 | ---
base_model: XingChen-AGI/Xing4.0-29B-A4B
model_name: Xing4.0-29B-A4B-4bit-MLX
library_name: mlx
pipeline_tag: text-generation
license: apache-2.0
tags:
- mlx
- quantization
- apple-silicon
- moe
- mla
- hyper-connections
- xing
- telechat
- base_model:quantized:XingChen-AGI/Xing4.0-29B-A4B
- base_model_size:10B to 100B
---
# Xing4.0-29B-A4B-4bit-MLX — 4-bit MLX quant of Xing4.0-29B-A4B
Unofficial Apple Silicon quantization of **[XingChen-AGI/Xing4.0-29B-A4B](https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B)**,
produced with `mlx_lm.convert` (MLX `affine` quantizer, group size 64) on an Apple M5 Max / 128 GB.
I am not affiliated with China Telecom AI. All upstream weights, benchmarks and license terms belong
to them, and the upstream Apache-2.0 license governs this repository too (see [License](#license)).
> **Format note:** these are **MLX safetensors**, not GGUF. They will not load in llama.cpp / Ollama /
> LM Studio, and they will not load in PyTorch/vLLM/SGLang either. Use an MLX runtime.
## Before you download: you need `xing4_0` support in mlx-lm
Xing4.0 uses a new architecture (`model_type: xing4_0`) that **mlx-lm does not implement yet**.
Without it any MLX runtime stops with:
```
ValueError: Model type xing4_0 not supported.
```
An implementation exists and is verified against the upstream PyTorch code (see
[Provenance](#provenance)), but it is not merged into mlx-lm at the time of writing. Until it is,
these weights will not load anywhere. If you need them now, open an issue here and I will point you
at the model file.
## Pick a variant
| | [4bit](https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX) | [6bit](https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-6bit-MLX) | [8bit](https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-8bit-MLX) |
|---|---|---|---|
| Weights on disk | 15.51 GiB (16.65 GB), 4 shards | 22.37 GiB (24.02 GB), 5 shards | 29.23 GiB (31.38 GB), 6 shards |
| Effective precision | 4.514 bits per weight | 6.512 bits per weight | 8.509 bits per weight |
| Peak RAM, short prompt | 15.6 GB | 22.4 GB | 29.3 GB |
| Measured generation | 21.2 tok/s | 21.0 tok/s | 21.1 tok/s |
| Choose when | prioritize memory headroom | balance size and weight precision | prioritize weight precision and have the memory |
All three drop the multi-token-prediction layer (see below). Throughput is nearly identical across
the three because only ~4B parameters are active per token; the difference shows up in memory, not
speed. Task-level accuracy after quantization has **not** been measured. Runtime memory also depends
on context length and KV cache.
## What is inside (read from the shipped `config.json`)
| Field | Value |
|---|---|
| Architecture | `Xing4_0ForCausalLM` (`model_type: xing4_0`) |
| Parameters served | 29.51 B total, 4 B active per token |
| Layers / hidden | 40 layers, `hidden_size` 3584, dense FFN 9216 |
| Attention | MLA — `q_lora_rank` 768, `kv_lora_rank` 512, `qk_nope_head_dim` 128, `qk_rope_head_dim` 64, `v_head_dim` 128, 32 heads |
| MoE | 64 routed experts (`moe_intermediate_size` 1024) + 1 shared, 4 active per token, `noaux_tc` routing with sigmoid scoring, first 2 layers dense |
| Residual stream | mHC hyper-connections: `hc_mult` 4 parallel streams mixed by a Sinkhorn-normalized matrix (`hc_sinkhorn_iters` 20), two per layer |
| Position encoding | YaRN, `factor` 64 over `original_max_position_embeddings` 4096, interleaved RoPE |
| Context | `max_position_embeddings: 262144` |
| Vocab | 131,072 (tokenizer, `tokenization_xing4_0.py` and `chat_template.jinja` copied unchanged) |
| MTP | **dropped** — `num_nextn_predict_layers` normalized to 0 |
## Quantization recipe
- `mlx_lm.convert -q --q-bits 4 --q-group-size 64`, mode `affine`, source BF16
- 476 modules quantized: attention projections, all 64 routed experts + shared expert per MoE layer, the dense FFNs of layers 0–1, `embed_tokens` and `lm_head`
- effective **4.514 bits per weight** (scales and biases included)
- **never quantized:** all RMSNorms, the MoE router (`mlp.gate`), the hyper-connection tables (`hc_fn`, `hc_base` in BF16), and `hc_scale` / `e_score_correction_bias` kept in FP32
### About the dropped MTP layer
The upstream checkpoint carries a 41st layer (1.71 B parameters) implementing multi-token
prediction: `eh_proj`, `enorm`, `hnorm`, its own `embed_tokens`, a full attention + MoE block and a
`shared_head`. MLX has no speculative-decoding path for this architecture, so those tensors are
dropped and `num_nextn_predict_layers` is set to 0 to keep the shipped config self-consistent. If
you want MTP, use the upstream BF16 checkpoint with a runtime that supports it.
## Requirements
- Apple Silicon, macOS 15+ (built and smoke-tested on M5 Max, 128 GB)
- `mlx-lm` **with `xing4_0` support** — see the warning above
- roughly 15.6 GB of free unified memory for a short prompt, more for long context
## Usage
```python
from mlx_lm import load, generate
# the custom tokenizer is loaded from the repo, so both flags are needed
model, tokenizer = load(
"TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX",
tokenizer_config={"trust_remote_code": True},
trust_remote_code=True,
)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Explain Sinkhorn normalization in one sentence."}],
add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True))
```
`mlx_lm.load` forwards `trust_remote_code` to the model but not to the tokenizer, which is why
`tokenizer_config` carries its own flag. Without it you get an unrelated-looking
`AttributeError: 'PreTrainedConfig' object has no attribute 'max_position_embeddings'`.
The chat template supports `enable_thinking` (on by default) and tool calls. The model tends to
write its reasoning trace in Chinese even for English prompts; that is upstream behaviour, not a
quantization artifact.
### oMLX
oMLX 0.6.4 cannot run this architecture yet, and its oQ mixed-precision quantizer fails on it for
the same reason (its sensitivity pass cannot load the model). These are plain MLX quants, not oQ
builds.
## Recommended sampling
Upstream recommends, and the shipped `generation_config.json` matches:
| Scenario | temperature | top_p | repetition_penalty |
|---|---|---|---|
| Complex reasoning / general | 1.0 | 0.95 | 1.05 |
| Coding / agent tasks | 0.8 | 0.95 | 1.05 |
Note that `mlx_lm.generate` does not apply a repetition penalty unless you pass a logits processor.
## Provenance
The MLX implementation used to produce and load these weights was validated before quantizing:
| Check | Result |
|---|---|
| Hyper-connection vs the upstream PyTorch module, float32 | max relative error < 1e-5 |
| Full model vs `modeling_xing4_0.py`, small random-weight config | **max relative error 2.6e-07 on logits, 100% argmax agreement** |
| Upstream BF16 checkpoint, 58 GB | loads and generates coherent text |
| Each quant in this family | loads and generates coherent text |
Two upstream bugs found along the way, neither affecting these weights: the reference
`_init_weights` initializes `module.fn/base/scale` while the class defines `hc_fn/hc_base/hc_scale`
(random init from config fails, loading pretrained weights is unaffected), and `mlx_lm.load` does
not forward `trust_remote_code` to the tokenizer.
## Benchmarks
I publish no numbers I have not measured myself. The table below is **upstream's**, measured on the
BF16 model, and is not a measurement of these quantized weights:
| Benchmark | Xing4.0-29B-A4B (BF16, upstream) | this quant |
|---|---|---|
| IFBench | 69.67 | not measured |
| AIME2026 | 90.00 | not measured |
| AA.LCR | 61.00 | not measured |
| Tau3-Bench | 64.63 | not measured |
| Claw-Eval | 76.55 | not measured |
| SWE-bench Verified | 75.00 | not measured |
| Terminal-Bench 2.1 | 57.50 | not measured |
| SWE-bench Multilingual | 66.00 | not measured |
| DeepresearchBII | 60.80 | not measured |
Measurements and issue reports ("quant X broke task Y") are welcome and will be merged into this
table.
## Known caveats
- Quantization is lossy. If you see a regression, compare against a higher-precision variant and the
BF16 source before filing a bug.
- The hyper-connection mixing runs in the residual path of every layer and is kept in BF16/FP32 here;
its sensitivity to weight quantization elsewhere in the model has not been studied.
- 262k context is the architecture's limit, not a promise: keep the KV cache inside your memory
budget or the machine swaps.
- No MTP head, so no self-speculative decoding.
- Agentic and long-context behaviour at 4-bit is untested.
## License
Distributed under the **Apache License 2.0**, inherited from
[XingChen-AGI/Xing4.0-29B-A4B](https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B). See `LICENSE-NOTICE.md` in this repository.
## Citation
```bibtex
@misc{xing4-29b-a4b-mlx-4bit,
title = {Xing4.0-29B-A4B-4bit-MLX: MLX 4-bit quantization of Xing4.0-29B-A4B},
author = {TokenAI-zer},
year = {2026},
howpublished = {\url{https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX}},
note = {Unofficial quantization of XingChen-AGI/Xing4.0-29B-A4B}
}
@misc{liu2025trainingreporttelechat3moe,
title = {Training Report of TeleChat3-MoE},
author = {Xinzhang Liu and others},
year = {2025},
eprint = {2512.24157},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2512.24157}
}
```
## Acknowledgements
- **China Telecom AI (XingChen-AGI)** for Xing4.0-29B-A4B and the mHC architecture.
- **Apple MLX team** for `mlx` and `mlx-lm`, whose DeepSeek-V3 implementation this architecture
builds on directly.
|