File size: 9,762 Bytes
2d29ae2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d9990c
2d29ae2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d9990c
2d29ae2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7d9990c
2d29ae2
7d9990c
2d29ae2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
---
base_model: XingChen-AGI/Xing4.0-29B-A4B
model_name: Xing4.0-29B-A4B-4bit-MLX
library_name: mlx
pipeline_tag: text-generation
license: apache-2.0
tags:
  - mlx
  - quantization
  - apple-silicon
  - moe
  - mla
  - hyper-connections
  - xing
  - telechat
  - base_model:quantized:XingChen-AGI/Xing4.0-29B-A4B
  - base_model_size:10B to 100B
---

# Xing4.0-29B-A4B-4bit-MLX — 4-bit MLX quant of Xing4.0-29B-A4B

Unofficial Apple Silicon quantization of **[XingChen-AGI/Xing4.0-29B-A4B](https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B)**,
produced with `mlx_lm.convert` (MLX `affine` quantizer, group size 64) on an Apple M5 Max / 128 GB.

I am not affiliated with China Telecom AI. All upstream weights, benchmarks and license terms belong
to them, and the upstream Apache-2.0 license governs this repository too (see [License](#license)).

> **Format note:** these are **MLX safetensors**, not GGUF. They will not load in llama.cpp / Ollama /
> LM Studio, and they will not load in PyTorch/vLLM/SGLang either. Use an MLX runtime.

## Before you download: you need `xing4_0` support in mlx-lm

Xing4.0 uses a new architecture (`model_type: xing4_0`) that **mlx-lm does not implement yet**.
Without it any MLX runtime stops with:

```
ValueError: Model type xing4_0 not supported.
```

An implementation exists and is verified against the upstream PyTorch code (see
[Provenance](#provenance)), but it is not merged into mlx-lm at the time of writing. Until it is,
these weights will not load anywhere. If you need them now, open an issue here and I will point you
at the model file.

## Pick a variant

| | [4bit](https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX) | [6bit](https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-6bit-MLX) | [8bit](https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-8bit-MLX) |
|---|---|---|---|
| Weights on disk | 15.51 GiB (16.65 GB), 4 shards | 22.37 GiB (24.02 GB), 5 shards | 29.23 GiB (31.38 GB), 6 shards |
| Effective precision | 4.514 bits per weight | 6.512 bits per weight | 8.509 bits per weight |
| Peak RAM, short prompt | 15.6 GB | 22.4 GB | 29.3 GB |
| Measured generation | 21.2 tok/s | 21.0 tok/s | 21.1 tok/s |
| Choose when | prioritize memory headroom | balance size and weight precision | prioritize weight precision and have the memory |

All three drop the multi-token-prediction layer (see below). Throughput is nearly identical across
the three because only ~4B parameters are active per token; the difference shows up in memory, not
speed. Task-level accuracy after quantization has **not** been measured. Runtime memory also depends
on context length and KV cache.

## What is inside (read from the shipped `config.json`)

| Field | Value |
|---|---|
| Architecture | `Xing4_0ForCausalLM` (`model_type: xing4_0`) |
| Parameters served | 29.51 B total, 4 B active per token |
| Layers / hidden | 40 layers, `hidden_size` 3584, dense FFN 9216 |
| Attention | MLA — `q_lora_rank` 768, `kv_lora_rank` 512, `qk_nope_head_dim` 128, `qk_rope_head_dim` 64, `v_head_dim` 128, 32 heads |
| MoE | 64 routed experts (`moe_intermediate_size` 1024) + 1 shared, 4 active per token, `noaux_tc` routing with sigmoid scoring, first 2 layers dense |
| Residual stream | mHC hyper-connections: `hc_mult` 4 parallel streams mixed by a Sinkhorn-normalized matrix (`hc_sinkhorn_iters` 20), two per layer |
| Position encoding | YaRN, `factor` 64 over `original_max_position_embeddings` 4096, interleaved RoPE |
| Context | `max_position_embeddings: 262144` |
| Vocab | 131,072 (tokenizer, `tokenization_xing4_0.py` and `chat_template.jinja` copied unchanged) |
| MTP | **dropped** — `num_nextn_predict_layers` normalized to 0 |

## Quantization recipe

- `mlx_lm.convert -q --q-bits 4 --q-group-size 64`, mode `affine`, source BF16
- 476 modules quantized: attention projections, all 64 routed experts + shared expert per MoE layer, the dense FFNs of layers 0–1, `embed_tokens` and `lm_head`
- effective **4.514 bits per weight** (scales and biases included)
- **never quantized:** all RMSNorms, the MoE router (`mlp.gate`), the hyper-connection tables (`hc_fn`, `hc_base` in BF16), and `hc_scale` / `e_score_correction_bias` kept in FP32

### About the dropped MTP layer

The upstream checkpoint carries a 41st layer (1.71 B parameters) implementing multi-token
prediction: `eh_proj`, `enorm`, `hnorm`, its own `embed_tokens`, a full attention + MoE block and a
`shared_head`. MLX has no speculative-decoding path for this architecture, so those tensors are
dropped and `num_nextn_predict_layers` is set to 0 to keep the shipped config self-consistent. If
you want MTP, use the upstream BF16 checkpoint with a runtime that supports it.

## Requirements

- Apple Silicon, macOS 15+ (built and smoke-tested on M5 Max, 128 GB)
- `mlx-lm` **with `xing4_0` support** — see the warning above
- roughly 15.6 GB of free unified memory for a short prompt, more for long context

## Usage

```python
from mlx_lm import load, generate

# the custom tokenizer is loaded from the repo, so both flags are needed
model, tokenizer = load(
    "TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX",
    tokenizer_config={"trust_remote_code": True},
    trust_remote_code=True,
)

prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Explain Sinkhorn normalization in one sentence."}],
    add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True))
```

`mlx_lm.load` forwards `trust_remote_code` to the model but not to the tokenizer, which is why
`tokenizer_config` carries its own flag. Without it you get an unrelated-looking
`AttributeError: 'PreTrainedConfig' object has no attribute 'max_position_embeddings'`.

The chat template supports `enable_thinking` (on by default) and tool calls. The model tends to
write its reasoning trace in Chinese even for English prompts; that is upstream behaviour, not a
quantization artifact.

### oMLX

oMLX 0.6.4 cannot run this architecture yet, and its oQ mixed-precision quantizer fails on it for
the same reason (its sensitivity pass cannot load the model). These are plain MLX quants, not oQ
builds.

## Recommended sampling

Upstream recommends, and the shipped `generation_config.json` matches:

| Scenario | temperature | top_p | repetition_penalty |
|---|---|---|---|
| Complex reasoning / general | 1.0 | 0.95 | 1.05 |
| Coding / agent tasks | 0.8 | 0.95 | 1.05 |

Note that `mlx_lm.generate` does not apply a repetition penalty unless you pass a logits processor.

## Provenance

The MLX implementation used to produce and load these weights was validated before quantizing:

| Check | Result |
|---|---|
| Hyper-connection vs the upstream PyTorch module, float32 | max relative error < 1e-5 |
| Full model vs `modeling_xing4_0.py`, small random-weight config | **max relative error 2.6e-07 on logits, 100% argmax agreement** |
| Upstream BF16 checkpoint, 58 GB | loads and generates coherent text |
| Each quant in this family | loads and generates coherent text |

Two upstream bugs found along the way, neither affecting these weights: the reference
`_init_weights` initializes `module.fn/base/scale` while the class defines `hc_fn/hc_base/hc_scale`
(random init from config fails, loading pretrained weights is unaffected), and `mlx_lm.load` does
not forward `trust_remote_code` to the tokenizer.

## Benchmarks

I publish no numbers I have not measured myself. The table below is **upstream's**, measured on the
BF16 model, and is not a measurement of these quantized weights:

| Benchmark | Xing4.0-29B-A4B (BF16, upstream) | this quant |
|---|---|---|
| IFBench | 69.67 | not measured |
| AIME2026 | 90.00 | not measured |
| AA.LCR | 61.00 | not measured |
| Tau3-Bench | 64.63 | not measured |
| Claw-Eval | 76.55 | not measured |
| SWE-bench Verified | 75.00 | not measured |
| Terminal-Bench 2.1 | 57.50 | not measured |
| SWE-bench Multilingual | 66.00 | not measured |
| DeepresearchBII | 60.80 | not measured |

Measurements and issue reports ("quant X broke task Y") are welcome and will be merged into this
table.

## Known caveats

- Quantization is lossy. If you see a regression, compare against a higher-precision variant and the
  BF16 source before filing a bug.
- The hyper-connection mixing runs in the residual path of every layer and is kept in BF16/FP32 here;
  its sensitivity to weight quantization elsewhere in the model has not been studied.
- 262k context is the architecture's limit, not a promise: keep the KV cache inside your memory
  budget or the machine swaps.
- No MTP head, so no self-speculative decoding.
- Agentic and long-context behaviour at 4-bit is untested.

## License

Distributed under the **Apache License 2.0**, inherited from
[XingChen-AGI/Xing4.0-29B-A4B](https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B). See `LICENSE-NOTICE.md` in this repository.

## Citation

```bibtex
@misc{xing4-29b-a4b-mlx-4bit,
  title        = {Xing4.0-29B-A4B-4bit-MLX: MLX 4-bit quantization of Xing4.0-29B-A4B},
  author       = {TokenAI-zer},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/TokenAI-zer/Xing4.0-29B-A4B-4bit-MLX}},
  note         = {Unofficial quantization of XingChen-AGI/Xing4.0-29B-A4B}
}

@misc{liu2025trainingreporttelechat3moe,
  title         = {Training Report of TeleChat3-MoE},
  author        = {Xinzhang Liu and others},
  year          = {2025},
  eprint        = {2512.24157},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2512.24157}
}
```

## Acknowledgements

- **China Telecom AI (XingChen-AGI)** for Xing4.0-29B-A4B and the mHC architecture.
- **Apple MLX team** for `mlx` and `mlx-lm`, whose DeepSeek-V3 implementation this architecture
  builds on directly.