--- license: mit pipeline_tag: text-generation library_name: transformers base_model: inclusionAI/Ling-3.0-flash base_model_relation: quantized tags: - moe - nvfp4 - w4a4 - compressed-tensors - vllm - hybrid-linear-attention - mtp --- # Ling-3.0-flash — NVFP4 **W4A4** (calibrated, MTP verified working) The first **W4A4** NVFP4 quantization of [Ling-3.0-flash](https://huggingface.co/inclusionAI/Ling-3.0-flash) (124B total / 5.1B active, hybrid KDA+MLA MoE) — weights **and activations** in NVFP4, calibrated with 128 samples via [llm-compressor](https://github.com/vllm-project/llm-compressor). | | | |---|---| | **Size** | **81.4 GB** (from 255 GB BF16) | | **Format** | compressed-tensors `nvfp4-pack-quantized`, W4A4 (E2M1 + FP8-E4M3 group-16 scales, static input global scales) | | **Measured** | **110 tok/s** single / **493 tok/s** @ 8-way generation, **~7,000 tok/s prefill** — 8×RTX PRO 2000 Blackwell (16 GB), TP=8, 32K context | | **Long context** | **131,072 tokens** verified (eager mode; prefill 5,900–8,000 tok/s) | | **MTP** | **Works: 93% acceptance** (k=1). Off by default — see below | | **Sanity** | 6/6 (math / logic / code / JP idiom / instruction-following / long-form JP); tool-calling 4/4 | ### Recommended: fp8 KV cache — 131K context at full speed With the bundled [`vllm_patch/triton_decode_attention.py`](./vllm_patch/triton_decode_attention.py) (1-line shared-memory fix for sm_120), `--kv-cache-dtype fp8` works and everything fits at once: | | context | KV budget | generation | prefill | |---|---|---|---|---| | **fp8 KV + patch** (recommended) | **131,072** | **458,752 tok** | **112.6 / 482–493 tok/s** (1 / 8-way) | 1.9k single / **6.5k @ 4-way** | | bf16 KV, 32K | 32,768 | ~187K tok | 110 / 493 tok/s | up to 7.6k single (4096 chunks) | | bf16 KV, 131K (`--enforce-eager`) | 131,072 | ~187K tok | 16.5 / 31.7 tok/s | 5.9k–8.0k | 458K tokens of KV = three concurrent 131K streams, or eight at 57K each. Quality battery under fp8 KV: 6/6 (math / logic / code / JP idiom / instruction-following / long-form JP) — no measured degradation vs bf16 KV. Single-stream prefill in the recommended config is chunking-bound (`--max-num-batched-tokens 2048`, the largest stable value at 131K on 16 GB); parallel prefill reaches ~6,500 tok/s. **Why context is cheap here but headroom is not.** Only 7 of 42 layers are MLA (the other 35 are KDA linear attention and carry no KV), so the KV cost is **8.2 KB/token** — roughly 9× cheaper per token than an all-MLA model like DeepSeek-V3. The constraint is the card: weights + non-torch overhead take **12.66 GB of a 16 GB GPU**, leaving ~1.5 GB for KV and KDA state — a total budget of about **187K tokens** to split between length and concurrency (131K × 1 stream, or 23K × 8 streams, etc.). On 80 GB cards this limit essentially vanishes. How this differs from [olka-fi/Ling-3.0-flash-NVFP4](https://huggingface.co/olka-fi/Ling-3.0-flash-NVFP4) (weight-only W4A16): this build quantizes activations too (W4A4) with real calibration data, and ships a working MTP path. ## Quantization boundary Same proven boundary as the W4A16 build: **routed experts only** (layers 2–41, `mlp.experts.*.{gate,up,down}_proj`, 120.8B of 127.4B params) → NVFP4 W4A4. Kept in BF16: shared experts, KDA & gated-MLA attention, router gate + expert_bias, dense layers 0–1, the MTP layer 42, embeddings, lm_head, norms. ## Serving (vLLM fork required) Upstream vLLM does not support `bailing_hybrid` v3. Use the prebuilt image `olkafi/vllm-bailing-v3` (inclusionAI/vllm @ `ling_3_0` + SwiGLU-clamp patch). This exact command reproduces every number in this card (131K context, fp8 KV, 458,752-token KV budget, ~112 tok/s single-stream). Both bundled patch files are mounted over the image's own copies — no rebuild needed, they are pure Python: ```bash REPO=/path/to/this-repo # e.g. $(huggingface-cli download sakamakismile/Ling-3.0-flash-W4A4-NVFP4) docker run --gpus all -d --name ling3 --ipc=host -p 8000:8000 \ -e NCCL_P2P_DISABLE=1 -e NCCL_CUMEM_ENABLE=0 \ -v $REPO:/models/ling3:ro \ -v $REPO/vllm_patch/unquantized.py:/opt/vllm-ling3/vllm/model_executor/layers/fused_moe/oracle/unquantized.py:ro \ -v $REPO/vllm_patch/triton_decode_attention.py:/opt/vllm-ling3/vllm/v1/attention/ops/triton_decode_attention.py:ro \ olkafi/vllm-bailing-v3:latest /models/ling3 \ --served-model-name ling3 --trust-remote-code --host 0.0.0.0 --port 8000 \ --tensor-parallel-size 8 --disable-custom-all-reduce \ --kernel-config '{"moe_backend":"marlin"}' --kv-cache-dtype fp8 \ --compilation-config '{"max_cudagraph_capture_size":16}' \ --gpu-memory-utilization 0.95 --max-model-len 131072 --max-num-seqs 8 \ --max-num-batched-tokens 2048 \ --enable-prefix-caching --mamba-cache-mode align \ --enable-auto-tool-choice --tool-call-parser ling3 --reasoning-parser ling3 ``` Notes for reproducers: - `--max-num-batched-tokens 2048`: 4096 OOMs at 131K on 16 GB cards (chunked-prefill activations); 2048 is stable. - On cards with more than 16 GB, drop the two patch mounts if you don't need them (`flashinfer_cutlass` may still be broken on sm_120 — keep `moe_backend=marlin` there), and raise `--max-num-batched-tokens` / `--max-num-seqs` freely. - NCCL flags are for machines without GPU P2P; harmless otherwise. Sampling (from the base card): `temperature=0.6, top_p=0.95, top_k=20`, thinking enabled. ### ⚠️ Known landmines (sm_120 / consumer Blackwell) 1. **`moe_backend=marlin` is mandatory.** The auto-selected `flashinfer_cutlass` backend **silently corrupts output** (endless `!!!!`) on sm_120 with EP, and plain TP=8 hits `NotImplementedError` (per-rank intermediate 768/8=96 needs unsupported padding). `cutlass` rejects EP. Marlin works and is the source of the numbers above. 2. **KV headroom is tight** — on 16 GB cards use `--gpu-memory-utilization 0.95`, `--max-num-seqs 8` and `--compilation-config '{"max_cudagraph_capture_size":16}'`. 32K fits comfortably that way; 131K needs `--enforce-eager` (see table above). 3. **`--kv-cache-dtype fp8` needs the bundled kernel patch on sm_120.** The TRITON_MLA fp8 decode kernel asks for 102,400 bytes of shared memory at `num_stages=2`; consumer Blackwell caps at 101,376 — 1 KB short. [`vllm_patch/triton_decode_attention.py`](./vllm_patch/triton_decode_attention.py) drops fp8-KV MLA to `num_stages=1`, which fits with no measured speed loss (112.6 tok/s single-stream, same as bf16 KV). 4. The bundled [`vllm_patch/unquantized.py`](./vllm_patch/unquantized.py) makes the *unquantized* MTP-layer MoE fall back to triton when `moe_backend=marlin` is forced globally — without it, serving with `--speculative-config` fails at startup. ## MTP / speculative decoding: verified working, off by default The W4A16 release reported 0% acceptance and shipped with MTP disabled. With this build + the bundled patch, MTP works: | config | single-stream | acceptance | |---|---|---| | no MTP (recommended) | **113.2 tok/s** | — | | MTP k=1 | 88.3 tok/s | **93.0%** | | MTP k=2 | 110.0 tok/s | 68%/tok | The draft (BF16 MTP layer on the triton path) currently costs more than the accepted tokens buy back on this hardware, so MTP is a proof-of-life, not a speed win — leave `--speculative-config` off for throughput. To try it: `--speculative-config '{"method":"mtp","num_speculative_tokens":1}'` (add `--compilation-config '{"max_cudagraph_capture_size":32}' --max-num-seqs 32` for memory). ## Bake provenance llm-compressor 0.12 (`lna-lab/llmc-bake:v0.24.0`), 128 calibration samples (`neuralmagic/calibration`, chat-templated) @ 2048 tokens, basic pipeline with CPU-resident weights and single-GPU onload calibration (fla/KDA kernels are triton/cuda-only, so pure-CPU forward is impossible). Recipe quirks that will bite reproducers: transformers 5.x needs three shims in the custom modeling file (`is_torch_fx_available`, `ROPE_INIT_FUNCTIONS['default']`, `_tied_weights_keys`); `rope_scaling` must be forced back to `None` after config normalization; llm-compressor drops `model_type` from the saved config (breaking vLLM's MLA detection — restore `"model_type": "bailing_hybrid"`); and post-calibration weight observation must be moved to CPU (`set_onload_device`) or 61,440 expert observers accumulate on one GPU and OOM. ## Acknowledgements Built on **inclusionAI's Ling-3.0-flash** (MIT). Quantization boundary and serving groundwork follow **olka-fi's** pioneering W4A16 release and vendor-fork patches. MIT license, inherited from the base model.