Instructions to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: ./llama-cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Use Docker
docker model run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- LM Studio
- Jan
- vLLM
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- Ollama
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Ollama:
ollama run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- Unsloth Desktop
- Pi
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Docker Model Runner:
docker model run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- Lemonade
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Run and chat with the model
lemonade run user.Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF-IQ2_XS
List all available models
lemonade list
- Hermes Agent
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Sakura — MiMo-V2.6-Flash-MOPD P160 (expert-pruned, SSD-streaming ready)
Sakura MiMo P160 is an experimental expert-pruned GGUF of Xiaomi's MiMo-V2.6-Flash-MOPD: 160 of 256 routed experts per MoE layer, about 53.79 GiB. It is built to run on a single 64 GB AMD Strix Halo (Radeon 8060S, Vulkan) with ~35 GiB on the GPU and the remaining experts computed from system RAM — the MiMo SSD Streaming setup described below. Focus: German instructions, long planning prompts and coding. Independent community build, not an official Xiaomi release.
Part of the Sakura Micro line: compact, locally runnable derivatives of large models (lines: Sakura, Sakura Mini, Sakura Micro).
Provenance
- Base model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD, the MOPD2 upgrade of MiMo-V2.6-Flash-RL (309B total / 15B active, 48 layers, 256 experts, top-8). Its routers are bit-identical to the RL checkpoint, so an expert selection measured on RL transfers exactly.
- Built from Xiaomi's original weights: only the kept experts were downloaded (HTTP ranges), the MXFP4 experts were repacked losslessly, FP8 tensors converted exactly like llama.cpp's converter (checked byte-identical against the reference on the RL release).
- Expert selection (per layer): protect the experts carrying 80 % of the full model's routing on its own German answers to long coding/planning prompts (on-policy measurement), then the decode-trace hot experts, then REAP saliency bands and routed-token share decide. Selections tuned only on foreign text scored better perplexity but looped in chat; measuring on the model's own outputs fixed that.
- Quantization: llama.cpp with an importance matrix (Baekpica's, cut to the kept experts). Experts IQ2_XXS/IQ2_XS in layers 12–35, IQ2_XS/IQ2_S elsewhere, down projections of layers 44/46/47 IQ3_XXS; attention IQ4_XS, dense FFN Q6_K, embeddings Q8_0, output Q6_K.
Release artifact
MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf
- Quantization: mixed low-bit, mostly IQ2_XS experts (details above);
IQ2_XSin the file name is the dominant type - Size: 57,761,339,136 Bytes (53.79 GiB)
- SHA-256:
90dbd32b5d41ec902da81d0c369a8649cb29c059d7d9624a0c2f49511fab3b1b - Parameters: ~195B total after pruning (from 309B), ~15B active per token (unchanged)
- The bundled
SHA256SUMSfile can be used to verify the download.
Measured results
| Build | Size | PPL German | PPL code | Chat checks |
|---|---|---|---|---|
| Sakura MiMo P160 (this release, MOPD) | 53.8 GiB | 6.02 | 5.80 | short 7/8 · German stress 6/6 · long texts 4/4 · roadmaps 2/2 at temperature 1.0 · no loops |
| same selection and recipe, RL checkpoint | 53.8 GiB | 6.02 | 5.87 | roadmap: 1/2 at temperature 0 (rumination loop), 2/2 at 0.6 |
| 144 experts, RL (cut from TrevorJS IQ2) | 53.5 GiB | 6.12 | 6.10 | short 7/8 · long 4/4 · roadmaps 2/2 at temperature 0 |
| full MiMo-V2.6-Flash-RL (TrevorJS IQ2, reference) | 91.9 GiB | 5.04 | 4.23 | — |
Decode speed of this release: 9–10.5 tok/s at 16k context and 5.6–7.7 tok/s at 32k, measured with the desktop in normal parallel use (details below).
Perplexity: held-out German and code text, 8 × 512 tokens, not used for selection or calibration (lower is better). Speed: AMD Ryzen AI Max+ 395 / Radeon 8060S, 64 GB, Vulkan, MiMo SSD Streaming setup below, measured with the desktop in normal parallel use (browser, chat apps; free RAM near zero) — see the speed table below for the context-size effect. Chat checks: short factual/arithmetic prompts, long free-form texts, six German mid-length answers and two long German roadmap/deployment prompts with fixed closing lines. These are small functional checks, not a claim of general quality versus the original model.
Why 160 experts
We built and tested several expert counts from the same MiMo-V2.6-Flash family on this machine (held-out PPL German / code; lower is better):
| Experts kept per layer | Size | PPL German | PPL code | Chat |
|---|---|---|---|---|
| 128 (hot-expert selection) | 48.0 GiB | 14.50 | 7.30 | German broken (word salad, loops) |
| 128 (German-protected) | 48.0 GiB | 5.97 | 6.99 | German fine, but code loops and format errors |
| 144 (on-policy) | 53.5 GiB | 6.12 | 6.10 | stable |
| 160 (on-policy, this release) | 53.8 GiB | 6.02 | 5.80 | stable, no loops at temperature 1.0 |
| 256 (full model, reference) | 91.9 GiB | 5.04 | 4.23 | — (2.6 tok/s here, mostly paged from NVMe) |
The 128/144-expert rows and the full-model reference are MiMo-V2.6-Flash-RL builds. The same 160-expert selection on the RL checkpoint scored 6.02 / 5.87, so the step from 144 to 160 helps on its own; MOPD adds a little on code.
Smaller REAP/pruning levels cost noticeable quality — at 128 experts German or code breaks down. 160 experts give the best quality among the variants that still run at a usable speed on a 64 GB machine: the file is barely larger than the 144-expert build because the middle layers use slightly smaller quant types. Keeping more experts would push more weights off the GPU and cost speed. 160 is the sweet spot among the variants we tested for 64 GB; we did not build 176+.
Running it with stock llama.cpp
Any recent llama.cpp with mimo2 support loads this GGUF. With enough memory for the whole file on the GPU:
llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 32768 --temp 1.0 --top-p 0.95
Use temperature 1.0 / top-p 0.95 (Xiaomi's recommendation): at 0–0.6, long thinking tasks sometimes fell into rumination loops in our tests; at 1.0 they did not.
MiMo SSD Streaming — running it with less VRAM
The model is larger than a 32 GB GPU. MiMo SSD Streaming keeps the attention, the outlier-heavy early layers and the
late layers on the GPU and lets the CPU compute the experts of a block of middle layers (14–33, ~20 GiB) from RAM,
with the NVMe as backing store. Components (all in runtime/ of this repo):
| Component | What it does |
|---|---|
-ot "…ffn_(gate|up|down)_exps\.weight=CPU" |
keeps the experts of chosen layers in host memory (mmap); fits the model into ~35 GiB GPU |
Swap-MoE --expert-streaming (ek15072809/Swap-MoE) + our GPU guard patch |
demand-pages those experts from NVMe instead of loading them all up front |
Op-offload on (no --no-op-offload) |
prompt processing copies the used experts to the GPU (fast when the experts are in RAM) |
Expert lock (LLAMA_EXPERT_LOCK_LIST, our patch, optional) |
pins the most-used CPU-side experts in RAM (VirtualLock) so multi-token batches do not page-fault |
| UTF-8 sanitizing in chat parsing (our patch) | pruned models sometimes emit a broken emoji byte sequence; stock llama-server then fails with HTTP 500 — the reply is kept, broken bytes become U+FFFD |
Speed on this machine (decode, tok/s; desktop in parallel use, free RAM near zero):
| Mode | Context | Lock | Decode | Prefill (4k-token prompt) |
|---|---|---|---|---|
| fast (default) | 16k | none | 9–10.5 | 13–25 |
| long | 32k | hottest 30 % (~6 GiB) | 5.6–7.7 | 14–18 |
| (24k, for reference) | 24k | none | 6.6–7.4 | 32 |
All rows use the default -ub 512, which gives the best overall speed. A larger ubatch trades decode for prefill:
-ub |
Prefill (2k prompt) | Decode |
|---|---|---|
| 512 (default) | 12.7 | 9.6 |
| 1024 | 22.3 | 5.9 |
| 2048 | 34–39 (4k prompt) | ~5.3 |
When the CPU-side experts do not fit in RAM they are read from the NVMe once per ubatch, so larger ubatches speed up
prefill; but their bigger GPU buffers spill into shared system memory and take RAM away from those experts, which
nearly halves decode. A larger ubatch only pays off when the prompt is more than about twice as long as the answer
(for example a one-off read of large files); for coding with long reasoning keep 512. -ub 4096 ran out of GPU
memory on this machine.
The context size matters because the KV cache and GPU buffers spill into shared system memory; at 16k enough RAM
stays free for the CPU-side experts. On a quiet machine with the RAM to spare, the same runtime reached 8.9 tok/s
decode, ~120 tok/s prefill and 15/18/17 tok/s on 2/4/8-token batches with the expert lock (measured with our
144-expert build). n-gram speculation (--spec-type ngram-mod) helped the RL-based builds when returning edited files,
but not this MOPD build (about 50 % acceptance), so it is off by default.
Example for 64 GB Strix Halo (runtime/start_server_example.bat, long as first argument for the 32k mode):
llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 16384 -fit off --expert-streaming ^
-ot "blk\.(1[4-9]|2[0-9]|3[0-3])\.ffn_(gate|up|down)_exps\.weight=CPU" ^
--temp 1.0 --top-p 0.95
Smaller GPUs: move more middle layers to the CPU by widening the -ot layer range (each middle layer holds
~1 GiB of experts; keep layers 0–11 and 36–47 on the GPU if you can, they are the most sensitive). Speed depends on
whether the CPU-side experts fit in RAM: up to ~20 GiB in RAM decode at 6–10 tok/s here; with 56 GiB paged from NVMe
(the full, unpruned model) it was 2.6 tok/s. The lock list (hotlock_L14-33.txt) only applies to layers 14–33;
regenerate it with hot_lock_list.py for other ranges or leave the variable unset.
The runtime patches and build steps are in runtime/patches/ (llama.cpp f5e85d43a + Swap-MoE + ours). Stock
llama.cpp ignores LLAMA_EXPERT_LOCK_LIST and runs the model without that feature.
Known limitations
- Pruning removes capacity: expect less world knowledge than the original (e.g. a wrong statute number in a German legal answer) and occasional character-level typos in long German texts.
- First-shot code is mediocre; review it.
- Reversing a word letter by letter fails (all pruned variants).
- Experimental release; speed figures are from one machine.
License
MIT, like the base model. The license text is provided in LICENSE; sources and credits in NOTICE.md.
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura — MiMo-V2.6-Flash-MOPD P160(专家剪枝,支持 SSD 流式加载)
Sakura MiMo P160 是 Xiaomi 的 MiMo-V2.6-Flash-MOPD 的一个实验性专家剪枝 GGUF:每个 MoE 层 256 个路由专家中保留 160 个, 约 53.79 GiB。它的设计目标是在单台 64 GB 的 AMD Strix Halo(Radeon 8060S,Vulkan)上运行, 约 35 GiB 放在 GPU 上,其余专家由系统内存计算,即下文介绍的 MiMo SSD Streaming 方案。重点:德语指令、长规划提示和编程。 这是独立的社区构建版本,不是 Xiaomi 的官方发布。
属于 Sakura Micro 系列: 大模型的紧凑、可本地运行的衍生版本(系列: Sakura, Sakura Mini, Sakura Micro)。
来源
- 基础模型:XiaomiMiMo/MiMo-V2.6-Flash-MOPD,即 MiMo-V2.6-Flash-RL 的 MOPD2 升级版(总计 309B / 激活 15B,48 层,256 个专家,top-8)。它的路由器与 RL 检查点逐位相同,因此在 RL 上测得的专家选择可以精确迁移。
- 由 Xiaomi 的原始权重构建:只下载了保留的专家(HTTP range),MXFP4 专家被无损重新打包,FP8 张量的转换方式与 llama.cpp 的转换器完全相同(已在 RL 发布版上与参照逐字节核对)。
- 专家选择(逐层):保护那些承载了完整模型在 它自己的德语回答(针对长编程/规划提示,on-policy 测量)上 80 % 路由量的专家,然后由解码轨迹中的热点专家、REAP 显著性分段和路由 token 占比来决定。仅在外来文本上调好的选择虽然困惑度更好,但在聊天中会循环;改为在模型自己的输出上测量后,这一问题得到解决。
- 量化:使用 llama.cpp 和重要性矩阵(Baekpica 的,已裁剪到保留的专家)。第 12–35 层的专家为 IQ2_XXS/IQ2_XS,其余层为 IQ2_XS/IQ2_S,第 44/46/47 层的 down 投影为 IQ3_XXS;注意力为 IQ4_XS,稠密 FFN 为 Q6_K,嵌入为 Q8_0,输出为 Q6_K。
发布文件
MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf
- 量化:混合低比特,专家大部分为 IQ2_XS(详见上文);文件名中的
IQ2_XS是占主导的类型 - 大小:57,761,339,136 字节(53.79 GiB)
- SHA-256:
90dbd32b5d41ec902da81d0c369a8649cb29c059d7d9624a0c2f49511fab3b1b - 参数量:剪枝后总计约 195B(原为 309B),每个 token 激活约 15B(不变)
- 可使用随附的
SHA256SUMS文件校验下载。
测量结果
| 构建 | 大小 | PPL 德语 | PPL 代码 | 聊天检查 |
|---|---|---|---|---|
| Sakura MiMo P160(本次发布,MOPD) | 53.8 GiB | 6.02 | 5.80 | 短提示 7/8 · 德语压力测试 6/6 · 长文本 4/4 · 路线图 2/2(temperature 1.0)· 无循环 |
| 相同的选择和配方,RL 检查点 | 53.8 GiB | 6.02 | 5.87 | 路线图:temperature 0 时 1/2(反刍式循环),0.6 时 2/2 |
| 144 个专家,RL(由 TrevorJS IQ2 裁剪) | 53.5 GiB | 6.12 | 6.10 | 短提示 7/8 · 长文本 4/4 · 路线图 2/2(temperature 0) |
| 完整的 MiMo-V2.6-Flash-RL(TrevorJS IQ2,参照) | 91.9 GiB | 5.04 | 4.23 | — |
本次发布的解码速度:在桌面正常并行使用的情况下测得,16k 上下文时 9–10.5 tok/s,32k 时 5.6–7.7 tok/s(详见下文)。
困惑度:留出的德语和代码文本,8 × 512 个 token,未用于选择或校准(越低越好)。 速度:AMD Ryzen AI Max+ 395 / Radeon 8060S,64 GB,Vulkan,采用下文的 MiMo SSD Streaming 方案,在桌面正常并行使用(浏览器、聊天应用;空闲内存接近零)的情况下测得,上下文大小的影响见下面的速度表。 聊天检查:简短的事实/算术提示、长篇自由文本、六个德语中等长度回答,以及两个带固定结束语的德语长路线图/部署提示。这些只是小规模的功能性检查,不是关于相对于原始模型的整体质量的声明。
为什么是 160 个专家
我们在这台机器上,用同一个 MiMo-V2.6-Flash 系列构建并测试了几种专家数量 (留出的 PPL 德语 / 代码;越低越好):
| 每层保留的专家数 | 大小 | PPL 德语 | PPL 代码 | 聊天 |
|---|---|---|---|---|
| 128(热点专家选择) | 48.0 GiB | 14.50 | 7.30 | 德语崩坏(词语杂烩、循环) |
| 128(保护德语) | 48.0 GiB | 5.97 | 6.99 | 德语正常,但代码循环且格式出错 |
| 144(on-policy) | 53.5 GiB | 6.12 | 6.10 | 稳定 |
| 160(on-policy,本次发布) | 53.8 GiB | 6.02 | 5.80 | 稳定,temperature 1.0 下无循环 |
| 256(完整模型,参照) | 91.9 GiB | 5.04 | 4.23 | —(此处 2.6 tok/s,大部分从 NVMe 换页) |
128/144 个专家的行以及完整模型参照都是 MiMo-V2.6-Flash-RL 的构建。相同的 160 专家选择在 RL 检查点上得分为 6.02 / 5.87,所以从 144 到 160 本身就有帮助;MOPD 在代码上又略有提升。
较小的 REAP/剪枝程度会造成明显的质量损失:在 128 个专家时,德语或代码会崩坏。在 64 GB 机器上仍能以可用速度运行的变体中,160 个专家的质量最好:该文件只比 144 专家的构建略大一点,因为中间层使用了稍小的量化类型。保留更多专家会把更多权重推出 GPU,代价是速度。在我们测试过的变体中,160 是 64 GB 的最佳平衡点;我们没有构建 176 及以上。
使用原版 llama.cpp 运行
任何支持 mimo2 的较新 llama.cpp 都能加载这个 GGUF。如果 GPU 上有足够内存放下整个文件:
llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 32768 --temp 1.0 --top-p 0.95
请使用 temperature 1.0 / top-p 0.95(Xiaomi 的推荐):在我们的测试中,0–0.6 时,长思考任务有时会陷入反刍式循环;1.0 时则没有。
MiMo SSD Streaming — 用更少的显存运行
该模型比 32 GB 的 GPU 更大。MiMo SSD Streaming 把注意力、离群值较多的早期层和后期层保留在 GPU 上,让 CPU 从内存中计算中间若干层(14–33,约 20 GiB)的专家,并以 NVMe 作为后备存储。各组件(都在本仓库的 runtime/ 中):
| 组件 | 作用 |
|---|---|
-ot "…ffn_(gate|up|down)_exps\.weight=CPU" |
将选定层的专家保留在主机内存(mmap)中;使模型适配约 35 GiB 的 GPU |
Swap-MoE --expert-streaming(ek15072809/Swap-MoE)+ 我们的 GPU 防护补丁 |
按需从 NVMe 调入这些专家,而不是一开始就全部加载 |
Op-offload 开启(不使用 --no-op-offload) |
处理提示时将用到的专家复制到 GPU(专家在内存中时很快) |
专家锁(LLAMA_EXPERT_LOCK_LIST,我们的补丁,可选) |
把最常用的 CPU 侧专家固定在内存中(VirtualLock),使多 token 批处理不会触发缺页 |
| 聊天解析中的 UTF-8 清理(我们的补丁) | 剪枝后的模型有时会输出损坏的 emoji 字节序列;原版 llama-server 随后会以 HTTP 500 失败。现在回复会保留,损坏的字节变为 U+FFFD |
在这台机器上的速度(解码,tok/s;桌面并行使用,空闲内存接近零):
| 模式 | 上下文 | 锁 | 解码 | 预填充(4k token 提示) |
|---|---|---|---|---|
| fast(默认) | 16k | 无 | 9–10.5 | 13–25 |
| long | 32k | 最热的 30 %(约 6 GiB) | 5.6–7.7 | 14–18 |
| (24k,仅供参考) | 24k | 无 | 6.6–7.4 | 32 |
所有行都使用默认的 -ub 512,它的整体速度最好。更大的 ubatch 用解码换预填充:
-ub |
预填充(2k 提示) | 解码 |
|---|---|---|
| 512(默认) | 12.7 | 9.6 |
| 1024 | 22.3 | 5.9 |
| 2048 | 34–39(4k 提示) | ~5.3 |
当 CPU 侧的专家放不进内存时,每个 ubatch 会从 NVMe 读取一次,因此更大的 ubatch 会加快预填充;但它们更大的 GPU 缓冲区会溢出到共享系统内存,抢占这些专家所需的内存,使解码速度几乎减半。只有当提示长度超过回答长度约两倍时(例如一次性读取大文件),更大的 ubatch 才划算;对于带有长推理的编程,请保持 512。-ub 4096 在这台机器上耗尽了 GPU 内存。
上下文大小很重要,因为 KV 缓存和 GPU 缓冲区会溢出到共享系统内存;在 16k 时,仍有足够的内存留给 CPU 侧的专家。在内存充裕的安静机器上,同一个运行时在使用专家锁的情况下达到了 8.9 tok/s 解码、约 120 tok/s 预填充,以及 2/4/8 个 token 批处理时的 15/18/17 tok/s(用我们的 144 专家构建测得)。n-gram 推测解码(--spec-type ngram-mod)在返回编辑后的文件时对基于 RL 的构建有帮助,但对这个 MOPD 构建没有帮助(接受率约 50 %),因此默认关闭。
64 GB Strix Halo 的示例(runtime/start_server_example.bat,第一个参数为 long 时使用 32k 模式):
llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 16384 -fit off --expert-streaming ^
-ot "blk\.(1[4-9]|2[0-9]|3[0-3])\.ffn_(gate|up|down)_exps\.weight=CPU" ^
--temp 1.0 --top-p 0.95
较小的 GPU: 通过加宽 -ot 的层范围,将更多中间层移到 CPU(每个中间层约含 1 GiB 的专家;如果可以,请把 0–11 层和 36–47 层保留在 GPU 上,它们最敏感)。速度取决于 CPU 侧专家能否放进内存:最多约 20 GiB 放在内存中时,这里解码为 6–10 tok/s;当 56 GiB 从 NVMe 换页时(完整的未剪枝模型)为 2.6 tok/s。锁列表(hotlock_L14-33.txt)只适用于 14–33 层;对于其他范围,请用 hot_lock_list.py 重新生成,或者不设置该变量。
运行时补丁和构建步骤在 runtime/patches/ 中(llama.cpp f5e85d43a + Swap-MoE + 我们的补丁)。原版 llama.cpp 会忽略 LLAMA_EXPERT_LOCK_LIST,在没有该功能的情况下运行模型。
已知局限
- 剪枝会去除容量:预计世界知识会少于原始模型(例如在一个德语法律回答中写错法条编号),并且在长德语文本中偶尔出现字符级的拼写错误。
- 第一次生成的代码质量一般;请自行审查。
- 逐字母反转一个单词会失败(所有剪枝变体都是如此)。
- 实验性发布;速度数据来自一台机器。
许可证
- Downloads last month
- 670
2-bit
Model tree for webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF
Base model
XiaomiMiMo/MiMo-V2.6-Flash-MOPD