Text Generation
Transformers
Safetensors
English
Chinese
mimo_v2
multimodal
vision-language
audio
agent
video-understanding
long-context
conversational
custom_code
Eval Results
8-bit precision
fp8
Instructions to use XiaomiMiMo/MiMo-V2.6-Pro-RL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use XiaomiMiMo/MiMo-V2.6-Pro-RL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="XiaomiMiMo/MiMo-V2.6-Pro-RL", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("XiaomiMiMo/MiMo-V2.6-Pro-RL", trust_remote_code=True, device_map="auto") - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use XiaomiMiMo/MiMo-V2.6-Pro-RL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "XiaomiMiMo/MiMo-V2.6-Pro-RL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XiaomiMiMo/MiMo-V2.6-Pro-RL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
- SGLang
How to use XiaomiMiMo/MiMo-V2.6-Pro-RL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "XiaomiMiMo/MiMo-V2.6-Pro-RL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XiaomiMiMo/MiMo-V2.6-Pro-RL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "XiaomiMiMo/MiMo-V2.6-Pro-RL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XiaomiMiMo/MiMo-V2.6-Pro-RL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use XiaomiMiMo/MiMo-V2.6-Pro-RL with Docker Model Runner:
docker model run hf.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Add model card
Browse files
README.md
CHANGED
|
@@ -1,3 +1,237 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
- zh
|
| 6 |
+
tags:
|
| 7 |
+
- text-generation
|
| 8 |
+
- multimodal
|
| 9 |
+
- vision-language
|
| 10 |
+
- audio
|
| 11 |
+
- agent
|
| 12 |
+
- video-understanding
|
| 13 |
+
- long-context
|
| 14 |
+
- mimo_v2
|
| 15 |
+
- transformers
|
| 16 |
+
library_name: transformers
|
| 17 |
---
|
| 18 |
+
|
| 19 |
+
<br/><br/>
|
| 20 |
+
|
| 21 |
+
<div align="center">
|
| 22 |
+
<picture>
|
| 23 |
+
<source srcset="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo_darkmode.png?raw=true" media="(prefers-color-scheme: dark)">
|
| 24 |
+
<img src="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo.png?raw=true" width="60%" alt="Xiaomi-MiMo" />
|
| 25 |
+
</picture>
|
| 26 |
+
</div>
|
| 27 |
+
|
| 28 |
+
<br/>
|
| 29 |
+
|
| 30 |
+
<div align="center" style="line-height: 1;">
|
| 31 |
+
|
|
| 32 |
+
<a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" target="_blank">🤗 HuggingFace</a>
|
| 33 |
+
|
|
| 34 |
+
<a href="https://mimo.xiaomi.com/mimo-v2-6" target="_blank">📰 Blog </a>
|
| 35 |
+
|
|
| 36 |
+
<a href="https://platform.xiaomimimo.com" target="_blank">🎨 Xiaomi MiMo API Platform </a>
|
| 37 |
+
|
|
| 38 |
+
<a href="https://aistudio.xiaomimimo.com" target="_blank">🗨️ Xiaomi MiMo Studio </a>
|
| 39 |
+
|
|
| 40 |
+
<a href="https://mimo.xiaomimimo.com/desktop/" target="_blank">💻 Xiaomi MiMo Desktop </a>
|
| 41 |
+
|
|
| 42 |
+
</div>
|
| 43 |
+
|
| 44 |
+
<br/>
|
| 45 |
+
|
| 46 |
+
<div align="center" style="line-height: 1.2;">
|
| 47 |
+
<strong>Community</strong><br/>
|
| 48 |
+
<a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro/blob/main/assets/wechat.jpg" target="_blank">WeChat Group</a>
|
| 49 |
+
|
|
| 50 |
+
<a href="https://discord.gg/kKC2kNnQEX" target="_blank">Discord</a>
|
| 51 |
+
|
|
| 52 |
+
<a href="https://t.me/+3T-I0pekOVIyNDBl" target="_blank">Telegram</a>
|
| 53 |
+
|
|
| 54 |
+
<a href="https://www.reddit.com/r/XiaomiMiMo_Official/" target="_blank">Reddit</a>
|
| 55 |
+
</div>
|
| 56 |
+
|
| 57 |
+
<br/>
|
| 58 |
+
|
| 59 |
+
# MiMo-V2.6-Pro-RL
|
| 60 |
+
|
| 61 |
+
**Scaling Reinforcement Learning Toward Self-Improvement**
|
| 62 |
+
|
| 63 |
+
<p align="center">
|
| 64 |
+
<a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf"><b>Technical Report</b></a>
|
| 65 |
+
</p>
|
| 66 |
+
|
| 67 |
+
## 1. Introduction
|
| 68 |
+
|
| 69 |
+
MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series. The series is built to **scale reinforcement learning toward self-improvement** — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include:
|
| 70 |
+
|
| 71 |
+
- **Native Omnimodal + Long Horizon**: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs.
|
| 72 |
+
- **You Only RL Once**: One mixed RL run across coding, general agents, visual, and cybersecurity — not separate per-domain runs. Tasks and multiple harnesses are mixed in the same batch so capabilities reinforce each other and strategies transfer to harnesses never seen in training.
|
| 73 |
+
- **Scaling RL Compute**: Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches — 1,568 prompts × 16 rollouts per step, billions of tokens per update.
|
| 74 |
+
- **Groupwise Agentic Grading (Self-Improvement Loop)**: Binary pass/fail cannot rank passing solutions, so the reward signal itself is scaled. An agentic grader compares rollouts *within each group*: **Groupwise Reward Synthesis (GRS)** builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes; **Groupwise Advantage Redistribution (GAR)** ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.
|
| 75 |
+
- **Aligned RL**: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
|
| 76 |
+
- **Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2)**: After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix), reusing histories from teacher trajectories and SFT demonstrations so decision points train without regenerating preceding turns — extending capabilities to hard-to-verify tasks.
|
| 77 |
+
|
| 78 |
+
## Model Summary
|
| 79 |
+
|
| 80 |
+
- **Architecture**: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters
|
| 81 |
+
- **Context Length**: 1M tokens
|
| 82 |
+
- **Modalities**: Text, Image, Video, Audio
|
| 83 |
+
- **Vision Encoder**: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
|
| 84 |
+
- **Audio Encoder**: 308M AudioTokenizer + 127M audio patch encoder
|
| 85 |
+
- **Multi-Token Prediction (MTP)**: 5-layer speculative decoder
|
| 86 |
+
|
| 87 |
+

|
| 88 |
+
|
| 89 |
+
*Figure 1. MiMo-V2.6 architecture.*
|
| 90 |
+
|
| 91 |
+
## 2. Downloads
|
| 92 |
+
|
| 93 |
+
| Model | Download |
|
| 94 |
+
| --- | --- |
|
| 95 |
+
| **MiMo-V2.6-Pro-RL** | [🤗 HuggingFace](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) · 🤖 ModelScope *(at release)* |
|
| 96 |
+
| **MiMo-V2.6-Flash-RL** | [🤗 HuggingFace](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) · 🤖 ModelScope *(at release)* |
|
| 97 |
+
|
| 98 |
+
## 3. Evaluation Results
|
| 99 |
+
|
| 100 |
+
| Benchmark | MiMo-V2.6 Pro | MiMo-V2.6 Flash | MiMo-V2.5 Pro | Claude Opus 5 | GPT-5.6 Sol | Claude Fable 5 |
|
| 101 |
+
| --- | --- | --- | --- | --- | --- | --- |
|
| 102 |
+
| **Code Agent** | | | | | | |
|
| 103 |
+
| DeepSWE v1.1 | 71.9 | 67.9 | 19.0 | 74.0 | 73.0 | 70.0 |
|
| 104 |
+
| ProgramBench | 26.5 | 26.0 | 12.5 | 37.0 | 25.0 | 33.0 |
|
| 105 |
+
| MiMo Code Bench | 63.2 | 61.2 | 40.4 | 68.6 | 59.3 | - |
|
| 106 |
+
| **General Agent** | | | | | | |
|
| 107 |
+
| AutomationBench v1.0.6 | 53.1 | 52.3 | 16.0 | 50.3 | 45.8 | 46.2 |
|
| 108 |
+
| Toolathlon-Verified | 76.9 | 73.6 | 49.1 | 80.6 | 74.9 | 77.9 |
|
| 109 |
+
| GDPval-AA 2.1 | 1673 | - | 1107 | 1708 | 1588 | 1595 |
|
| 110 |
+
| Agents’ Last Exam | 31.6 | 27.6 | 13.2 | 31.6 | 30.8 | 25.7 |
|
| 111 |
+
| Terminal Bench 4.0 | 34.9 | 28.8 | 1.5 | 49.0 | 39.9 | 42.4 |
|
| 112 |
+
| Terminal Bench 2.1 | 89.9 | 87.6 | 65.2 | 89.1 | 88.8 | 84.3 |
|
| 113 |
+
| OSWorld-Verified | 82.0 | 80.8 | - | 83.4 | 83.0 | 86.0 |
|
| 114 |
+
| JobBench | 62.0 | 61.2 | 25.0 | 65.7 | 45.4 | 57.4 |
|
| 115 |
+
| **Cybersecurity** | | | | | | |
|
| 116 |
+
| CyberGym | 94.0 | 95.1 | 40.0 | - | - | - |
|
| 117 |
+
| MiMo Cyber Bench | 80.2 | 77.2 | 0.0 | - | - | - |
|
| 118 |
+
| ExploitGym | 17.8 | 6.0 | 0.2 | 22.1 | 30.3 | 28.4 |
|
| 119 |
+
| ExploitBench | 47.9 | 25.3 | 16.6 | 70.0 | 78.5 | 78.0 |
|
| 120 |
+
| SEC Bench Pro | 66.3 | 47.5 | 17.7 | - | 79.1 | - |
|
| 121 |
+
| **Visual Agent** | | | | | | |
|
| 122 |
+
| MiMo VisualCoding | 72.3 | 71.5 | - | 70.0 | 73.4 | 69.1 |
|
| 123 |
+
|
| 124 |
+
## 4. Model Architecture
|
| 125 |
+
|
| 126 |
+
### LLM Backbone
|
| 127 |
+
|
| 128 |
+
| Component | MiMo-V2.6-Pro-RL |
|
| 129 |
+
| --- | --- |
|
| 130 |
+
| Layers (Total / SWA / GA) | 70 / 60 / 10 |
|
| 131 |
+
| Hidden Size | 6144 |
|
| 132 |
+
| SWA Heads (Q/KV) | 128 / 8 |
|
| 133 |
+
| GA Heads (Q/KV) | 128 / 8 |
|
| 134 |
+
| Head Dimensions (QK / V) | 192 / 128 |
|
| 135 |
+
| Sliding Window Size | 128 |
|
| 136 |
+
| Routed Experts (Total / Activated) | 384 / 8 |
|
| 137 |
+
| Max Context Length | 1M |
|
| 138 |
+
| MTP / Speculative Decoder | 5 SWA layers, window 1024 |
|
| 139 |
+
|
| 140 |
+
The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.
|
| 141 |
+
|
| 142 |
+
### Vision Encoder (MiMo ViT)
|
| 143 |
+
|
| 144 |
+
| Configuration | Value |
|
| 145 |
+
| --- | --- |
|
| 146 |
+
| Layers (Total / SWA / GA) | 28 / 24 / 4 |
|
| 147 |
+
| Hidden Size | 1280 |
|
| 148 |
+
| Attention Heads (Q / KV) | 32 / 8 |
|
| 149 |
+
| Head Dimension | 64 |
|
| 150 |
+
| Patch Size (T × H × W) | 2 × 16 × 16 |
|
| 151 |
+
| Sliding Window (Left / Right) | 64 / 64 |
|
| 152 |
+
| Spatial Merge Size | 2 × 2 |
|
| 153 |
+
| Parameters | 681M |
|
| 154 |
+
|
| 155 |
+
### Audio Encoders
|
| 156 |
+
|
| 157 |
+
AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).
|
| 158 |
+
|
| 159 |
+
### Speculative Decoder
|
| 160 |
+
|
| 161 |
+
5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.
|
| 162 |
+
|
| 163 |
+
## 5. Deployment
|
| 164 |
+
|
| 165 |
+
For best performance, follow the [SGLang MiMo cookbook](https://docs.sglang.io/cookbook/autoregressive/Xiaomi/MiMo-V2.5). Docker image: `lmsysorg/sglang:latest`.
|
| 166 |
+
|
| 167 |
+
### SGLang
|
| 168 |
+
|
| 169 |
+
```bash
|
| 170 |
+
sglang serve \
|
| 171 |
+
--trust-remote-code \
|
| 172 |
+
--model-path XiaomiMiMo/MiMo-V2.6-Pro-RL \
|
| 173 |
+
--tp 16 \
|
| 174 |
+
--dp 2 \
|
| 175 |
+
--enable-dp-attention \
|
| 176 |
+
--mm-enable-dp-encoder \
|
| 177 |
+
--ep 16 \
|
| 178 |
+
--moe-a2a-backend deepep \
|
| 179 |
+
--moe-dense-tp-size 1 \
|
| 180 |
+
--mem-fraction-static 0.7 \
|
| 181 |
+
--max-running-requests 128 \
|
| 182 |
+
--chunked-prefill-size 32768 \
|
| 183 |
+
--page-size 64 \
|
| 184 |
+
--swa-full-tokens-ratio 0.3 \
|
| 185 |
+
--speculative-algorithm EAGLE \
|
| 186 |
+
--speculative-num-steps 3 \
|
| 187 |
+
--speculative-eagle-topk 1 \
|
| 188 |
+
--speculative-num-draft-tokens 4 \
|
| 189 |
+
--enable-multi-layer-eagle \
|
| 190 |
+
--reasoning-parser mimo \
|
| 191 |
+
--tool-call-parser mimo \
|
| 192 |
+
--host 0.0.0.0 \
|
| 193 |
+
--port 30000 \
|
| 194 |
+
--nnodes 2 \
|
| 195 |
+
--node-rank <node-rank> \
|
| 196 |
+
--dist-init-addr <node0-ip>:20000
|
| 197 |
+
```
|
| 198 |
+
|
| 199 |
+
### vLLM
|
| 200 |
+
|
| 201 |
+
Follow the [vLLM MiMo-V2.5 recipe](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.5). Pre-built image: `docker pull vllm/vllm-openai:mimov25-cu129`.
|
| 202 |
+
|
| 203 |
+
```bash
|
| 204 |
+
vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL \
|
| 205 |
+
--tensor-parallel-size 8 \
|
| 206 |
+
--trust-remote-code \
|
| 207 |
+
--gpu-memory-utilization 0.95 \
|
| 208 |
+
--max-model-len auto \
|
| 209 |
+
--reasoning-parser mimo \
|
| 210 |
+
--tool-call-parser mimo \
|
| 211 |
+
--enable-auto-tool-choice \
|
| 212 |
+
--generation-config vllm
|
| 213 |
+
```
|
| 214 |
+
|
| 215 |
+
Recommended sampling: `temperature=1.0`, `top_p=0.95`.
|
| 216 |
+
|
| 217 |
+
Also available in AI Studio, MiMo Code, Xiaomi MiMo Desktop, Xiaomi MiMo Open Platform API, and OpenRouter.
|
| 218 |
+
|
| 219 |
+
## Citation
|
| 220 |
+
|
| 221 |
+
```bibtex
|
| 222 |
+
@misc{mimo2026v26pro,
|
| 223 |
+
title={MiMo-V2.6-Pro-RL},
|
| 224 |
+
author={{Xiaomi MiMo Team}},
|
| 225 |
+
year={2026},
|
| 226 |
+
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
|
| 227 |
+
}
|
| 228 |
+
```
|
| 229 |
+
|
| 230 |
+
## Contact
|
| 231 |
+
|
| 232 |
+
For questions or feedback, reach us at [[email protected]](mailto:[email protected]) or join our community:
|
| 233 |
+
|
| 234 |
+
- [WeChat Group](https://work.weixin.qq.com/apph5/external_room/join/group_mng?plg_id=c417f99bd9014b5dd894daa8bfe19790&)
|
| 235 |
+
- [Discord](https://discord.gg/WX2R2uNp)
|
| 236 |
+
- [Telegram](https://t.me/+3T-I0pekOVIyNDBl)
|
| 237 |
+
- [Reddit](https://www.reddit.com/r/XiaomiMiMo_Official/)
|