bwshen-mi commited on
Commit
c165b4c
·
verified ·
1 Parent(s): a949623

Add model card

Browse files
Files changed (1) hide show
  1. README.md +234 -0
README.md CHANGED
@@ -1,3 +1,237 @@
1
  ---
2
  license: mit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ language:
4
+ - en
5
+ - zh
6
+ tags:
7
+ - text-generation
8
+ - multimodal
9
+ - vision-language
10
+ - audio
11
+ - agent
12
+ - video-understanding
13
+ - long-context
14
+ - mimo_v2
15
+ - transformers
16
+ library_name: transformers
17
  ---
18
+
19
+ <br/><br/>
20
+
21
+ <div align="center">
22
+ <picture>
23
+ <source srcset="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo_darkmode.png?raw=true" media="(prefers-color-scheme: dark)">
24
+ <img src="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo.png?raw=true" width="60%" alt="Xiaomi-MiMo" />
25
+ </picture>
26
+ </div>
27
+
28
+ <br/>
29
+
30
+ <div align="center" style="line-height: 1;">
31
+ |
32
+ <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" target="_blank">🤗 HuggingFace</a>
33
+ &nbsp;|
34
+ <a href="https://mimo.xiaomi.com/mimo-v2-6" target="_blank">📰 Blog </a>
35
+ &nbsp;|
36
+ <a href="https://platform.xiaomimimo.com" target="_blank">🎨 Xiaomi MiMo API Platform </a>
37
+ &nbsp;|
38
+ <a href="https://aistudio.xiaomimimo.com" target="_blank">🗨️ Xiaomi MiMo Studio </a>
39
+ &nbsp;|
40
+ <a href="https://mimo.xiaomimimo.com/desktop/" target="_blank">💻 Xiaomi MiMo Desktop </a>
41
+ &nbsp;|
42
+ </div>
43
+
44
+ <br/>
45
+
46
+ <div align="center" style="line-height: 1.2;">
47
+ <strong>Community</strong><br/>
48
+ <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro/blob/main/assets/wechat.jpg" target="_blank">WeChat Group</a>
49
+ &nbsp;|&nbsp;
50
+ <a href="https://discord.gg/kKC2kNnQEX" target="_blank">Discord</a>
51
+ &nbsp;|&nbsp;
52
+ <a href="https://t.me/+3T-I0pekOVIyNDBl" target="_blank">Telegram</a>
53
+ &nbsp;|&nbsp;
54
+ <a href="https://www.reddit.com/r/XiaomiMiMo_Official/" target="_blank">Reddit</a>
55
+ </div>
56
+
57
+ <br/>
58
+
59
+ # MiMo-V2.6-Pro-RL
60
+
61
+ **Scaling Reinforcement Learning Toward Self-Improvement**
62
+
63
+ <p align="center">
64
+ <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf"><b>Technical Report</b></a>
65
+ </p>
66
+
67
+ ## 1. Introduction
68
+
69
+ MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series. The series is built to **scale reinforcement learning toward self-improvement** — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include:
70
+
71
+ - **Native Omnimodal + Long Horizon**: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs.
72
+ - **You Only RL Once**: One mixed RL run across coding, general agents, visual, and cybersecurity — not separate per-domain runs. Tasks and multiple harnesses are mixed in the same batch so capabilities reinforce each other and strategies transfer to harnesses never seen in training.
73
+ - **Scaling RL Compute**: Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches — 1,568 prompts × 16 rollouts per step, billions of tokens per update.
74
+ - **Groupwise Agentic Grading (Self-Improvement Loop)**: Binary pass/fail cannot rank passing solutions, so the reward signal itself is scaled. An agentic grader compares rollouts *within each group*: **Groupwise Reward Synthesis (GRS)** builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes; **Groupwise Advantage Redistribution (GAR)** ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.
75
+ - **Aligned RL**: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
76
+ - **Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2)**: After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix), reusing histories from teacher trajectories and SFT demonstrations so decision points train without regenerating preceding turns — extending capabilities to hard-to-verify tasks.
77
+
78
+ ## Model Summary
79
+
80
+ - **Architecture**: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters
81
+ - **Context Length**: 1M tokens
82
+ - **Modalities**: Text, Image, Video, Audio
83
+ - **Vision Encoder**: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
84
+ - **Audio Encoder**: 308M AudioTokenizer + 127M audio patch encoder
85
+ - **Multi-Token Prediction (MTP)**: 5-layer speculative decoder
86
+
87
+ ![Figure 1: MiMo-V2.6 architecture — omni encoders, hybrid SWA backbone, and MTP blocks](assets/architecture.png)
88
+
89
+ *Figure 1. MiMo-V2.6 architecture.*
90
+
91
+ ## 2. Downloads
92
+
93
+ | Model | Download |
94
+ | --- | --- |
95
+ | **MiMo-V2.6-Pro-RL** | [🤗 HuggingFace](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) · 🤖 ModelScope *(at release)* |
96
+ | **MiMo-V2.6-Flash-RL** | [🤗 HuggingFace](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) · 🤖 ModelScope *(at release)* |
97
+
98
+ ## 3. Evaluation Results
99
+
100
+ | Benchmark | MiMo-V2.6 Pro | MiMo-V2.6 Flash | MiMo-V2.5 Pro | Claude Opus 5 | GPT-5.6 Sol | Claude Fable 5 |
101
+ | --- | --- | --- | --- | --- | --- | --- |
102
+ | **Code Agent** | | | | | | |
103
+ | DeepSWE v1.1 | 71.9 | 67.9 | 19.0 | 74.0 | 73.0 | 70.0 |
104
+ | ProgramBench | 26.5 | 26.0 | 12.5 | 37.0 | 25.0 | 33.0 |
105
+ | MiMo Code Bench | 63.2 | 61.2 | 40.4 | 68.6 | 59.3 | - |
106
+ | **General Agent** | | | | | | |
107
+ | AutomationBench v1.0.6 | 53.1 | 52.3 | 16.0 | 50.3 | 45.8 | 46.2 |
108
+ | Toolathlon-Verified | 76.9 | 73.6 | 49.1 | 80.6 | 74.9 | 77.9 |
109
+ | GDPval-AA 2.1 | 1673 | - | 1107 | 1708 | 1588 | 1595 |
110
+ | Agents’ Last Exam | 31.6 | 27.6 | 13.2 | 31.6 | 30.8 | 25.7 |
111
+ | Terminal Bench 4.0 | 34.9 | 28.8 | 1.5 | 49.0 | 39.9 | 42.4 |
112
+ | Terminal Bench 2.1 | 89.9 | 87.6 | 65.2 | 89.1 | 88.8 | 84.3 |
113
+ | OSWorld-Verified | 82.0 | 80.8 | - | 83.4 | 83.0 | 86.0 |
114
+ | JobBench | 62.0 | 61.2 | 25.0 | 65.7 | 45.4 | 57.4 |
115
+ | **Cybersecurity** | | | | | | |
116
+ | CyberGym | 94.0 | 95.1 | 40.0 | - | - | - |
117
+ | MiMo Cyber Bench | 80.2 | 77.2 | 0.0 | - | - | - |
118
+ | ExploitGym | 17.8 | 6.0 | 0.2 | 22.1 | 30.3 | 28.4 |
119
+ | ExploitBench | 47.9 | 25.3 | 16.6 | 70.0 | 78.5 | 78.0 |
120
+ | SEC Bench Pro | 66.3 | 47.5 | 17.7 | - | 79.1 | - |
121
+ | **Visual Agent** | | | | | | |
122
+ | MiMo VisualCoding | 72.3 | 71.5 | - | 70.0 | 73.4 | 69.1 |
123
+
124
+ ## 4. Model Architecture
125
+
126
+ ### LLM Backbone
127
+
128
+ | Component | MiMo-V2.6-Pro-RL |
129
+ | --- | --- |
130
+ | Layers (Total / SWA / GA) | 70 / 60 / 10 |
131
+ | Hidden Size | 6144 |
132
+ | SWA Heads (Q/KV) | 128 / 8 |
133
+ | GA Heads (Q/KV) | 128 / 8 |
134
+ | Head Dimensions (QK / V) | 192 / 128 |
135
+ | Sliding Window Size | 128 |
136
+ | Routed Experts (Total / Activated) | 384 / 8 |
137
+ | Max Context Length | 1M |
138
+ | MTP / Speculative Decoder | 5 SWA layers, window 1024 |
139
+
140
+ The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.
141
+
142
+ ### Vision Encoder (MiMo ViT)
143
+
144
+ | Configuration | Value |
145
+ | --- | --- |
146
+ | Layers (Total / SWA / GA) | 28 / 24 / 4 |
147
+ | Hidden Size | 1280 |
148
+ | Attention Heads (Q / KV) | 32 / 8 |
149
+ | Head Dimension | 64 |
150
+ | Patch Size (T × H × W) | 2 × 16 × 16 |
151
+ | Sliding Window (Left / Right) | 64 / 64 |
152
+ | Spatial Merge Size | 2 × 2 |
153
+ | Parameters | 681M |
154
+
155
+ ### Audio Encoders
156
+
157
+ AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).
158
+
159
+ ### Speculative Decoder
160
+
161
+ 5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.
162
+
163
+ ## 5. Deployment
164
+
165
+ For best performance, follow the [SGLang MiMo cookbook](https://docs.sglang.io/cookbook/autoregressive/Xiaomi/MiMo-V2.5). Docker image: `lmsysorg/sglang:latest`.
166
+
167
+ ### SGLang
168
+
169
+ ```bash
170
+ sglang serve \
171
+ --trust-remote-code \
172
+ --model-path XiaomiMiMo/MiMo-V2.6-Pro-RL \
173
+ --tp 16 \
174
+ --dp 2 \
175
+ --enable-dp-attention \
176
+ --mm-enable-dp-encoder \
177
+ --ep 16 \
178
+ --moe-a2a-backend deepep \
179
+ --moe-dense-tp-size 1 \
180
+ --mem-fraction-static 0.7 \
181
+ --max-running-requests 128 \
182
+ --chunked-prefill-size 32768 \
183
+ --page-size 64 \
184
+ --swa-full-tokens-ratio 0.3 \
185
+ --speculative-algorithm EAGLE \
186
+ --speculative-num-steps 3 \
187
+ --speculative-eagle-topk 1 \
188
+ --speculative-num-draft-tokens 4 \
189
+ --enable-multi-layer-eagle \
190
+ --reasoning-parser mimo \
191
+ --tool-call-parser mimo \
192
+ --host 0.0.0.0 \
193
+ --port 30000 \
194
+ --nnodes 2 \
195
+ --node-rank <node-rank> \
196
+ --dist-init-addr <node0-ip>:20000
197
+ ```
198
+
199
+ ### vLLM
200
+
201
+ Follow the [vLLM MiMo-V2.5 recipe](https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.5). Pre-built image: `docker pull vllm/vllm-openai:mimov25-cu129`.
202
+
203
+ ```bash
204
+ vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL \
205
+ --tensor-parallel-size 8 \
206
+ --trust-remote-code \
207
+ --gpu-memory-utilization 0.95 \
208
+ --max-model-len auto \
209
+ --reasoning-parser mimo \
210
+ --tool-call-parser mimo \
211
+ --enable-auto-tool-choice \
212
+ --generation-config vllm
213
+ ```
214
+
215
+ Recommended sampling: `temperature=1.0`, `top_p=0.95`.
216
+
217
+ Also available in AI Studio, MiMo Code, Xiaomi MiMo Desktop, Xiaomi MiMo Open Platform API, and OpenRouter.
218
+
219
+ ## Citation
220
+
221
+ ```bibtex
222
+ @misc{mimo2026v26pro,
223
+ title={MiMo-V2.6-Pro-RL},
224
+ author={{Xiaomi MiMo Team}},
225
+ year={2026},
226
+ howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
227
+ }
228
+ ```
229
+
230
+ ## Contact
231
+
232
+ For questions or feedback, reach us at [[email protected]](mailto:[email protected]) or join our community:
233
+
234
+ - [WeChat Group](https://work.weixin.qq.com/apph5/external_room/join/group_mng?plg_id=c417f99bd9014b5dd894daa8bfe19790&)
235
+ - [Discord](https://discord.gg/WX2R2uNp)
236
+ - [Telegram](https://t.me/+3T-I0pekOVIyNDBl)
237
+ - [Reddit](https://www.reddit.com/r/XiaomiMiMo_Official/)