Text Generation
Safetensors
GGUF
English
Chinese
vllm
mimo_v2
nvfp4
modelopt
mixture-of-experts
lossless-weight-transcode
conversational
custom_code
8-bit precision
Instructions to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16 # Run inference directly in the terminal: llama cli -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16 # Run inference directly in the terminal: llama cli -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16 # Run inference directly in the terminal: ./llama-cli -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Use Docker
docker model run hf.co/ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
- LM Studio
- Jan
- vLLM
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ProCreations/MiMo-V2.6-Flash-RL-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ProCreations/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
- Ollama
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with Ollama:
ollama run hf.co/ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
- Unsloth Desktop
- Pi
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with Docker Model Runner:
docker model run hf.co/ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
- Lemonade
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Run and chat with the model
lemonade run user.MiMo-V2.6-Flash-RL-NVFP4-BF16
List all available models
lemonade list
- Hermes Agent
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ProCreations/MiMo-V2.6-Flash-RL-NVFP4 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ProCreations/MiMo-V2.6-Flash-RL-NVFP4:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download auxiliary-verification.json from ProCreations/MiMo-V2.6-Flash-RL-NVFP4: direct link, hf CLI and curl.
- Browser
- Download file 4.3 kB
-
https://huggingface.co/ProCreations/MiMo-V2.6-Flash-RL-NVFP4/resolve/main/auxiliary-verification.json
- Command line
-
hf download hf://ProCreations/MiMo-V2.6-Flash-RL-NVFP4/auxiliary-verification.json
-
curl -L -o auxiliary-verification.json https://huggingface.co/ProCreations/MiMo-V2.6-Flash-RL-NVFP4/resolve/main/auxiliary-verification.json
4.3 kB
| { | |
| "status": "passed", | |
| "completed_at": 1790041282.8461897, | |
| "files": [ | |
| { | |
| "path": "MiMo_V2_6_technical_report.pdf", | |
| "sha256": "7fe42601dc952cd2b74996e5a24f8e85eab6fcbf471f73ba95aafd559d4ef39b", | |
| "bytes": 3046685, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "assets/architecture.png", | |
| "sha256": "d288768e1771fec19b39ed7dbac4adcbbd2e490384d4ad3c58d259b0c6c6bdcc", | |
| "bytes": 405440, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "audio_tokenizer/chat_template.jinja", | |
| "sha256": "cf1a0a0e5cbc6a6a1b609f19f6db5483fddf978e6c8c8453ca2549bed02b7425", | |
| "bytes": 5588, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "audio_tokenizer/config.json", | |
| "sha256": "e0702adae37947e0c980c38bae58ffa0d48bd492d523afe814aa4d73f008c7d1", | |
| "bytes": 1215, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "audio_tokenizer/generation_config.json", | |
| "sha256": "5e14358a9bd50424310fe2c5afabf6a131c3a9b452c9322f7d9ee6f9ff6b8ef9", | |
| "bytes": 149, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "audio_tokenizer/model.safetensors", | |
| "sha256": "077033345d80eef3a315e8d394e0589667e80e4cdaba9bc5a7488410c6657265", | |
| "bytes": 1872618384, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "audio_tokenizer/tokenizer_config.json", | |
| "sha256": "ac41378c3257a15e31d5afa7a8e3e6ca9c2ec6702adf7a93064f8dae384939df", | |
| "bytes": 6059, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "chat_template.jinja", | |
| "sha256": "853650bee57bf95020373e4c928bd5a4b41b9915adf964a77711d2b49a291887", | |
| "bytes": 3868, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "configuration_mimo_v2.py", | |
| "sha256": "773062ac9850b908eb54751b3e4dbe00e653c0e80595599409e97d0c1af2ce3e", | |
| "bytes": 10049, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "dflash/config.json", | |
| "sha256": "3a65305bd8d9a8bef6ca46866567be1c56f0a867fc802f4749f1333a976590ef", | |
| "bytes": 1242, | |
| "operation": "syntax-only JSON repair; identical parsed configuration" | |
| }, | |
| { | |
| "path": "dflash/dflash.py", | |
| "sha256": "da5ab1738b954800950405131f1d1d97c3345f37e32676d511d3a25dfddd9d75", | |
| "bytes": 14255, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "dflash/dflash_draft_model.safetensors", | |
| "sha256": "94d9c02c17e0b469f88699e98424825ebe13a17f07eb3b86727b5e0c8143d287", | |
| "bytes": 2936121080, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "dflash/mask_embedding.pt", | |
| "sha256": "b35b379fe0497ffdc0d6407f502622237d9346e0d03f457ec574be60f6c2cee0", | |
| "bytes": 9882, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "dflash/model.safetensors.index.json", | |
| "sha256": "a951ae2b40c8ea394482e74dd1a35488e0fc3eb60b3359857daa60e7dfd849f0", | |
| "bytes": 4689, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "generation_config.json", | |
| "sha256": "eecf00d9701921271c904a916c425ba6db61bcd190902e200d0a322bfd5572dd", | |
| "bytes": 195, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "merges.txt", | |
| "sha256": "599bab54075088774b1733fde865d5bd747cbcc7a547c5bc12610e874e26f5e3", | |
| "bytes": 1671839, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "modeling_mimo_v2.py", | |
| "sha256": "a8c3cb3aae473bcc15f023010547c919f15eba6546e6ed7efb61a8937b12f3ad", | |
| "bytes": 85516, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "preprocessor_config.json", | |
| "sha256": "b269e51bdc1c53ef7c82522984378513f462869cf83c4c75ec3f98bd92aa7972", | |
| "bytes": 350, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "tokenizer.json", | |
| "sha256": "ff15eb925890d6b71b5160de4b846fbd13178438ab463b38ecc953e8cd1dcb3e", | |
| "bytes": 11423823, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "tokenizer_config.json", | |
| "sha256": "413a7845f52943ccf4de0e5c838414507d16c44dbf573da9e20bc8902b384d06", | |
| "bytes": 12150, | |
| "operation": "identical bytes" | |
| }, | |
| { | |
| "path": "vocab.json", | |
| "sha256": "ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910", | |
| "bytes": 2776833, | |
| "operation": "identical bytes" | |
| } | |
| ] | |
| } | |