Instructions to use olaverse/PurpleMIST-Flash-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use olaverse/PurpleMIST-Flash-1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="olaverse/PurpleMIST-Flash-1.0")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("olaverse/PurpleMIST-Flash-1.0") model = AutoModel.from_pretrained("olaverse/PurpleMIST-Flash-1.0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
PurpleMIST-Flash-1.0
PurpleMIST-Flash-1.0 Highlights
PurpleMIST-Flash-1.0 is a System One decision model. Give it a piece of unstructured state (a message, ticket, record or transcript) and any number of typed questions about it. It returns a calibrated probability distribution for every answer, in one forward pass. It is built for African languages as well as English.
- As accurate as DeepSeek-V4-Pro, a model about 200× its size, with far better probabilities. On African Typed Decisions (8 languages) it scores 0.668 accuracy, against 0.669 for DeepSeek-V4-Pro (1.6T parameters). Its distributions are about 5× closer to the reference (KL 0.214 vs 1.125), and its Brier score is 0.115 vs 0.199.
- Holds up outside English. Accuracy falls by 0.053 from the English original to the eight African languages. OpenDecider-small falls by 0.171.
- Strong zero-shot in English. It scores 0.720 accuracy on
LocalLLaMA/typed-decisions (KL 0.189, Brier 0.098)
without training on that benchmark's
trainsplit. That puts it 4th in accuracy and 2nd in KL and Brier against the zero-shot models listed there. - Calibrated. Each question type has its own temperature, fitted on held-out states from every training source. ECE is 0.049 on the African languages and 0.056 in English.
- A smaller sibling: PurpleMIST-Mini-1.0 (1.9B) has the same interface, for English workloads.
- One pass per state. All questions about a state share one sequence. Each option is scored from its own text, so any number of options works and options do not need to be fixed in advance.
See it in action
Every probability in these videos comes from live calls to PurpleMIST-Flash-1.0.
English, and a game. It routes a support ticket, catches a scam text and holds a mismatched invoice. Then it plays a lane-runner game in which every move is a live decision: 60 moves, 0 crashes, 24 coins.
African languages. It works through a fake bank SMS in Yorùbá, a lost mobile-money transfer in Hausa, an M-Pesa scam in Kiswahili, a market-levy text in Nigerian Pidgin and a return request in Amharic. The cases come from African Typed Decisions.
Model Overview
- Type: decision model. It scores options and does not generate text.
- Backbone: the text model of Qwen/Qwen3.5-9B-Base, with the vision tower and language-model head removed. It has 32 layers (24 Gated DeltaNet linear-attention and 8 full attention), hidden size 4096 and about 7.9B parameters.
- Head: a pointer head. Each option's logit is a scaled dot product between projections (4096 → 1024) of the answer position and the option's own line. The head has 8.4M parameters.
- Fine-tuning: LoRA rank 32 on every linear layer, merged into the weights. The unmerged adapter is in
adapter/. - Question types:
| Type | You give | You get back |
|---|---|---|
choice |
instructions and named options (criteria: key → description) |
{"choice": key, "probabilities": {key: p}} |
noul |
a yes/no statement | {"noul": p_yes} |
score |
instructions and ordered level descriptions (lowest first) | {"score": expected level, "probabilities": {"0": p, ...}} |
- Weights: bf16 safetensors, 15.9 GB. A 24 GB GPU is enough.
- Context: up to 65,536 tokens per pass. On the Jev Decision Index it answered all 140,620 requests, including
ones over 65,000 tokens, without truncation. When the questions do not fit alongside the state, they are split
across several passes automatically.
Decideruses 2,048 tokens by default, to keep memory low; passmax_len=65536for long inputs.
Quickstart
pip install "transformers>=5.19" torch huggingface_hub
pip install flash-linear-attention causal-conv1d # optional: fast kernels for the linear-attention layers
import os, sys
from huggingface_hub import hf_hub_download
repo = "olaverse/PurpleMIST-Flash-1.0"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "purplemist.py")))
from purplemist import Decider
d = Decider.from_pretrained(repo, device="cuda")
state = {"channel": "whatsapp",
"message": "Ẹ jọ̀ọ́, wọ́n ti yọ owó lẹ́ẹ̀mejì lórí káàdì mi fún ọjà kan náà. Mo fẹ́ kí ẹ dá owó mi padà."}
questions = {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments, charges and refunds", "delivery": "orders in transit",
"technical": "app or account problems"}},
"refund": {"type": "noul", "instructions": "The customer is asking for money back."},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["can wait", "normal queue", "today", "immediately"]},
}
print(d.decide(state, questions))
# {"team": {"choice": ..., "probabilities": {...}}, "refund": {"noul": ...}, "urgency": {"score": ..., "probabilities": {...}}}
state can be a string or any JSON-serialisable object. Instructions and option descriptions can be in English or
in the state's language.
Serving over HTTP
serve.py runs the model as a small local server with the System One request shape (POST /v1/systemone), so
clients built for that API can call it. It needs only the packages from the quickstart.
hf download olaverse/PurpleMIST-Flash-1.0 serve.py purplemist.py --local-dir purplemist
python purplemist/serve.py --model olaverse/PurpleMIST-Flash-1.0 --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": {"message": "I was charged twice, please refund me"},
"questions": {
"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "payments", "tech": "app problems"}},
"refund": {"type": "noul", "instructions": "The customer asks for money back."}}}'
# {"model": "...", "answers": {"team": {"choice": ..., "probabilities": {...}}, "refund": {"noul": ...}}, "latency_ms": ...}
The server also offers GET /health and GET /v1/models. It handles one request at a time. The answers match Ollama's decision API. Ollama 0.35+
serves decision models at the same POST /v1/systemone, and serve.py returns the same answer fields (choice,
noul, score, probabilities, confidence, legend), so code written for Ollama's decision endpoint reads its
answers the same way. PurpleMIST is not in the Ollama library yet, so for now run it with serve.py.
Evaluation
African Typed Decisions
All six models were run through one harness. Each request contains a whole case: the state and all of its questions. Every answer is turned into a probability vector over the same options, and every model is scored by the same code. LLMs were prompted for JSON probabilities at temperature 0 with reasoning off. Any answer that is missing or cannot be parsed counts as a uniform guess.
| Model | Size | English original | African, translated | English → African drop | African, native | KL ↓ | Brier ↓ | ECE ↓ |
|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Pro (prompted) | 1.6T | 0.720 | 0.669 | 0.051 | 0.955 † | 1.125 | 0.199 | 0.121 |
| PurpleMIST-Flash-1.0 | 7.9B | 0.720 | 0.668 | 0.053 | 0.932 ‡ | 0.214 | 0.115 | 0.049 |
| Qwen3.8-Flash (prompted) | undisclosed | 0.712 | 0.643 | 0.070 | 0.916 † | 0.862 | 0.216 | 0.132 |
| OpenDecider-small | 4B | 0.648 | 0.477 | 0.171 | 0.635 | 0.365 | 0.199 | 0.027 |
| PurpleMIST-Mini-1.0 | 1.9B | 0.601 | 0.416 | 0.185 | 0.769 ‡ | 0.399 | 0.221 | 0.050 |
| laya-multilingual | 0.32B | 0.348 | 0.322 | 0.026 | 0.280 | 5.053 | 0.545 | 0.360 |
KL, Brier and ECE are on the translated set: 3,010 cases, 15,050 decisions. The translated cases are the 400 LocalLLaMA/typed-decisions test cases in eight languages, so they compare directly with the English column.
† DeepSeek-V4-Pro wrote the Nigerian-language native cases, and Qwen3.8-Flash wrote or checked all of them. Both are scored partly against their own answers on that column. ‡ PurpleMIST was trained on native cases generated by the same pipeline. They were separate cases from those in the benchmark.
By language (accuracy, translated cases):
| Model | Amharic | Hausa | Igbo | Pidgin | Somali | Swahili | Yorùbá | isiZulu |
|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Pro | 0.650 | 0.675 | 0.673 | 0.688 | 0.678 | 0.682 | 0.651 | 0.657 |
| PurpleMIST-Flash-1.0 | 0.649 | 0.649 | 0.654 | 0.710 | 0.683 | 0.690 | 0.644 | 0.671 |
| Qwen3.8-Flash | 0.622 | 0.625 | 0.627 | 0.681 | 0.641 | 0.661 | 0.629 | 0.664 |
| OpenDecider-small | 0.491 | 0.421 | 0.473 | 0.619 | 0.443 | 0.510 | 0.439 | 0.456 |
| PurpleMIST-Mini-1.0 | 0.330 | 0.414 | 0.422 | 0.589 | 0.371 | 0.418 | 0.420 | 0.406 |
| laya-multilingual | 0.334 | 0.326 | 0.317 | 0.351 | 0.299 | 0.311 | 0.352 | 0.299 |
PurpleMIST leads on Pidgin, Somali, Swahili and isiZulu, and is level with DeepSeek-V4-Pro on Amharic and Yorùbá (within 0.007). It trails on Hausa (−0.026) and Igbo (−0.019).
LocalLLaMA/typed-decisions (English, zero-shot)
These are the published zero-shot rows from the benchmark card, with our measured rows added. The test split has 400 cases and 2,000 decisions.
| Model | Size | Accuracy ↑ | KL ↓ | Brier ↓ | ECE ↓ | Source |
|---|---|---|---|---|---|---|
| meraGPT Decider 1 | undisclosed | 0.768 | 0.096 | 0.052 | 0.180 | benchmark card |
| Liquid AI d1 | undisclosed | 0.742 | 0.475 | 0.155 | 0.124 | benchmark card |
| TypeSafe Jev 1.13.0 | undisclosed | 0.727 | 1.442 | 0.148 | 0.144 | benchmark card |
| PurpleMIST-Flash-1.0 | 7.9B | 0.720 | 0.189 | 0.098 | 0.056 | measured by us |
| DeepSeek-V4-Pro (prompted) | 1.6T | 0.720 | 1.359 | 0.185 | 0.108 | measured by us |
| Featherless Simple Jev | 35B (3B active) | 0.716 | 0.488 | 0.176 | – | benchmark card |
| Qwen3.8-Flash (prompted) | undisclosed | 0.712 | 0.781 | 0.172 | 0.078 | measured by us |
| prima-ratio + 12B | 12B | 0.702 | 0.564 | 0.234 | 0.146 | benchmark card (self-reported) |
| OpenDecider-small | 4B | 0.671 | 0.211 | 0.117 | – | benchmark card (self-reported) |
| OpenDecider-small | 4B | 0.648 | 0.224 | 0.126 | 0.059 | measured by us |
| PurpleMIST-Mini-1.0 | 1.9B | 0.601 | 0.280 | 0.156 | 0.058 | measured by us |
| Bongard-mini | 7.5B | 0.594 | 0.256 | 0.132 | 0.067 | benchmark card (self-reported) |
| Jeff-Gemma4-E2B | 4.6B | 0.561 | 0.403 | 0.219 | 0.188 | benchmark card |
| Jeff-Qwen3.5-2B | 2.2B | 0.511 | 0.460 | 0.237 | 0.203 | benchmark card |
| Jeff-Qwen3.5-0.8B | 0.85B | 0.483 | 0.679 | 0.313 | 0.251 | benchmark card |
| Prior (ignores the input) | – | 0.470 | 0.347 | 0.189 | 0.088 | benchmark card |
Sizes are total parameters; "undisclosed" means the provider does not publish one. Submitters compute ECE differently, so compare KL and Brier rather than ECE.
By question type. PurpleMIST's strongest type is yes/no (0.822), within 0.02 of the two leaders. Its gap to
meraGPT Decider 1 is mostly on choice and score. The other models' figures are from the benchmark card. PurpleMIST's
come from the evaluation script described below (overall 0.717).
Jev Decision Index (public suite)
The Jev Decision Index tests decision models on 37
public benchmarks in five areas: knowledge and reasoning, language understanding, retrieval and classification, tools
and automation, and arts and human taste. Each score is chance-corrected, so 0 is random guessing and 100 is perfect.
PurpleMIST-Flash-1.0 was run with the maintainers' own kit (apolinario/decision-index,
edition 0.3) on one RTX PRO 6000. It was served by serve.py --no-truncate --max-len 65536, and all 140,620 requests
were answered, with none truncated or declined.
Its public index is 46.7. Compared with the 13 models of similar size (7–10.5B parameters) on the board:
- It is ahead of 8 of the 13.
- It is 4.8 points above their median (41.9).
- It is above the median in four of the five areas: language understanding (+6.4), knowledge and reasoning (+4.3), tools (+3.8) and retrieval (+1.4).
The two leaders, Bespoke Nimble 9B v3 and Cloudflare clef-flash, are still about 10 points ahead. The board's Full score adds the maintainers' private tests (80% of it) after review, so it is still pending.
| Model | Size | Public index ↑ | Knowledge | Language | Retrieval | Tools | Arts & taste |
|---|---|---|---|---|---|---|---|
| Bespoke Nimble 9B v3 | 9.7B | 57.2 | 41.0 | 60.7 | 55.5 | 84.2 | 44.0 |
| Cloudflare clef-flash | 9.7B | 56.1 | 53.0 | 47.2 | 52.2 | 81.9 | 48.4 |
| JevK5 9B v0.3.3 | 9.7B | 47.6 | 31.6 | 54.7 | 50.0 | 64.1 | 35.7 |
| JPT-9B | 9.7B | 47.0 | 30.9 | 57.7 | 44.6 | 67.0 | 28.8 |
| vLLM-SR Decision 2.0 Lux 9B | 9.7B | 46.8 | 35.5 | 46.9 | 55.6 | 60.4 | 33.1 |
| PurpleMIST-Flash-1.0 | 7.9B | 46.7 | 35.0 | 53.1 | 47.5 | 65.1 | 25.5 |
| JEV-9B | 9.7B | 43.3 | 30.7 | 44.1 | 46.5 | 62.1 | 32.8 |
| Sieve 9B | 9.7B | 41.9 | 29.0 | 45.6 | 46.1 | 61.3 | 21.9 |
| Kev 9B v2 | 9.7B | 41.5 | 30.8 | 46.7 | 46.8 | 56.9 | 17.5 |
| Winnow-E4B | 8.0B | 40.9 | 22.8 | 46.8 | 43.8 | 62.5 | 26.9 |
| Bespoke Nimble 9B v2 | 9.7B | 39.9 | 30.5 | 44.3 | 34.9 | 59.2 | 27.6 |
| Jevstral 8B | 8.0B | 34.0 | 24.7 | 33.6 | 39.2 | 52.9 | 14.1 |
| Gevva E4B | 8.0B | 32.4 | 17.9 | 27.6 | 40.0 | 56.9 | 22.6 |
| CLM-v0.1-8B | 8.2B | 7.4 | 2.4 | 5.7 | 11.3 | 14.0 | 5.2 |
Bold marks the best score in each column. Other models' scores are as published on the board. PurpleMIST's are from its own complete run, not yet reviewed.
- Strongest areas against models its size: language understanding (53.1, 4th of 14), and tools and knowledge (4th of 14 each).
- Weakest: arts and human taste (25.5, 9th of 14). Its training mix had little preference or taste data.
- Speed: median 63 ms and mean 94 ms per request over the full run (self-measured; the board's latencies are measured by its maintainers).
Other held-out sets
These numbers come from the training repository's evaluation script, which reads the held-out splits of the training data. On African Typed Decisions that script reads the questions from the training-data format and gives slightly different figures (accuracy 0.674 rather than 0.668), so the tables above use the benchmark harness.
| Set | Decisions | Accuracy | Brier | KL | ECE |
|---|---|---|---|---|---|
| Open-Jev test | 6,972 | 0.787 | 0.312 | 0.605 | 0.090 |
| Open-Jev out-of-distribution | 8,325 | 0.767 | 0.360 | 0.703 | 0.122 |
| tasksource test | 3,800 | 0.810 | 0.227 | 0.438 | 0.024 |
| Najd | 4,368 | 0.813 | 0.257 | 0.640 | 0.015 |
| HelpSteer2 (score) | 5,000 | 0.604 | 0.522 | 0.939 | 0.014 |
Open-Jev's yes/no questions remain the least calibrated (ECE about 0.20), because temperatures are fitted across all sources rather than to Open-Jev alone.
Training
Data: 300,000 states sampled from a 1.71M-decision training set, with every African state included. The sources are:
- public System One and preference data: tasksource, Open-Jev
train, HelpSteer2, jev-decisions-clean50k, two synthetic sets and Najd development; - 86,134 African decisions. 36,398 are cases written natively in the eight languages and 49,736 are translated from the public training states.
Every source allows commercial use. LocalLLaMA/typed-decisions
trainwas not used, which keeps the English result zero-shot. Corpora distilled from the Jev API were excluded.- public System One and preference data: tasksource, Open-Jev
Objective: soft-target cross-entropy against each question's reference distribution. Options are shuffled during training.
Run: 1,501 steps at learning rate 1.5e-4, on 4× H100.
Calibration: after training, one temperature per question type was fitted on 2,751 states (5,345 questions) that training never used. The states are spread across eight sources, each weighted equally. The temperatures are choice 0.938, noul 1.035 and score 0.907, and they are stored in
purplemist_config.json.
Limitations
- Gold comes from LLM teachers. Reference answers were generated by language models, so a score measures agreement with those teachers, not ground truth.
- African data is machine-written. It was machine-written and machine-checked, not reviewed by native speakers.
- Text only. Images in the state are not read.
Citation
@misc{purplemist-flash-1.0,
title = {PurpleMIST-Flash-1.0: a calibrated System One decision model for English and African languages},
author = {Olaverse},
year = {2026},
url = {https://huggingface.co/olaverse/PurpleMIST-Flash-1.0}
}
- Downloads last month
- 22
Model tree for olaverse/PurpleMIST-Flash-1.0
Base model
Qwen/Qwen3.5-9B-BaseDatasets used to train olaverse/PurpleMIST-Flash-1.0
olaverse/african-typed-decisions
Collection including olaverse/PurpleMIST-Flash-1.0
Evaluation results
- LocalLLaMA/typed-decisions leaderboard
- Accuracy View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions, all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options. One forward pass per state, no generated tokens. bf16, calibrated temperatures from purplemist_config.json.0.72 * - Kl From Gold View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions, all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options. One forward pass per state, no generated tokens. bf16, calibrated temperatures from purplemist_config.json.0.19 * - Brier View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions, all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options. One forward pass per state, no generated tokens. bf16, calibrated temperatures from purplemist_config.json.0.1 * - olaverse/african-typed-decisions
- Accuracy Translated View evaluation resultssource
whole case per request; calibrated temperatures from purplemist_config.json0.67 *




