Instructions to use yunicro/MiDM-4B-q35-e1-bx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use yunicro/MiDM-4B-q35-e1-bx with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
MiDM-4B-q35-e1-bx
v0.2.1 benchmark and architecture analysis
Documentation update; this repository's existing weights and inference code are unchanged. Full benchmark atlas · GitHub release · Version DOI.
The atlas separates matched measurements from historical/published/API comparisons. JevBench public accuracy is not its official composite. Missing MiDM DeepSWE and local Terminal-Bench measurements remain explicit; their candidate pools were not found. Unknown proprietary parameter counts are not guessed. SQL/Python regressions remain documented.
MiDM (Minimal Decision Model): a typed-decision model. The input is one state plus typed questions
(choice, noul yes/no, score), and the output is a probability distribution over each question's options.
It is a LoRA adapter plus a linear pointer head on Qwen/Qwen3.5-4B-Base. The model reads the state, the question and
every option line in one pass. It scores each option from the final hidden state at the end of that option's
line.
Use
from midm import MiDM # midm.py ships in this repo
m = MiDM.from_pretrained("yunicro/MiDM-4B-q35-e1-bx") # 4-bit NF4 base by default (as in training)
m.predict(state={"order_total": 1240, "po_total": 1420},
questions={"action": {"type": "choice", "instructions": "What should happen to this invoice?",
"criteria": {"approve": "Approve and pay.", "hold": "Hold for review.",
"reject": "Reject as fraud."}}})
Requests follow the "System One" typed-decision shape. requirements.txt lists the dependencies.
A GPU with about 9 GB free is needed for the 4B model in 4-bit.
Evaluation (each suite evaluated once; accuracy = argmax over options)
| suite | MiDM-4B-q35-e1-bx | MiDM-4B-q35-e1 | change (pp) |
|---|---|---|---|
| typed-decisions test (2000) | 0.788 | 0.805 | -1.7 |
| Kev decision-v7 dev (1468) | 0.856 | 0.861 | -0.5 |
| Kev transfer-v4 dev (764) | 0.791 | 0.755 | +3.5 |
| Kev transfer-v4 test (764) | 0.798 | 0.764 | +3.4 |
| Kev transfer-v9 dev (1264) | 0.705 | 0.675 | +3.0 |
| JevBench public (231) | 0.710 | 0.710 | +0.0 |
Kev transfer suites contain sources never seen in training (e.g. MMLU, SciQ, emotion, PAWS, QNLI). JevBench "public" is the 231 public items only; sealed-set performance is typically much lower on that board.
Comparison with CLM-8B and TypeSafe Jev
| system | typed-decisions test | Kev transfer-v4 dev | Kev transfer-v4 test | JevBench public |
|---|---|---|---|---|
| TypeSafe Jev (commercial, zero-shot) | 0.727 | 0.857 | – | 0.866 |
| CLM-8B (frozen Qwen3-8B + heads; typed-dec. from CLM PR #2, JevBench board) | 0.685 | – | – | 0.407 |
| CLM recipe re-trained by us (frozen Qwen3-8B, z-score, 3 seeds) | 0.766 | – | – | – |
| CLM-style frozen Qwen3-4B + heads (ours, same data as MiDM) | 0.759 | 0.366 | 0.411 | 0.411 |
| Kev-9B (reported) | – | 0.822 | 0.852 | – |
| MiDM-4B-q35-e1-bx (this model) | 0.788 | 0.791 | 0.798 | 0.710 |
Rows not marked "ours"/"by us" are numbers reported by their authors or boards. They are not paired with ours and use different protocols. Jev is zero-shot, while this model was trained on typed-decisions train, so the typed-decisions column favours this model. On unseen sources (Kev transfer) and JevBench public, Jev and Kev-9B remain ahead. CLM-style frozen-encoder heads are competitive in-distribution but fall to about 0.4 on unseen sources.
CLM-8B DeepSWE claim (best-of-4 verifier, 38 held-out tasks). We reproduce 31/38 = 81.6% exactly. Only 13 tasks can be changed by the selector. On those, the result is 10/13 against a random expectation of 7.0 (exact one-sided p = 0.062). The CLM base head before DeepSWE fine-tuning gets 27/38 = 71.1%, below random (73.7%).
Robustness checks
- Seeds. The recipe was trained with 3 seeds. Kev transfer-v4 test is 0.796 ± 0.002, against 0.763 ± 0.005 for the same recipe without the extra data (MiDM-4B-q35-e1). Typed-decisions test is 0.796 ± 0.007, against 0.801 ± 0.004.
- Package check. Loaded with this repo's
midm.pyonly, it scores td_holdout 0.835 (training repo: 0.833). - Latency. Batch 1 on an RTX 3090 in bf16 takes about 300 ms per question. That is without flash-linear-attention or causal-conv1d; installing them speeds up the Qwen3.5 linear-attention layers. Batched evaluation runs at about 70 ms per question.
- Known weakness. On JevBench public hard items it scores 0.45 against Jev's 0.73. Long inputs (4096 tokens) and long-document training did not close that gap.
Training
- Base:
Qwen/Qwen3.5-4B-Base, frozen, 4-bit NF4. LoRA r=16, alpha=32 on all attention/MLP (and linear-attention) projections. Pointer head: Linear(2560, 1). - Recipe: soft-target cross-entropy over options, options shuffled, 1 epoch(s), lr 2e-4, 4096-token batches.
- Data:
td_train: LocalLLaMA/typed-decisionsalltrain (Apache-2.0; synthetic, teacher-labelled)kev_train: jaredpalmer/kev-suites decision-v7 train (derived from agnews, yelp, dbpedia14, banking77, trec, mnli, imdb, sst5, boolq, amazon reviews, plus generated policy data; each source keeps its own terms)breadth_v1: built for this release from glue/mrpc, glue/qqp, glue/rte, go_emotions, civil_comments, winogrande (train splits; each keeps its own terms)pp6_new: jaredpalmer/kev-suites public-pool-v6: arc, openbookqa, csqa rows- Held-out evaluation suites were checked for zero text overlap with all training rows.
Limitations
- typed-decisions is synthetic and teacher-labelled, with public test labels.
- Gains on unseen sources are partly near-transfer. The largest gain is on knowledge multiple-choice (MMLU).
- Accuracy on the in-distribution typed-decisions test is lower than the narrower-data model.
- The model scores only the options you give it. It cannot answer open questions and does not abstain.
- Training data include datasets with their own licences or terms (e.g. Yelp, Amazon reviews). Review them before commercial use. The Apache-2.0 licence applies to this adapter and code only.
Citation
Authors: YeoHoon Yoon and Kyung-Sung Kim (Graduate School of AI, aSSIST University, Seoul, Republic of Korea).
Code, results and the audit: https://github.com/YeoHoonYun/midm-decision-models (Zenodo DOI 10.5281/zenodo.23084584).
@software{yoon_kim_2026_midm,
author = {Yoon, YeoHoon and Kim, Kyung-Sung},
title = {MiDM: Minimal Decision Models and an audit of CLM-8B},
year = {2026},
version = {0.1.0},
publisher = {Zenodo},
doi = {10.5281/zenodo.23084584}
}
Related release: MiDM v0.2.0 (9B)
The 4B weights and v0.1.0 tag are unchanged. A separately versioned 9B model and completed matched evaluation report are available. Some decision tasks improve, while SQL/Python selection does not improve consistently. The report integration study releases aggregate checks only and establishes no financial-accuracy superiority. Version DOI.
- Downloads last month
- 27
Model tree for yunicro/MiDM-4B-q35-e1-bx
Base model
Qwen/Qwen3.5-4B-Base

