--- license: apache-2.0 base_model: Qwen/Qwen3.5-4B base_model_relation: finetune tags: - multilingual - decision-model - classification - qwen3_5 - pytorch - rocm --- ![Open Decision Foundation Models — Nox-4B](assets/decision-nox-4b-header.png) # Decision-1.0-Nox-4B *Nox, Latin for night.* Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime. [Decision family](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9) | Type | Use it for | Output | |---|---|---| | **Choice** | Route a request or choose among 2–255 actions. | Selected ID + distribution | | **Noul** | Check a condition against supplied evidence. | P(true) | | **Score** | Apply 2–10 ordered rubric descriptions. | Expected index + distribution | ## Measured capability **73.09% weighted accuracy** across 3,766 decisions and 54 tasks: **+2.99 points over Kev-4B** on the same benchmark. | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall | |---|---:|---:|---:|---:|---:|---:|---:| | Nox-4B | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 69.60 | **73.09** | | Lux-9B | 9B | **84.38** | **52.75** | 90.16 | **91.46** | 77.72 | **77.40** | | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 | | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 | | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 | | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 | | Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 | | Sol-2B | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 | | Eos-0.8B | 0.8B | 65.94 | 46.04 | 70.31 | 81.67 | 52.01 | 61.89 | | Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 | | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 | | Kai-0.6B | 0.6B | 57.96 | 40.83 | 54.69 | 69.79 | 48.37 | 53.52 | | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 | | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 | | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 | Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models. ![Decision model ranking](assets/decision-ranking.png) ![Capability matrix](assets/decision-matrix.png) [All 54 tasks](evaluation/TASKS.md) · [Order, missing-evidence and calibration diagnostics](evaluation/DIAGNOSTICS.md) · [Methods and uncertainty](evaluation/EVALUATION.md) ## More questions, one request ![Question-count latency](assets/decision-question-scaling.png) Distinct Choice questions at a fixed **499 input tokens per question**. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. These measurements precede null-description normalization and use explicit descriptions. [p50, p95 and measurement scope](evaluation/QUESTION-SCALING.md). ## Download the complete model repository ```bash hf download llm-semantic-router/Decision-1.0-Nox-4B --local-dir Decision-1.0-Nox-4B ``` This downloads the complete model release. The root `config.json` lists the backbone, tokenizer, decision head, and calibration files. ## Serve with vLLM Semantic Router This repository contains model data only. Use the vLLM Semantic Router Decision runtime to load `llm-semantic-router/Decision-1.0-Nox-4B` and serve Choice, Noul, and Score requests. The serving implementation and its dependencies live in vLLM Semantic Router; this release does not bundle executable model code. `transformers.AutoModel.from_pretrained` cannot load the custom Decision head directly. After configuring a compatible Decision endpoint, send a [SystemOne request](https://docs.typesafe.ai/api) (replace the placeholder URL and key): ```bash curl -X POST 'https://your-decision-endpoint.example/v1/systemone' \ -H 'Authorization: Bearer YOUR_ENDPOINT_API_KEY' \ -H 'Content-Type: application/json' \ --data-raw '{"model":"Decision-1.0-Nox-4B","state":"Customer requests a refund.","questions":{"route":{"type":"choice","instructions":"Which team should handle this?","criteria":{"billing":"Payments and refunds","technical":"Product faults"}}}}' ``` The Hugging Face repository is a model download, not a hosted inference endpoint. The published model's complete state, question, and candidates have a 16,384-token input limit. See the [evaluation scope](evaluation/EVALUATION.md) for measured conditions. ## Architecture ![Decision decoder architecture](assets/architecture.png) A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. The serving runtime schedules questions according to available hardware and request load. [Candidate head](assets/readout.png) · [Vector architecture](assets/architecture.svg) Adapted from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. [License](LICENSE) · [Attributions](ATTRIBUTIONS.md).