Decision-1.0-Nox-4B
Nox, Latin for night.
Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime.
| Type | Use it for | Output |
|---|---|---|
| Choice | Route a request or choose among 2–255 actions. | Selected ID + distribution |
| Noul | Check a condition against supplied evidence. | P(true) |
| Score | Apply 2–10 ordered rubric descriptions. | Expected index + distribution |
Measured capability
73.09% weighted accuracy across 3,766 decisions and 54 tasks: +2.99 points over Kev-4B on the same benchmark.
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|---|---|---|---|---|---|---|---|
| Nox-4B | 4B | 83.00 | 51.79 | 79.06 | 86.25 | 69.60 | 73.09 |
| Lux-9B | 9B | 84.38 | 52.75 | 90.16 | 91.46 | 77.72 | 77.40 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
| Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
| Sol-2B | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
| Eos-0.8B | 0.8B | 65.94 | 46.04 | 70.31 | 81.67 | 52.01 | 61.89 |
| Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
| Kai-0.6B | 0.6B | 57.96 | 40.83 | 54.69 | 69.79 | 48.37 | 53.52 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
Accuracy (%). Overall weights: Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
All 54 tasks · Order, missing-evidence and calibration diagnostics · Methods and uncertainty
More questions, one request
Distinct Choice questions at a fixed 499 input tokens per question. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. These measurements precede null-description normalization and use explicit descriptions. p50, p95 and measurement scope.
Download the complete model repository
hf download llm-semantic-router/Decision-1.0-Nox-4B --local-dir Decision-1.0-Nox-4B
This downloads the complete model release. The root config.json lists the backbone, tokenizer, decision head, and calibration files.
Serve with vLLM Semantic Router
This repository contains model data only. Use the vLLM Semantic Router Decision runtime to load llm-semantic-router/Decision-1.0-Nox-4B and serve Choice, Noul, and Score requests. The serving implementation and its dependencies live in vLLM Semantic Router; this release does not bundle executable model code. transformers.AutoModel.from_pretrained cannot load the custom Decision head directly.
After configuring a compatible Decision endpoint, send a SystemOne request (replace the placeholder URL and key):
curl -X POST 'https://your-decision-endpoint.example/v1/systemone' \
-H 'Authorization: Bearer YOUR_ENDPOINT_API_KEY' \
-H 'Content-Type: application/json' \
--data-raw '{"model":"Decision-1.0-Nox-4B","state":"Customer requests a refund.","questions":{"route":{"type":"choice","instructions":"Which team should handle this?","criteria":{"billing":"Payments and refunds","technical":"Product faults"}}}}'
The Hugging Face repository is a model download, not a hosted inference endpoint.
The published model's complete state, question, and candidates have a 16,384-token input limit. See the evaluation scope for measured conditions.
Architecture
A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. The serving runtime schedules questions according to available hardware and request load.
Candidate head · Vector architecture
Adapted from Qwen3.5-4B. It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. License · Attributions.
- Downloads last month
- 31




