OpenDecision-Small
📄 Technical report: OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding (PDF; arXiv version coming soon) · Code
OpenDecision-Small is a 70.6M-parameter encoder that makes decisions over options you define at call time: intent, routing, triage, sentiment, document type, yes/no gates and multi-label tags, several questions about the same text in one call. It is the small sibling of OpenDecision-Large (DeBERTa-v3-xsmall instead of DeBERTa-v3-large) and runs comfortably on CPU.
Usage
pip install "opendecision @ git+https://github.com/tokz-labs/OpenDecision"
from opendecision import OpenDecision, Choice, Bool, MultiLabel
model = OpenDecision.from_pretrained("Tokz-labs/OpenDecision-Small")
out = model.decide(
"I was charged twice for my May invoice and the app keeps logging me out.",
{
"queue": Choice(["billing", "technical", "account", "sales"]),
"tags": MultiLabel(["double_charge", "login_issue", "refund_request", "outage"]),
"urgent": Bool(),
},
)
print(out["queue"].best, out["tags"].selected, out["urgent"].policy)
Benchmarks
| OpenDecision-Small (71M) | GLiNER2.5-small (74M) | GLiNER2.5-Decide (340M) | |
|---|---|---|---|
| Classification suite, full test sets, avg macro-F1* | 79.26 | 74.39 | 72.21 |
| Held-out label sets (TREC, Emotion, Subjectivity), avg macro-F1 | 44.16 | 43.84 | 56.12 |
| GPU-seconds per 1,000 suite decisions (NVIDIA L4) | 17.5 | — | — |
* AG News, CLINC150, IMDb, Rotten Tomatoes, XNLI; XNLI is not zero-shot (NLI data in training). Larger open models score higher (GLiFormer large-v1 80.53, Laya 81.47, OpenDecision-Large 87.91); OpenDecision-Small is the fastest model we measured (17.5 GPU-seconds per 1,000 decisions against 43.5 for Laya and 77.4 for OpenDecision-Large).
Accuracy is highest on decision types covered by the training data (support routing, intents, topics,
sentiment, triage, NLI). The final training stage adds 15,942 classification tasks invented by
an LLM (with the held-out task families excluded), which raised the held-out average from
43.51 to 44.16. GoEmotions multi-label exact match is
17.49 (GLiNER2.5-small 27.14). For a new schema, fine-tune on
a few hundred labeled examples (python -m opendecision.cli.train).
Training data
Fine-tuned from DeBERTa-v3-xsmall in two stages on LLM-generated decisions and LLM-invented classification tasks (OpenAI models), the Fast Decisions public development split, MultiNLI, WANLI and samples of DBpedia-14, BANKING77, SNLI, WikiText-103 and 20 Newsgroups (Mitchell 1997, CC BY 4.0). Every source allows commercial use under its license; sources and licenses are listed in the technical report.
Scope and limitations
English only; inputs up to 512 tokens; options truncated to 32 tokens. Not a chat or reasoning model. Probabilities are overconfident without temperature scaling. Most training text was generated or labeled by OpenAI models.
Citation
@misc{alwarawreh2026opendecision,
title = {OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding},
author = {Alwarawreh, Abdallah},
year = {2026},
note = {Tokz Labs technical report},
url = {https://github.com/tokz-labs/OpenDecision/releases/latest/download/OpenDecision.pdf}
}
- Downloads last month
- 18
Model tree for Tokz-labs/OpenDecision-Small
Base model
microsoft/deberta-v3-xsmall