OpenDecision-Small

📄 Technical report: OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding (PDF; arXiv version coming soon) · Code

OpenDecision-Small is a 70.6M-parameter encoder that makes decisions over options you define at call time: intent, routing, triage, sentiment, document type, yes/no gates and multi-label tags, several questions about the same text in one call. It is the small sibling of OpenDecision-Large (DeBERTa-v3-xsmall instead of DeBERTa-v3-large) and runs comfortably on CPU.

Usage

pip install "opendecision @ git+https://github.com/tokz-labs/OpenDecision"
from opendecision import OpenDecision, Choice, Bool, MultiLabel

model = OpenDecision.from_pretrained("Tokz-labs/OpenDecision-Small")
out = model.decide(
    "I was charged twice for my May invoice and the app keeps logging me out.",
    {
        "queue": Choice(["billing", "technical", "account", "sales"]),
        "tags": MultiLabel(["double_charge", "login_issue", "refund_request", "outage"]),
        "urgent": Bool(),
    },
)
print(out["queue"].best, out["tags"].selected, out["urgent"].policy)

Benchmarks

OpenDecision-Small (71M) GLiNER2.5-small (74M) GLiNER2.5-Decide (340M)
Classification suite, full test sets, avg macro-F1* 79.26 74.39 72.21
Held-out label sets (TREC, Emotion, Subjectivity), avg macro-F1 44.16 43.84 56.12
GPU-seconds per 1,000 suite decisions (NVIDIA L4) 17.5 — —

* AG News, CLINC150, IMDb, Rotten Tomatoes, XNLI; XNLI is not zero-shot (NLI data in training). Larger open models score higher (GLiFormer large-v1 80.53, Laya 81.47, OpenDecision-Large 87.91); OpenDecision-Small is the fastest model we measured (17.5 GPU-seconds per 1,000 decisions against 43.5 for Laya and 77.4 for OpenDecision-Large).

Accuracy is highest on decision types covered by the training data (support routing, intents, topics, sentiment, triage, NLI). The final training stage adds 15,942 classification tasks invented by an LLM (with the held-out task families excluded), which raised the held-out average from 43.51 to 44.16. GoEmotions multi-label exact match is 17.49 (GLiNER2.5-small 27.14). For a new schema, fine-tune on a few hundred labeled examples (python -m opendecision.cli.train).

Training data

Fine-tuned from DeBERTa-v3-xsmall in two stages on LLM-generated decisions and LLM-invented classification tasks (OpenAI models), the Fast Decisions public development split, MultiNLI, WANLI and samples of DBpedia-14, BANKING77, SNLI, WikiText-103 and 20 Newsgroups (Mitchell 1997, CC BY 4.0). Every source allows commercial use under its license; sources and licenses are listed in the technical report.

Scope and limitations

English only; inputs up to 512 tokens; options truncated to 32 tokens. Not a chat or reasoning model. Probabilities are overconfident without temperature scaling. Most training text was generated or labeled by OpenAI models.

Citation

@misc{alwarawreh2026opendecision,
  title  = {OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding},
  author = {Alwarawreh, Abdallah},
  year   = {2026},
  note   = {Tokz Labs technical report},
  url    = {https://github.com/tokz-labs/OpenDecision/releases/latest/download/OpenDecision.pdf}
}
Downloads last month
18
Safetensors
Model size
70.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tokz-labs/OpenDecision-Small

Finetuned
(61)
this model

Datasets used to train Tokz-labs/OpenDecision-Small