OpenDecision-Large-Packed

📄 Technical report: OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding (PDF; arXiv version coming soon) · Code

OpenDecision-Large-Packed is the low-cost variant of OpenDecision-Large: the same 433.9M-parameter DeBERTa-v3-large decision model, using a packed candidate encoder that scores up to 16 options in one forward pass. The input text is encoded once per pack; attention masks and per-sequence relative positions give every option the same view it would have alone, so answers do not depend on option order.

Usage

pip install "opendecision @ git+https://github.com/tokz-labs/OpenDecision"
from opendecision import OpenDecision, Choice, Bool, MultiLabel

model = OpenDecision.from_pretrained("Tokz-labs/OpenDecision-Large-Packed")
out = model.decide(
    "I was charged twice for my May invoice and the app keeps logging me out.",
    {
        "queue": Choice(["billing", "technical", "account", "sales"]),
        "tags": MultiLabel(["double_charge", "login_issue", "refund_request", "outage"]),
        "urgent": Bool(),
    },
)
print(out["queue"].best, out["tags"].selected, out["urgent"].policy)

Efficiency

GPU-seconds per 1,000 decisions on the complete classification-suite test sets, one NVIDIA L4:

Model All five tasks Without CLINC150 (151 options) $ / 1k decisions (L4 list price)
OpenDecision-Large-Packed 69.1 34.8 0.0153
OpenDecision-Large (cross-encoder) 77.4 58.9 0.0172
Laya 43.5 41.2 0.0097
JevK5 v0.3 (4B) 267.4 119.9 0.0594
SemIf (4B) 275.0 130.4 0.0611

Packing pays off for questions with up to ~16 options; with many more (CLINC150's 151 intents) the cross-encoder is faster. Batch-size-1 latency (L4, bf16, p50): 66 ms for a short text with 10 options; 69 ms for a 512-token text with 10 options (cross-encoder: 248 ms).

Benchmarks

Classification suite, complete test sets, macro-F1 (open models, identical inputs; GLiNER2.5 with its best interface per dataset):

Model AG News CLINC150 IMDb Rotten Tomatoes XNLI Average
OpenDecision-Large 87.37 76.38 94.84 92.03 88.93* 87.91
OpenDecision-Large-Packed 86.70 75.95 94.59 89.96 88.97* 87.23
SemIf (Qwen3.5-4B, 4B) 87.28 78.22 95.27 88.46 82.36 86.32
JevK5 v0.3 (4B) 85.45 79.86 95.06 87.60 83.40 86.27
Laya (421M) 92.42 47.57 92.16 87.99 87.22 81.47
GLiFormer large-v1 (576M) 82.88 59.49 94.19 84.06 82.03 80.53
GLiNER2.5-Decide (340M) 71.37 66.19 90.02 85.83 47.65 72.21

* NLI data is in the training data: XNLI is not zero-shot.

On label sets that no training data targeted, averaged over TREC question types, Emotion and Subjectivity, it scores 55.64 macro-F1 (OpenDecision-Large 57.00, Laya 62.11, SemIf 64.29, JevK5 72.07); its final training stage, which adds 15,942 LLM-invented classification tasks with these task families excluded, raised this from 45.70. GoEmotions multi-label exact match is low (13.30; the model selects too many labels). For a new kind of decision, fine-tune on a few hundred labeled examples (python -m opendecision.cli.train).

Training data

Two stages from DeBERTa-v3-large (the packed stage starts from the clean cross-encoder) on LLM-generated decisions and LLM-invented classification tasks (OpenAI models), the Fast Decisions public development split, MultiNLI, WANLI and samples of DBpedia-14, BANKING77, SNLI, WikiText-103 and 20 Newsgroups (Mitchell 1997, CC BY 4.0). Every source allows commercial use under its license; sources and licenses are listed in the technical report.

Scope and limitations

English only; inputs up to 512 tokens; options truncated to 48 tokens. Not a chat or reasoning model. The benchmarks above were also used to select training mixtures. Most training text was generated or labeled by OpenAI models.

Citation

@misc{alwarawreh2026opendecision,
  title  = {OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding},
  author = {Alwarawreh, Abdallah},
  year   = {2026},
  note   = {Tokz Labs technical report},
  url    = {https://github.com/tokz-labs/OpenDecision/releases/latest/download/OpenDecision.pdf}
}
Downloads last month
22
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tokz-labs/OpenDecision-Large-Packed

Finetuned
(316)
this model

Datasets used to train Tokz-labs/OpenDecision-Large-Packed