--- license: apache-2.0 language: - en library_name: opendecision pipeline_tag: zero-shot-classification base_model: microsoft/deberta-v3-large tags: - decision - text-classification - intent-classification - routing - zero-shot-classification - multi-label - deberta-v3 datasets: - fastino/fast-decisions - stanfordnlp/snli - fancyzhx/dbpedia_14 - mteb/banking77 --- # OpenDecision-Large-Packed ๐Ÿ“„ Technical report: [*OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding*](https://github.com/tokz-labs/OpenDecision/releases/latest/download/OpenDecision.pdf) (PDF; arXiv version coming soon) ยท [Code](https://github.com/tokz-labs/OpenDecision) OpenDecision-Large-Packed is the low-cost variant of [OpenDecision-Large](https://huggingface.co/Tokz-labs/OpenDecision-Large): the same 433.9M-parameter DeBERTa-v3-large decision model, using a **packed candidate encoder** that scores up to 16 options in one forward pass. The input text is encoded once per pack; attention masks and per-sequence relative positions give every option the same view it would have alone, so answers do not depend on option order. ## Usage ```bash pip install "opendecision @ git+https://github.com/tokz-labs/OpenDecision" ``` ```python from opendecision import OpenDecision, Choice, Bool, MultiLabel model = OpenDecision.from_pretrained("Tokz-labs/OpenDecision-Large-Packed") out = model.decide( "I was charged twice for my May invoice and the app keeps logging me out.", { "queue": Choice(["billing", "technical", "account", "sales"]), "tags": MultiLabel(["double_charge", "login_issue", "refund_request", "outage"]), "urgent": Bool(), }, ) print(out["queue"].best, out["tags"].selected, out["urgent"].policy) ``` ## Efficiency GPU-seconds per 1,000 decisions on the complete classification-suite test sets, one NVIDIA L4: | Model | All five tasks | Without CLINC150 (151 options) | $ / 1k decisions (L4 list price) | |---|---|---|---| | **OpenDecision-Large-Packed** | 69.1 | **34.8** | 0.0153 | | OpenDecision-Large (cross-encoder) | 77.4 | 58.9 | 0.0172 | | Laya | 43.5 | 41.2 | 0.0097 | | JevK5 v0.3 (4B) | 267.4 | 119.9 | 0.0594 | | SemIf (4B) | 275.0 | 130.4 | 0.0611 | Packing pays off for questions with up to ~16 options; with many more (CLINC150's 151 intents) the cross-encoder is faster. Batch-size-1 latency (L4, bf16, p50): 66 ms for a short text with 10 options; 69 ms for a 512-token text with 10 options (cross-encoder: 248 ms). ## Benchmarks Classification suite, complete test sets, macro-F1 (open models, identical inputs; GLiNER2.5 with its best interface per dataset): | Model | AG News | CLINC150 | IMDb | Rotten Tomatoes | XNLI | Average | |---|---|---|---|---|---|---| | OpenDecision-Large | 87.37 | 76.38 | 94.84 | 92.03 | 88.93* | 87.91 | | **OpenDecision-Large-Packed** | 86.70 | 75.95 | 94.59 | 89.96 | 88.97* | 87.23 | | SemIf (Qwen3.5-4B, 4B) | 87.28 | 78.22 | 95.27 | 88.46 | 82.36 | 86.32 | | JevK5 v0.3 (4B) | 85.45 | 79.86 | 95.06 | 87.60 | 83.40 | 86.27 | | Laya (421M) | 92.42 | 47.57 | 92.16 | 87.99 | 87.22 | 81.47 | | GLiFormer large-v1 (576M) | 82.88 | 59.49 | 94.19 | 84.06 | 82.03 | 80.53 | | GLiNER2.5-Decide (340M) | 71.37 | 66.19 | 90.02 | 85.83 | 47.65 | 72.21 | \* NLI data is in the training data: XNLI is not zero-shot. On label sets that no training data targeted, averaged over TREC question types, Emotion and Subjectivity, it scores 55.64 macro-F1 (OpenDecision-Large 57.00, Laya 62.11, SemIf 64.29, JevK5 72.07); its final training stage, which adds 15,942 LLM-invented classification tasks with these task families excluded, raised this from 45.70. GoEmotions multi-label exact match is low (13.30; the model selects too many labels). For a new kind of decision, fine-tune on a few hundred labeled examples (`python -m opendecision.cli.train`). ## Training data Two stages from DeBERTa-v3-large (the packed stage starts from the clean cross-encoder) on LLM-generated decisions and LLM-invented classification tasks (OpenAI models), the Fast Decisions public development split, MultiNLI, WANLI and samples of DBpedia-14, BANKING77, SNLI, WikiText-103 and 20 Newsgroups (Mitchell 1997, CC BY 4.0). Every source allows commercial use under its license; sources and licenses are listed in the technical report. ## Scope and limitations English only; inputs up to 512 tokens; options truncated to 48 tokens. Not a chat or reasoning model. The benchmarks above were also used to select training mixtures. Most training text was generated or labeled by OpenAI models. ## Citation ```bibtex @misc{alwarawreh2026opendecision, title = {OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding}, author = {Alwarawreh, Abdallah}, year = {2026}, note = {Tokz Labs technical report}, url = {https://github.com/tokz-labs/OpenDecision/releases/latest/download/OpenDecision.pdf} } ```