Instructions to use gabrielfior/decider-q2b-1.1b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use gabrielfior/decider-q2b-1.1b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="gabrielfior/decider-q2b-1.1b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("gabrielfior/decider-q2b-1.1b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
decider-q2b-1.1b
A 1.12B-parameter decision model. It takes a state and a typed question with developer-supplied options, and returns a probability for each option in one forward pass. It generates no text.
It ties Strands Decider 2B on JevBench's public set (167/231 each) with about 10k training rows (Strands: ~123k, per their architecture doc) and 0.6x the parameters. Strands has the better Brier score (calibration).
Write-up: Can a smaller model match AWS Strands Decider 2B? · Code and logs: decision-models-experiments
How it works
- Torso: Qwen3.5-2B-Base with a rank-16 LoRA, merged into the weights.
- Head: a pointer head (~1M parameters, LayerNorm, 4-head cross-attention, residual pointer) that scores each option against the
<answer>position. - Read-out at layer 16 of 24. Layers 17-24 never influence a decision, so they are deleted (exact, no change in logits).
- Vocabulary trimmed from ~248k to 98,304 ids plus specials. This cost one JevBench decision.
- Result: 1.88B → 1.12B parameters. Weights are bf16 (2.1 GB).
The model is not a standard transformers checkpoint (custom head and read-out). The code/ folder holds the inference code.
Usage
pip install torch transformers numpy huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("gabrielfior/decider-q2b-1.1b")
sys.path.insert(0, f"{path}/code")
from serve import Decider
model = Decider(f"{path}/torso", __import__("pathlib").Path(path), None, device="cpu", tta=2)
answer = model.decide(
"Customer message: 'My card was charged twice and I want my money back today.'",
{
"type": "choice",
"instructions": "Route this message to the right team.",
"criteria": {
"billing": "Payments, refunds, invoices",
"tech": "Bugs and outages",
"sales": "New purchases and pricing",
},
},
)
print(answer) # {'type': 'choice', 'choice': 'billing', 'probabilities': {...}}
Question types:
type |
criteria |
Output |
|---|---|---|
noul (yes/no) |
{"true": "...", "false": "..."} |
{"noul": p_yes} |
choice |
{key: description, ...} |
probability per key |
score |
["level 0", "level 1", ...] (ordinal) |
probability per level |
tta=2 averages two option orders for choice questions (about +2 decisions on JevBench, 2x cost). serve.py also exposes the JevBench /v1/systemone HTTP format: python code/serve.py --torso <path>/torso --run <path> --tta 2.
On a laptop CPU one decision takes about 2 s without optimised kernels; on a GPU the median is ~0.05-0.2 s.
Results
JevBench public set, 231 decisions (48 easy, 72 standard, 111 hard):
| Model | Params | Train rows | Correct | Accuracy | Brier ↓ |
|---|---|---|---|---|---|
| Strands Decider 2B v19 | 1.9B | ~123k | 167 | 0.723 | 0.342 |
| This model, one option order | 1.12B | 9.9k | 165 | 0.714 | 0.374 |
This model, two option orders (tta=2) |
1.12B | 9.9k | 167 | 0.723 | 0.372 |
On our 2,000-row dev split the model scores 0.769 accuracy / 0.311 Brier. A 3-seed ensemble of the untrimmed 1.88B model reaches 171/231 (not this checkpoint).
Training data
decision-v7 from the Kev project: public NLP datasets (AG News, Amazon, Banking77, BoolQ, DBpedia, IMDB, MNLI, SST-5, TREC, Yelp) plus generated policy and rule examples. This checkpoint trained on 9.9k rows (8,000 train rows plus the 1,896 rows of two families, TREC and legacy policy, that were held out in earlier runs). No distillation from other decision models.
We checked all 231 public JevBench items against all 12,576 decision-v7 records for shared runs of 8 words and found no overlap.
Limitations
- Small benchmark. A four-decision gap on 231 items is within one standard error. Read the result as a tie.
- Selection on the benchmark. We looked at JevBench at milestones to choose the torso, the two-order averaging and the 9.9k-row training set, so scores are somewhat adapted to these items.
- Newer Strands version. I compare against the released v19. Strands' results page also lists a v20 at 169/231.
- Calibration. Brier is worse than Strands' (0.372 vs 0.342). Per-type temperatures are baked into
config.json. - Style match. Generated policy and rule rows may resemble JevBench's policy and routing items. We did not measure the effect.
- Public split only. The private split could not be scored.
- English, short inputs. Like Strands, long multi-step documents are a weak spot, and we have not tested other languages.
- Inherits the biases and label noise of its public training datasets.
License and data terms
Weights are released under Apache 2.0, like the base model and the Kev models trained on the same data. decision-v7 carries no licence of its own and includes rows derived from sources with restrictive terms (for example Yelp). Check the terms of the underlying datasets before commercial use.
Citation
@misc{fior2026decider,
author = {Gabriel Fior},
title = {decider-q2b-1.1b: a 1.12B decision model},
year = {2026},
url = {https://huggingface.co/gabrielfior/decider-q2b-1.1b}
}
- Downloads last month
- 25
Model tree for gabrielfior/decider-q2b-1.1b
Base model
Qwen/Qwen3.5-2B-Base