decider-q2b-1.1b

A 1.12B-parameter decision model. It takes a state and a typed question with developer-supplied options, and returns a probability for each option in one forward pass. It generates no text.

It ties Strands Decider 2B on JevBench's public set (167/231 each) with about 10k training rows (Strands: ~123k, per their architecture doc) and 0.6x the parameters. Strands has the better Brier score (calibration).

Write-up: Can a smaller model match AWS Strands Decider 2B? · Code and logs: decision-models-experiments

How it works

  • Torso: Qwen3.5-2B-Base with a rank-16 LoRA, merged into the weights.
  • Head: a pointer head (~1M parameters, LayerNorm, 4-head cross-attention, residual pointer) that scores each option against the <answer> position.
  • Read-out at layer 16 of 24. Layers 17-24 never influence a decision, so they are deleted (exact, no change in logits).
  • Vocabulary trimmed from ~248k to 98,304 ids plus specials. This cost one JevBench decision.
  • Result: 1.88B → 1.12B parameters. Weights are bf16 (2.1 GB).

The model is not a standard transformers checkpoint (custom head and read-out). The code/ folder holds the inference code.

Usage

pip install torch transformers numpy huggingface_hub
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("gabrielfior/decider-q2b-1.1b")
sys.path.insert(0, f"{path}/code")
from serve import Decider

model = Decider(f"{path}/torso", __import__("pathlib").Path(path), None, device="cpu", tta=2)

answer = model.decide(
    "Customer message: 'My card was charged twice and I want my money back today.'",
    {
        "type": "choice",
        "instructions": "Route this message to the right team.",
        "criteria": {
            "billing": "Payments, refunds, invoices",
            "tech": "Bugs and outages",
            "sales": "New purchases and pricing",
        },
    },
)
print(answer)  # {'type': 'choice', 'choice': 'billing', 'probabilities': {...}}

Question types:

type criteria Output
noul (yes/no) {"true": "...", "false": "..."} {"noul": p_yes}
choice {key: description, ...} probability per key
score ["level 0", "level 1", ...] (ordinal) probability per level

tta=2 averages two option orders for choice questions (about +2 decisions on JevBench, 2x cost). serve.py also exposes the JevBench /v1/systemone HTTP format: python code/serve.py --torso <path>/torso --run <path> --tta 2.

On a laptop CPU one decision takes about 2 s without optimised kernels; on a GPU the median is ~0.05-0.2 s.

Results

JevBench public set, 231 decisions (48 easy, 72 standard, 111 hard):

Model Params Train rows Correct Accuracy Brier ↓
Strands Decider 2B v19 1.9B ~123k 167 0.723 0.342
This model, one option order 1.12B 9.9k 165 0.714 0.374
This model, two option orders (tta=2) 1.12B 9.9k 167 0.723 0.372

On our 2,000-row dev split the model scores 0.769 accuracy / 0.311 Brier. A 3-seed ensemble of the untrimmed 1.88B model reaches 171/231 (not this checkpoint).

Training data

decision-v7 from the Kev project: public NLP datasets (AG News, Amazon, Banking77, BoolQ, DBpedia, IMDB, MNLI, SST-5, TREC, Yelp) plus generated policy and rule examples. This checkpoint trained on 9.9k rows (8,000 train rows plus the 1,896 rows of two families, TREC and legacy policy, that were held out in earlier runs). No distillation from other decision models.

We checked all 231 public JevBench items against all 12,576 decision-v7 records for shared runs of 8 words and found no overlap.

Limitations

  • Small benchmark. A four-decision gap on 231 items is within one standard error. Read the result as a tie.
  • Selection on the benchmark. We looked at JevBench at milestones to choose the torso, the two-order averaging and the 9.9k-row training set, so scores are somewhat adapted to these items.
  • Newer Strands version. I compare against the released v19. Strands' results page also lists a v20 at 169/231.
  • Calibration. Brier is worse than Strands' (0.372 vs 0.342). Per-type temperatures are baked into config.json.
  • Style match. Generated policy and rule rows may resemble JevBench's policy and routing items. We did not measure the effect.
  • Public split only. The private split could not be scored.
  • English, short inputs. Like Strands, long multi-step documents are a weak spot, and we have not tested other languages.
  • Inherits the biases and label noise of its public training datasets.

License and data terms

Weights are released under Apache 2.0, like the base model and the Kev models trained on the same data. decision-v7 carries no licence of its own and includes rows derived from sources with restrictive terms (for example Yelp). Check the terms of the underlying datasets before commercial use.

Citation

@misc{fior2026decider,
  author = {Gabriel Fior},
  title  = {decider-q2b-1.1b: a 1.12B decision model},
  year   = {2026},
  url    = {https://huggingface.co/gabrielfior/decider-q2b-1.1b}
}
Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gabrielfior/decider-q2b-1.1b

Finetuned
(101)
this model