Instructions to use StrandsAgents/strands-decider-12B-gemma4-v1-2610 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use StrandsAgents/strands-decider-12B-gemma4-v1-2610 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
strands-decider-12B-gemma4-v1-2610
The main home of this model is
StrandsAgents/strands-decider-12B-gemma4-v1-2610. The same files are also atamazon/strands-decider-12B-gemma4-v1-2610.
Strands decider is a small, fast decision model for agentic AI. Unlike an LLM, which can generate arbitrary text, a decision model picks between sets of options and rates things on a scale, and every answer carries a calibrated confidence. It fits the decisions inside an agentic workflow: model routing, tool selection, argument checking, triage, guardrails, evaluations, and the rote decisions of a hybrid agent that leaves the hard ones to an LLM.
This is v1, the first Strands decider on Gemma 4, on google/gemma-4-12B-it. It trains on
hobson-v21's data (123,339 rows) plus 12,263 code-task rows. The teacher is Qwen/Qwen3.5-4B: its
output distributions are training targets where its answer agrees with the gold label. A frozen-KL
anchor keeps the answers close to those of StrandsAgents/strands-decider-2B-hobson-v19. The
checkpoint is a soup of three seeds of this recipe:
each layer's LoRA update is the exact mean of the seeds' updates (stored as one rank-48 adapter),
and the score of an option is the mean of the three heads' scores. The soup was then calibrated.
Model names: strands-decider-<size>-<base>-v<N>-<YYMM>, that is the base model's size, the
base model family, the release on that family (v1 is the first) and the release month
(docs/naming.md in the code repository).
This repository holds a LoRA adapter on google/gemma-4-12B-it plus a small readout head that
scores the options of a typed question (noul, a yes/no question; choice, one of N options;
score, a level on an ordered scale). The code, the training recipe, the data inventory and the
evaluations are at https://github.com/strands-labs/strands-decider, under the same Apache-2.0
license as this model (see LICENSE.md).
Use
pip install "strands-decider @ git+https://github.com/strands-labs/strands-decider"
This installs the code repository's main, which has Gemma 4 support; the package on PyPI does not have it yet.
Ask one or more questions about a state. Pass --device cuda, mps or cpu; without it, the
best available device is used. The base weights download from the Hub at first use.
strands-decider ask StrandsAgents/strands-decider-12B-gemma4-v1-2610 \
--state "Help! My payouts have been failing for 3 days! " \
--choice "Which team should handle this?=billing,sales,retail" \
--noul "Does this convey urgency?" \
--score "How frustrated is the writer?=calm,frustrated,depressed"
The output of this checkpoint (--device cpu; other devices differ in the last digits):
noul_0 noul = 0.956
choice_0 -> billing (confidence 0.909)
billing 0.939
sales 0.040
retail 0.021
score_0 score = 0.99 (confidence 0.799)
0: calm 0.106
1: frustrated 0.794
2: depressed 0.101
Or serve it over HTTP. The server binds to 127.0.0.1 and has no authentication: use it for
local experiments.
strands-decider serve StrandsAgents/strands-decider-12B-gemma4-v1-2610 --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "Help! My payouts have been failing for 3 days!",
"questions": {"is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}}
}'
In Python, strands_decider.modeling.StrandsDeciderModel.load("StrandsAgents/strands-decider-12B-gemma4-v1-2610"). The adapter in lora/ is a
standard PEFT adapter on the Gemma 4 text decoder
(transformers.Gemma4ForConditionalGeneration(...).model.language_model); the head is head.safetensors.
Results
| evaluation | tasks right | Brier | ECE | served |
|---|---|---|---|---|
| JevBench public, window 4096 | 200/231 | 0.173 | 0.061 | these files |
| benchmark | score |
|---|---|
| Decision Index 0.2.1, balanced skill (window 32,768) | 47.14 |
| internal set | accuracy | n |
|---|---|---|
| eval: held-out short tasks | 0.708 | 6,000 |
| eval: boardgame | 0.891 | 900 |
| eval: contractnli | 0.876 | 1,026 |
| eval: hotpotqa (held out) | 0.911 | 959 |
| eval: musique | 0.936 | 1,199 |
| eval: musique, answerable | 0.947 | 599 |
| eval: musique, unanswerable | 0.925 | 600 |
| eval: [4] adequacy_hs2 | 0.795 | 234 |
| eval: [5] gen:adequacy | 0.884 | 302 |
| eval: [6] code/bug_swapped_operands | 0.910 | 89 |
| eval: [6] code/bug_variable_misuse | 0.831 | 89 |
| eval: [6] code/bug_wrong_binary_operator | 0.798 | 89 |
| eval: [6] code/commit_message | 0.989 | 89 |
| eval: [6] code/correct_solution | 0.539 | 89 |
| eval: [6] code/cwe_choice | 0.955 | 89 |
| eval: [6] code/docstring_match | 0.955 | 89 |
| eval: [6] code/exception_type | 0.876 | 89 |
| eval: [6] code/exec_input | 0.810 | 142 |
| eval: [6] code/exec_output | 0.863 | 662 |
| eval: [6] code/needle_function | 0.990 | 391 |
| eval: [6] code/review_needed | 0.607 | 89 |
JevBench is the external benchmark (231 public tasks); the Decision Index is the decision-model
index of the decision-index kit. The internal sets are this recipe's held-out short tasks,
multi-step documents, answer-adequacy judgements and code tasks. evaluation/README.md in the
code repository describes each set. Per-skill results and the raw reports and logs are in
eval/summary.json and eval/internal/.
Limitations
- Questions are read less than documents. With the state and options fixed, a changed question often gets the same answer. Phrase a question so the obvious reading is the intended one.
- Long, multi-step documents are the weak spot: JevBench's hard tier scores far below its easy tier.
scoreandnoultransfer poorly to rubrics and yes/no tasks unlike the training mix. Train on a rubric resembling yours.- Calibration is one temperature per primitive, fitted on held-out short classification. The confidence bands are established there only: measure on your own traffic before you trust a threshold.
- Trained on public datasets, so it inherits their domains and their label noise.
evaluation/README.md, section "Limitations", has the measured figures behind each point.
Training data
The datasets list in the metadata holds the 25 Hub datasets that the training rows come from:
18 short-task classification datasets, tasksource/Boardgame-QA, nvidia/HelpSteer2, and five code
datasets. The other sources are:
- ContractNLI and MuSiQue, from their authors' releases (not from the Hub).
- Synthetic rows that open-weight language models generated and checked.
- Code-task rows built here: generated programs with execution-checked answers, and functions from permissively licensed GitHub repositories. The five code datasets give the other code tasks (CuBERT ETH Py150 Open, CommitPackFT, CodeReviewer, Juliet C/C++ 1.3, CodeContests). Code rows carry gold labels only.
- The output distributions of the frozen teacher model
Qwen/Qwen3.5-4B. They are training targets, next to an anchor toward the distributions ofStrandsAgents/strands-decider-2B-hobson-v19.
These Hub datasets are for calibration and evaluation only, not training: dair-ai/emotion,
mteb/amazon_massive_intent, raquiba/Sarcasm_News_Headline, ucberkeley-dlab/measuring-hate-speech
and hotpotqa/hotpot_qa.
In the code repository, data/sources.md lists every source with its revision, role, license and
attribution. data/README.md states what is committed,
what is downloaded and the reproduction contract. Read a source's own terms before you redistribute
data built from it.
Training
Three seeds (0, 1, 2) of one config (train_config.json; the seeds differ only in seed), each
trained by training/recipe.sh train on one host with 8 NVIDIA H100 80GB GPUs (NGPU=8), one epoch
of 4,109 steps. The soup was made by strands-decider soup and calibrated and evaluated by
training/recipe.sh calibrate eval. training/README.md has the setup, the stages and the hardware notes.
Provenance
provenance.json: base model and revision (inferred: the hosts did not pin one), and the
sha256 of the config, adapter, head and source pickle. MANIFEST.sha256 lists every file;
python -m strands_decider.hf_export verify <folder> checks them. The run records keep
their timings and results; host paths, cloud identifiers and cost fields are removed.
Changelog
- 2026-10-09 (Hub tag
261009): first release. - 2026-10-09: card corrected (the recipe, the install line, the dataset list). The weights did not change.
- Downloads last month
- -
Model tree for StrandsAgents/strands-decider-12B-gemma4-v1-2610
Datasets used to train StrandsAgents/strands-decider-12B-gemma4-v1-2610
google-research-datasets/paws
fancyzhx/ag_news
Space using StrandsAgents/strands-decider-12B-gemma4-v1-2610 1
Evaluation results
- accuracy (200/231) on JevBench public, served at 4096self-reported0.866
- accuracy (n=6000) on eval: held-out short tasksself-reported0.708
- accuracy (n=900) on eval: boardgameself-reported0.891
- accuracy (n=1026) on eval: contractnliself-reported0.876
- accuracy (n=959) on eval: hotpotqa (held out)self-reported0.911
- accuracy (n=1199) on eval: musiqueself-reported0.936
- accuracy (n=599) on eval: musique, answerableself-reported0.947
- accuracy (n=600) on eval: musique, unanswerableself-reported0.925