# BTZSC-22 add-on: DiffusionGemma-26B-A4B-it (2026-10-06) **Added after the `btzsc22-ext-v1` freeze.** The protocol, the manifest (2,200 examples) and the scorer are unchanged; this directory adds one more model on **Track A (native) only**. Track B (one-vs-rest) was not run for it. | | | |---|---| | model | `RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic` (FP8 of Google's DiffusionGemma-26B-A4B-it) | | serving | `razorback16/openjev:0.4.0` (vLLM structured reads, `/v1/systemone`), model `openjev-0.1`, `think: 0` (the server default), canvas 64, 1× H200 | | request | identical to the other models: `state = {"text": }`, one Choice question `"Which single label best describes the input text?"`, options `label_000…` with the label texts as descriptions | | interface limit | 255 options, so all 22 datasets are a **single request** (no knockout) | | failures | 0 of 2,200 | ## Files - `trackA_native.gemma.jsonl`: one prediction per example (full native probability vector). `latency_seconds` was recorded under 16 concurrent clients and is **not** a single-request latency (see `../../speed.md`). - `code/run_gemma.py`: the client used. - `code/score_addon_gemma.py`: `../code/score_final.py` with Gemma added to the model list. The Track A logic, the paired hierarchical bootstrap and the calibration code are copied verbatim. Rerunning it reproduces the four original models' numbers in `../results.json` exactly (0.474 / 0.525 / 0.530 / 0.537). - `results.json`, `per_dataset.csv`: the output. Gemma's Track B fields are `null` / `not_run`. ## Result (Track A, equal-weight mean macro-F1 over datasets) | group | Ines-1 | Julia-1 | laya-multilingual | GLiNER2.5-multi | DiffusionGemma-26B-A4B | |---|---:|---:|---:|---:|---:| | all 22 | 0.474 | 0.525 | 0.530 | 0.537 | **0.799** | | 2 labels (13) | 0.562 | 0.583 | 0.722 | 0.627 | **0.897** | | 3–20 labels (4) | 0.489 | 0.634 | 0.559 | 0.541 | **0.695** | | 21–72 labels (5) | 0.234 | 0.287 | 0.008 | 0.297 | **0.628** | | absent from Ines-1's decision FT mix (15) | 0.486 | 0.536 | 0.515 | 0.528 | **0.765** | Gemma − Ines-1, all 22: **+0.326** (95 % CI +0.253 to +0.392, paired hierarchical bootstrap). Native calibration (17 datasets with ≤ 20 labels): ECE 0.116 (Ines-1 0.127), Brier 0.250 (Ines-1 0.501), coverage at 5 % error budget 0.694 (Ines-1 0.089). DiffusionGemma-26B-A4B has ~26B total parameters (~4B active per token), about 16× Ines-1's total and 10× its active parameters. Speed and memory for the same models are in `../../speed.md` (run 3).