Endikavi commited on
Commit
23b15ec
·
verified ·
1 Parent(s): 2e965de

Comparison tables: BTZSC-22 with DiffusionGemma-26B-A4B add-on (Track A), speed/concurrency/GPU-memory run 3 (Ines-1, laya, GLiNER2.5-multi, Gemma)

Browse files
README.md CHANGED
@@ -36,7 +36,7 @@ datasets:
36
  > `tokenizer/`; the documentation (`README.md` and the other `.md` files, `paper/`, `assets/`, `tasksource_license_audit.csv`); the
37
  > training records (`training/*.json`); the model's own evaluation outputs in `eval/`; `SHA256SUMS`.
38
  > - **MIT** ([LICENSE-CODE](LICENSE-CODE)): the code — `mini_v41/`, `mini_v41_jev/`, `scripts/`, `examples/`,
39
- > `training/code/`, `eval/btzsc22/code/` — and the environment and build files: `Dockerfile`, `.dockerignore`, `requirements.txt`,
40
  > `requirements.lock`, `.gitattributes`, `.gitignore`.
41
  > - **Third-party terms, not relicensed:** the gold labels and teacher distributions of the public test sets that
42
  > `eval/items/`, `eval/reference_rows.json` and `eval/btzsc22/` (BTZSC label texts and targets) include keep their
@@ -199,6 +199,21 @@ learning the teacher's quirks.
199
  We evaluated Ines-1 on 22 public BTZSC classification datasets (100 examples per dataset), using an extension of the
200
  jev-benchmarks protocol. Fifteen of the 22 datasets are absent from Ines-1's decision fine-tuning mixture.
201
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
202
  Under the model-specific classification track, mean macro-F1 across datasets was 0.474 for Ines-1, 0.525 for Julia-1,
203
  0.530 for laya-multilingual, and 0.537 for GLiNER2.5. Using a paired hierarchical bootstrap with datasets as the unit of
204
  inference, the difference relative to Ines-1 was resolved for GLiNER2.5 (+0.063, 95% CI +0.014 to +0.107), but not for
@@ -244,19 +259,26 @@ Cost: 21,423 domain cases, 23,608 questions per epoch, 28.5M prompt tokens, 2.25
244
  "Engram removed at inference" = every Engram module returns its input (gate 0) on this checkpoint, which was trained
245
  with Engram: it measures this model's reliance on Engram, not Engram's contribution to training.
246
 
247
- ### Speed
 
 
 
248
 
249
- [eval/speed.md](eval/speed.md) has the full protocol and both runs. One H200, bf16, HTTP closed loop, same requests and
250
- load generator for every model. Two regimes, kept apart:
 
 
 
 
251
 
252
- | | single request p50 / p95 | max throughput (questions/s) |
253
- |---|---:|---:|
254
- | short request (2 questions, ~60 prompt tokens each) — Ines-1, 1 process | 9.6 / 9.7 ms | ~1,900 |
255
- | same, 144M / 322M encoders, 4 processes each | 12.7 / 21–26 ms | ~1,700–1,850 |
256
- | typed-decisions case (5 questions, ~315 tokens each) — Ines-1 | 29.6 / 38.3 ms | ~540 |
257
- | same, 144M / 322M encoders | 15–16 / 87–88 ms | ~760–1,110 |
258
 
259
- On ~300-token questions **the encoders are 1.4–2.1× faster**.
 
 
 
260
 
261
  ## Training
262
 
 
36
  > `tokenizer/`; the documentation (`README.md` and the other `.md` files, `paper/`, `assets/`, `tasksource_license_audit.csv`); the
37
  > training records (`training/*.json`); the model's own evaluation outputs in `eval/`; `SHA256SUMS`.
38
  > - **MIT** ([LICENSE-CODE](LICENSE-CODE)): the code — `mini_v41/`, `mini_v41_jev/`, `scripts/`, `examples/`,
39
+ > `training/code/`, `eval/btzsc22/code/`, `eval/btzsc22/addon_gemma/code/` — and the environment and build files: `Dockerfile`, `.dockerignore`, `requirements.txt`,
40
  > `requirements.lock`, `.gitattributes`, `.gitignore`.
41
  > - **Third-party terms, not relicensed:** the gold labels and teacher distributions of the public test sets that
42
  > `eval/items/`, `eval/reference_rows.json` and `eval/btzsc22/` (BTZSC label texts and targets) include keep their
 
199
  We evaluated Ines-1 on 22 public BTZSC classification datasets (100 examples per dataset), using an extension of the
200
  jev-benchmarks protocol. Fifteen of the 22 datasets are absent from Ines-1's decision fine-tuning mixture.
201
 
202
+ | mean macro-F1 (Track A, native interface) | Ines-1 | Julia-1 | laya-multilingual | GLiNER2.5-multi | DiffusionGemma-26B-A4B¹ |
203
+ |---|---:|---:|---:|---:|---:|
204
+ | parameters (total / active per token) | 1.59B / 405M | 144M | 322M | 287M | ~26B / ~4B |
205
+ | **all 22 datasets** | 0.474 | 0.525 | 0.530 | 0.537 | **0.799** |
206
+ | 2 labels (13) | 0.562 | 0.583 | 0.722 | 0.627 | **0.897** |
207
+ | 3–20 labels (4) | 0.489 | 0.634 | 0.559 | 0.541 | **0.695** |
208
+ | 21–72 labels (5) | 0.234 | 0.287 | 0.008 | 0.297 | **0.628** |
209
+ | absent from Ines-1's decision FT mix (15) | 0.486 | 0.536 | 0.515 | 0.528 | **0.765** |
210
+ | difference vs Ines-1, all 22 (95 % CI) | — | +0.051 (−0.026, +0.126) | +0.056 (−0.040, +0.163) | +0.063 (+0.014, +0.107) | +0.326 (+0.253, +0.392) |
211
+ | native ECE, ≤ 20 labels (17 datasets) | 0.127 | 0.298 | 0.093 | 0.159 | 0.116 |
212
+
213
+ ¹ Added on 2026-10-06, after the protocol freeze, on Track A only (FP8 weights served by OpenJev 0.4.0, `think: 0`,
214
+ same request): [`eval/btzsc22/addon_gemma/`](eval/btzsc22/addon_gemma/). The scorer reproduces the other four columns
215
+ exactly. It is a ~16× larger generalist model; speed and memory for the same models are under *Speed* below.
216
+
217
  Under the model-specific classification track, mean macro-F1 across datasets was 0.474 for Ines-1, 0.525 for Julia-1,
218
  0.530 for laya-multilingual, and 0.537 for GLiNER2.5. Using a paired hierarchical bootstrap with datasets as the unit of
219
  inference, the difference relative to Ines-1 was resolved for GLiNER2.5 (+0.063, 95% CI +0.014 to +0.107), but not for
 
259
  "Engram removed at inference" = every Engram module returns its input (gate 0) on this checkpoint, which was trained
260
  with Engram: it measures this model's reliance on Engram, not Engram's contribution to training.
261
 
262
+ ### Speed and memory
263
+
264
+ [eval/speed.md](eval/speed.md) has the full protocol and all runs. One H200 per model in turn (nothing else on it), HTTP
265
+ closed loop, same requests and load generator for every model (run 3, 2026-10-06):
266
 
267
+ | model | short request¹ c=1 p50 / p95 | short: max questions/s | typed case² c=1 p50 / p95 | typed: max questions/s | GPU memory idle / peak |
268
+ |---|---:|---:|---:|---:|---:|
269
+ | **Ines-1** (1 process, CUDA graphs) | **9.4 / 9.7 ms** | **1,896** | 30.6 / **40.5 ms** | 544 | 9.7 / 21.9 GB |
270
+ | laya-multilingual architecture (4 processes) | 13.0 / 25.9 ms | 1,848 | **20.6** / 89.2 ms | **840** | 8.0 / 24.1 GB |
271
+ | GLiNER2.5-multi architecture (4 workers, no batching) | 48.8 / 49.3 ms | 68 | 123.1 / 129.3 ms | 49 | 6.7 / 26.5 GB |
272
+ | DiffusionGemma-26B-A4B FP8 (OpenJev / vLLM) | 91.7 / 94.0 ms | 79 | 158.8 / 170.2 ms | 143 | 25.8 GiB weights; 120.6 / 135.1 GB reserved³ |
273
 
274
+ ¹ 2 questions, ~60 prompt tokens each (Ines-1 template). ² typed-decisions case, 5 questions, ~300 prompt tokens each
275
+ (Ines-1 template; each server counts tokens differently, see eval/speed.md).
276
+ ³ vLLM preallocates KV cache (`gpu_memory_utilization` 0.80); the floor is the weights.
 
 
 
277
 
278
+ Ines-1 has the lowest single-request latency and matches the encoder's peak short-request throughput (within the
279
+ ±10 % run-to-run variation); on ~300-token questions the
280
+ 322M encoder is faster (1.5× the throughput, lower median) while Ines-1 keeps the tighter tail. The GLiNER row measures
281
+ a server without batching, not the architecture's ceiling. Different serving stacks: not a general speed ranking.
282
 
283
  ## Training
284
 
SHA256SUMS CHANGED
@@ -8,7 +8,7 @@ cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
8
  6a7ddcad077839a44697c005ea7e6f455caecaa8054412913108d6e21a1c9dbe LICENSE-CODE
9
  121b31536257bd7147d9c9b054f6fb4e2bb68734024415c081ac0b50c7cf48f1 LICENSING_NOTES.md
10
  1faada60db1ccfc073ae063fc474d9f75cba0c9d2683d03ad66b2074efeb166a NOTICE
11
- 8b3d5abc94f861cd0c3622a42f3403c7bfb33938009a629b05c88a9fad3f2d2a README.md
12
  6ec8a2bb55f9331d77b9adbeae7fd71eeb5614a988cd21eee18c7bde775df5ba TASKSOURCE_LICENSE_AUDIT.md
13
  1a4df8f10bd30b01e6acc19ef2d935faee5485f0c25452ff64ba6c8598fabd97 THIRD_PARTY_DATA_NOTICE.md
14
  9708385cf095b543e009b0dd6013aa9985ef68a5b05c5ebef4c14c6a7f97cfa2 assets/architecture.png
@@ -16,6 +16,13 @@ cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
16
  757003b29ec1d935f236c06ea1e39cf9362c43de358f8b3df74326f8bfdbd40d config.json
17
  8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b eval/btzsc22/PROTOCOL.md
18
  5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 eval/btzsc22/REPORT.md
 
 
 
 
 
 
 
19
  95a4d4c459940195f03c6b937d750cd8e678527c4e2acaee1f7f2f3a9adc512c eval/btzsc22/btzsc_test_files.json
20
  bb7dbba94cc130a09c77226e51cb7b72e5228685fbf8c2f78522290f485eff1e eval/btzsc22/code/README.md
21
  afac77fa429f2f37c45e6cf46d0201551e9c98ddf9c01e7ba0cf1ac1feca0ef7 eval/btzsc22/code/btzsc22.yaml
@@ -57,8 +64,10 @@ be0892eb9d9e45f2173e59052d503b26d809253b01bc44a8a70c0e7c48add480 eval/items/v3_
57
  d0b5cb709b104e8fafcdba8f8c676896ce048c2a009f93c1b19ad3e949fcb128 eval/reference_rows.json
58
  508a3e33bf77cd4ab9f559d5dcd2ec3e473c30b3b1d54c90a7de5e6490a435a2 eval/report.json
59
  d5f3f5d214ac20264176a552c7779ddbbee79f97524b64f8c6d1804e1f305ba2 eval/report.md
60
- 6b38584ac67bb856448275789e45bef4bf104fd9462abaec984d232c05dc31ab eval/speed.md
61
  d0b6fcd59367eb58c02eb3b21e7d6c50caa6eccf5d822e8e7ecec22f0e3a275b eval/speed_raw.log
 
 
62
  22974cc87fd92b5c21168e350ff3db492209a98293ec28a7f5d70396f13c970d examples/quickstart.py
63
  6c4fc03dc8ec311323338f6f4f6c0a8a3514f417b57c99f70827d797fe2ead28 mini_v41/__init__.py
64
  7c57633d0d69fe8a709e0629cbe36b07590d94d0333039c5378fa66df4a74f05 mini_v41/attention.py
 
8
  6a7ddcad077839a44697c005ea7e6f455caecaa8054412913108d6e21a1c9dbe LICENSE-CODE
9
  121b31536257bd7147d9c9b054f6fb4e2bb68734024415c081ac0b50c7cf48f1 LICENSING_NOTES.md
10
  1faada60db1ccfc073ae063fc474d9f75cba0c9d2683d03ad66b2074efeb166a NOTICE
11
+ aed9b754a211decd074a0ddccd654852fc3797a04376bc80b882ea47cc200f58 README.md
12
  6ec8a2bb55f9331d77b9adbeae7fd71eeb5614a988cd21eee18c7bde775df5ba TASKSOURCE_LICENSE_AUDIT.md
13
  1a4df8f10bd30b01e6acc19ef2d935faee5485f0c25452ff64ba6c8598fabd97 THIRD_PARTY_DATA_NOTICE.md
14
  9708385cf095b543e009b0dd6013aa9985ef68a5b05c5ebef4c14c6a7f97cfa2 assets/architecture.png
 
16
  757003b29ec1d935f236c06ea1e39cf9362c43de358f8b3df74326f8bfdbd40d config.json
17
  8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b eval/btzsc22/PROTOCOL.md
18
  5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 eval/btzsc22/REPORT.md
19
+ 775c02f424c71a5b944ccaf19ba466ad1444644e585cf4785e9fe39d1a0e6140 eval/btzsc22/SHA256SUMS
20
+ cc8ce27ca26efcdc9985323e75fdbad20b602d69b1602c2883a1c80d37dbe15e eval/btzsc22/addon_gemma/README.md
21
+ 80bc6a7b10e75160f88265611a7be4fc07d625980a55697cee3ff3a60dfc5632 eval/btzsc22/addon_gemma/code/run_gemma.py
22
+ d4a2cb890861b2076e79934b5a1dd0f184c5b4418dec4916c8288def58a279e9 eval/btzsc22/addon_gemma/code/score_addon_gemma.py
23
+ 7107cfe4a47f306829ee4e30fa50ddeecc9f18f107c5eb300657877b3ce9aded eval/btzsc22/addon_gemma/per_dataset.csv
24
+ 2eca0b816e2ba974b31fd25043ab7ea38b39306f5ed97b16091a6905098a4a2f eval/btzsc22/addon_gemma/results.json
25
+ 5e853fc07f7ab18b695be6b356f8d2c1e6ae23a61e3d79fed0f3be310fe3e26f eval/btzsc22/addon_gemma/trackA_native.gemma.jsonl
26
  95a4d4c459940195f03c6b937d750cd8e678527c4e2acaee1f7f2f3a9adc512c eval/btzsc22/btzsc_test_files.json
27
  bb7dbba94cc130a09c77226e51cb7b72e5228685fbf8c2f78522290f485eff1e eval/btzsc22/code/README.md
28
  afac77fa429f2f37c45e6cf46d0201551e9c98ddf9c01e7ba0cf1ac1feca0ef7 eval/btzsc22/code/btzsc22.yaml
 
64
  d0b5cb709b104e8fafcdba8f8c676896ce048c2a009f93c1b19ad3e949fcb128 eval/reference_rows.json
65
  508a3e33bf77cd4ab9f559d5dcd2ec3e473c30b3b1d54c90a7de5e6490a435a2 eval/report.json
66
  d5f3f5d214ac20264176a552c7779ddbbee79f97524b64f8c6d1804e1f305ba2 eval/report.md
67
+ 2845d32914cc266a9b8a14f20b9ea7b37ad6febf2ae52a084c46ba51af140165 eval/speed.md
68
  d0b6fcd59367eb58c02eb3b21e7d6c50caa6eccf5d822e8e7ecec22f0e3a275b eval/speed_raw.log
69
+ da6b8271d68d79a13d4e79e5c5e90e6722d91754bf15e797670a43a6120e99d0 eval/speed_runs/2026-10-06_part1_gemma_laya.txt
70
+ 34b9bd9bf06e21f67ad73fc13e8d0deb49efa95a9f181d820b6eced9aa133d6d eval/speed_runs/2026-10-06_part2_gliner_ines.txt
71
  22974cc87fd92b5c21168e350ff3db492209a98293ec28a7f5d70396f13c970d examples/quickstart.py
72
  6c4fc03dc8ec311323338f6f4f6c0a8a3514f417b57c99f70827d797fe2ead28 mini_v41/__init__.py
73
  7c57633d0d69fe8a709e0629cbe36b07590d94d0333039c5378fa66df4a74f05 mini_v41/attention.py
eval/btzsc22/SHA256SUMS CHANGED
@@ -22,3 +22,9 @@ d0ca40988b8ad697223f7ccfc18ac35840c1390a5d46fabc219dab0c02069ed3 ./predictions/
22
  8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b ./PROTOCOL.md
23
  5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 ./REPORT.md
24
  707f23cd6b4ae30a38573d4f29ce852a82878b23ee062c61a8025d4bae1360ff ./results.json
 
 
 
 
 
 
 
22
  8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b ./PROTOCOL.md
23
  5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 ./REPORT.md
24
  707f23cd6b4ae30a38573d4f29ce852a82878b23ee062c61a8025d4bae1360ff ./results.json
25
+ cc8ce27ca26efcdc9985323e75fdbad20b602d69b1602c2883a1c80d37dbe15e ./addon_gemma/README.md
26
+ 80bc6a7b10e75160f88265611a7be4fc07d625980a55697cee3ff3a60dfc5632 ./addon_gemma/code/run_gemma.py
27
+ d4a2cb890861b2076e79934b5a1dd0f184c5b4418dec4916c8288def58a279e9 ./addon_gemma/code/score_addon_gemma.py
28
+ 7107cfe4a47f306829ee4e30fa50ddeecc9f18f107c5eb300657877b3ce9aded ./addon_gemma/per_dataset.csv
29
+ 2eca0b816e2ba974b31fd25043ab7ea38b39306f5ed97b16091a6905098a4a2f ./addon_gemma/results.json
30
+ 5e853fc07f7ab18b695be6b356f8d2c1e6ae23a61e3d79fed0f3be310fe3e26f ./addon_gemma/trackA_native.gemma.jsonl
eval/btzsc22/addon_gemma/README.md ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # BTZSC-22 add-on: DiffusionGemma-26B-A4B-it (2026-10-06)
2
+
3
+ **Added after the `btzsc22-ext-v1` freeze.** The protocol, the manifest (2,200 examples) and the scorer are unchanged; this
4
+ directory adds one more model on **Track A (native) only**. Track B (one-vs-rest) was not run for it.
5
+
6
+ | | |
7
+ |---|---|
8
+ | model | `RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic` (FP8 of Google's DiffusionGemma-26B-A4B-it) |
9
+ | serving | `razorback16/openjev:0.4.0` (vLLM structured reads, `/v1/systemone`), model `openjev-0.1`, `think: 0` (the server default), canvas 64, 1× H200 |
10
+ | request | identical to the other models: `state = {"text": <text>}`, one Choice question `"Which single label best describes the input text?"`, options `label_000…` with the label texts as descriptions |
11
+ | interface limit | 255 options, so all 22 datasets are a **single request** (no knockout) |
12
+ | failures | 0 of 2,200 |
13
+
14
+ ## Files
15
+
16
+ - `trackA_native.gemma.jsonl`: one prediction per example (full native probability vector). `latency_seconds` was
17
+ recorded under 16 concurrent clients and is **not** a single-request latency (see `../../speed.md`).
18
+ - `code/run_gemma.py`: the client used.
19
+ - `code/score_addon_gemma.py`: `../code/score_final.py` with Gemma added to the model list. The Track A logic,
20
+ the paired hierarchical bootstrap and the calibration code are copied verbatim. Rerunning it reproduces the four
21
+ original models' numbers in `../results.json` exactly (0.474 / 0.525 / 0.530 / 0.537).
22
+ - `results.json`, `per_dataset.csv`: the output. Gemma's Track B fields are `null` / `not_run`.
23
+
24
+ ## Result (Track A, equal-weight mean macro-F1 over datasets)
25
+
26
+ | group | Ines-1 | Julia-1 | laya-multilingual | GLiNER2.5-multi | DiffusionGemma-26B-A4B |
27
+ |---|---:|---:|---:|---:|---:|
28
+ | all 22 | 0.474 | 0.525 | 0.530 | 0.537 | **0.799** |
29
+ | 2 labels (13) | 0.562 | 0.583 | 0.722 | 0.627 | **0.897** |
30
+ | 3–20 labels (4) | 0.489 | 0.634 | 0.559 | 0.541 | **0.695** |
31
+ | 21–72 labels (5) | 0.234 | 0.287 | 0.008 | 0.297 | **0.628** |
32
+ | absent from Ines-1's decision FT mix (15) | 0.486 | 0.536 | 0.515 | 0.528 | **0.765** |
33
+
34
+ Gemma − Ines-1, all 22: **+0.326** (95 % CI +0.253 to +0.392, paired hierarchical bootstrap).
35
+ Native calibration (17 datasets with ≤ 20 labels): ECE 0.116 (Ines-1 0.127), Brier 0.250 (Ines-1 0.501),
36
+ coverage at 5 % error budget 0.694 (Ines-1 0.089).
37
+
38
+ DiffusionGemma-26B-A4B has ~26B total parameters (~4B active per token), about 16× Ines-1's total and 10× its active
39
+ parameters. Speed and memory for the same models are in `../../speed.md` (run 3).
eval/btzsc22/addon_gemma/code/run_gemma.py ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """BTZSC-22 backend for DiffusionGemma-26B-A4B-it (FP8) served by OpenJev 0.4.0 (/v1/systemone).
2
+ Same request as run_model.py (state={"text": text}, one Choice, options label_000...), think 0 (OpenJev default),
3
+ native mode only (OpenJev allows 255 options, every dataset fits). argv: url manifest out.jsonl
4
+ """
5
+ import json, os, sys, time, urllib.request
6
+ from concurrent.futures import ThreadPoolExecutor
7
+ URL, MAN, OUT = sys.argv[1:4]
8
+ Q = "Which single label best describes the input text?"
9
+
10
+ def call(text, labels):
11
+ crit = {"label_%03d" % i: l for i, l in enumerate(labels)}
12
+ body = {"model": "openjev-0.1", "state": {"text": text}, "questions": {"label": {"type": "choice", "instructions": Q, "criteria": crit}}}
13
+ req = urllib.request.Request(URL, json.dumps(body).encode(), {"Content-Type": "application/json"})
14
+ with urllib.request.urlopen(req, timeout=300) as r:
15
+ a = json.loads(r.read())["answers"]["label"]
16
+ return [float(a["probabilities"][k]) for k in crit]
17
+
18
+ def one(e):
19
+ t = time.perf_counter()
20
+ try:
21
+ p = call(e["text"], e["labels"]); pred = max(range(len(p)), key=p.__getitem__); err = None
22
+ except Exception as x: # failures stay in the denominator
23
+ p, pred, err = None, -1, repr(x)[:300]
24
+ return {"backend": "gemma", "mode": "native", "group": 0, "dataset": e["dataset"], "example_id": e["example_id"],
25
+ "target_index": e["target_index"], "predicted_index": pred, "n_labels": len(e["labels"]), "probabilities": p,
26
+ "calls": 1, "latency_seconds": time.perf_counter() - t, "error": err}
27
+
28
+ done = set()
29
+ if os.path.exists(OUT):
30
+ done = {json.loads(l)["example_id"] for l in open(OUT)}
31
+ ex = [e for e in map(json.loads, open(MAN)) if e["example_id"] not in done]
32
+ call(ex[0]["text"], ex[0]["labels"][:2]) # warm-up
33
+ with open(OUT, "a") as f, ThreadPoolExecutor(16) as pool:
34
+ for r in pool.map(one, ex):
35
+ f.write(json.dumps(r) + "\n"); f.flush()
36
+ print("done gemma native")
eval/btzsc22/addon_gemma/code/score_addon_gemma.py ADDED
@@ -0,0 +1,123 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """ADD-ON (6-oct, after the freeze): score_final.py + DiffusionGemma-26B-A4B-it (OpenJev) on Track A only. Track A logic, bootstrap and calibration copied verbatim; Track B not run for Gemma.
2
+ Original: btzsc22-ext-v1 analysis (frozen before reading Track B). argv: run_dir(v1) exploratory_dir out_dir
3
+ Run with the gliner venv (uses jev_benchmarks.metrics unchanged)."""
4
+ import csv, json, sys
5
+ from pathlib import Path
6
+ import numpy as np
7
+ from jev_benchmarks.metrics import score_predictions, _macro_f1
8
+ from jev_benchmarks.models import Prediction
9
+
10
+ V1, EXP, OUT = Path(sys.argv[1]), Path(sys.argv[2]), Path(sys.argv[3]); OUT.mkdir(parents=True, exist_ok=True)
11
+ MODELS = ["ines", "julia", "laya", "gliner", "gemma"]
12
+ GEMMA = Path(__file__).resolve().parent / "runs/btzsc22/addon_gemma/gemma.native.jsonl"
13
+ NAME = {"ines": "Ines-1", "julia": "Julia-1", "laya": "laya-multilingual", "gliner": "GLiNER2.5", "gemma": "DiffusionGemma-26B-A4B-it"}
14
+ IN_MIX = {"amazonpolarity", "banking77", "massive", "wikitoxic_insult", "wikitoxic_obscene", "wikitoxic_threat", "wikitoxic_toxicaggregated"}
15
+ man = [json.loads(l) for l in open(EXP / "manifest.jsonl")]
16
+ DS = list(dict.fromkeys(e["dataset"] for e in man)); K = {e["dataset"]: len(e["labels"]) for e in man}
17
+ IDS = {d: [e["example_id"] for e in man if e["dataset"] == d] for d in DS}
18
+ T = {d: np.array([e["target_index"] for e in man if e["dataset"] == d]) for d in DS}
19
+ LAB = {e["example_id"]: tuple(e["labels"]) for e in man}
20
+ rd = lambda p: {json.loads(l)["example_id"]: json.loads(l) for l in open(p)} if Path(p).exists() else {}
21
+ native = {m: rd(V1 / ("%s.native.jsonl" % m)) for m in MODELS if m != "gemma"}; native["gemma"] = rd(GEMMA)
22
+ ovr = {m: rd(V1 / ("%s.ovr.jsonl" % m)) for m in MODELS}
23
+ ines_k20, ines_k2 = rd(EXP / "ines.k20.jsonl"), rd(EXP / "ines.k2.jsonl")
24
+
25
+ def preds(src, d):
26
+ return np.array([(src[i]["predicted_index"] if i in src and not src[i].get("error") else -1) for i in IDS[d]])
27
+ def trackA(m, d):
28
+ if m == "ines" and K[d] > 26:
29
+ return preds(ines_k20, d), "knockout-20 (no native > 26 interface)"
30
+ r = native[m]; s = "single request"
31
+ if m == "julia" and K[d] > 20: s = "Router(width=20, survivors=2)"
32
+ return preds(r, d), s
33
+ P = {"A": {m: {d: trackA(m, d)[0] for d in DS} for m in MODELS}, "B": {m: {d: preds(ovr[m], d) for d in DS} for m in MODELS}}
34
+ missing = {tr: {m: int(sum((P[tr][m][d] < 0).sum() for d in DS)) for m in MODELS} for tr in P}
35
+ f1 = lambda tr, m, d, idx=None: _macro_f1(T[d] if idx is None else T[d][idx], P[tr][m][d] if idx is None else P[tr][m][d][idx], K[d])
36
+ acc = lambda tr, m, d: float(np.mean(P[tr][m][d] == T[d]))
37
+
38
+ GROUPS = {"all 22": DS, "2 labels (13)": [d for d in DS if K[d] == 2], "3-20 labels (4)": [d for d in DS if 2 < K[d] <= 20],
39
+ "21-72 labels (5)": [d for d in DS if K[d] > 20], "absent from Ines-1 decision FT mixture (15)": [d for d in DS if d not in IN_MIX],
40
+ "families in Ines-1 decision FT mixture (7)": [d for d in DS if d in IN_MIX]}
41
+ rng = np.random.default_rng(20260917)
42
+ STRATA = {d: [np.where(T[d] == c)[0] for c in np.unique(T[d])] for d in DS}
43
+ def hboot(sel, pairs, B=2000):
44
+ """paired hierarchical bootstrap: resample datasets, then examples stratified by target; equal-weight mean macro-F1"""
45
+ out = {p: [] for p in pairs}
46
+ for _ in range(B):
47
+ draw = rng.integers(0, len(sel), len(sel))
48
+ acc_ = {p: [] for p in pairs}
49
+ for j in draw:
50
+ d = sel[j]; idx = np.concatenate([rng.choice(c, len(c)) for c in STRATA[d]])
51
+ cache = {}
52
+ for p in pairs:
53
+ (ta, ma), (tb, mb) = p
54
+ for key in ((ta, ma), (tb, mb)):
55
+ if key not in cache: cache[key] = f1(key[0], key[1], d, idx)
56
+ acc_[p].append(cache[(tb, mb)] - cache[(ta, ma)])
57
+ for p in pairs: out[p].append(np.mean(acc_[p]))
58
+ res = {}
59
+ for p, v in out.items():
60
+ (ta, ma), (tb, mb) = p
61
+ obs = float(np.mean([f1(tb, mb, d) - f1(ta, ma, d) for d in sel]))
62
+ res["%s:%s - %s:%s" % (tb, mb, ta, ma)] = {"difference": obs, "ci95": [float(np.quantile(v, .025)), float(np.quantile(v, .975))]}
63
+ return res
64
+
65
+ R = {"protocol": "btzsc22-ext-v1", "missing_predictions": missing, "per_dataset": {}, "global": {}, "bootstrap": {}, "calibration": {},
66
+ "ines_ablation": {}, "anchors": {}, "ovr_ties": {}}
67
+ for d in DS:
68
+ R["per_dataset"][d] = {"labels": K[d], "in_ines_ft_mixture": d in IN_MIX,
69
+ **{tr: {m: {"accuracy": acc(tr, m, d), "macro_f1": f1(tr, m, d)} for m in MODELS} for tr in ("A", "B")},
70
+ "trackA_strategy": {m: trackA(m, d)[1] for m in MODELS}}
71
+ for g, sel in GROUPS.items():
72
+ R["global"][g] = {tr: {m: float(np.mean([f1(tr, m, d) for d in sel])) for m in MODELS} for tr in ("A", "B")}
73
+ cross = [(("A", "ines"), ("A", m)) for m in MODELS[1:]] + [(("B", "ines"), ("B", m)) for m in MODELS[1:]] + [(("A", m), ("B", m)) for m in MODELS]
74
+ for g, sel in GROUPS.items():
75
+ R["bootstrap"][g] = hboot(sel, cross, B=2000 if g == "all 22" else 1000)
76
+
77
+ def calib(srcs, field, sel):
78
+ out = {}
79
+ for m in MODELS:
80
+ per = []
81
+ for d in sel:
82
+ rows = []
83
+ for i in IDS[d]:
84
+ r = srcs[m].get(i)
85
+ if not r or r.get("error") or not r.get(field): break
86
+ rows.append(Prediction(experiment_id="x", backend=m, model_requested=m, model_resolved=m, dataset=d, example_id=i,
87
+ target_index=int(T[d][IDS[d].index(i)]), predicted_index=r["predicted_index"], labels=LAB[i],
88
+ probabilities=tuple(r[field]), latency_seconds=0.0))
89
+ else:
90
+ per.append(score_predictions(rows))
91
+ if per:
92
+ out[m] = {k: float(np.mean([s[k] for s in per])) for k in ("brier", "nll", "ece", "coverage_at_error_budget", "mean_confidence")} | {"datasets": len(per)}
93
+ return out
94
+ native_full = [d for d in DS if K[d] <= 20]
95
+ R["calibration"]["native (Track A, full native distributions, <=20 labels, 17 datasets)"] = calib(native, "probabilities", native_full)
96
+ R["calibration"]["derived (Track B softmax of log-odds, same 17 datasets)"] = calib(ovr, "derived_probabilities", native_full)
97
+ R["calibration"]["derived (Track B softmax of log-odds, all 22 datasets)"] = calib(ovr, "derived_probabilities", DS)
98
+ R["ovr_ties"] = {m: {"examples_with_tie_at_max": sum(1 for r in ovr[m].values() if r.get("ties_at_max", 1) > 1), "n": len(ovr[m])} for m in MODELS}
99
+
100
+ multi = [d for d in DS if K[d] > 2]
101
+ for d in multi:
102
+ row = {"native(<=26)": acc("A", "ines", d) if K[d] <= 26 else None, "knockout-20": float(np.mean(preds(ines_k20, d) == T[d])) if K[d] > 20 else None,
103
+ "knockout-2": float(np.mean(preds(ines_k2, d) == T[d])), "one-vs-rest": acc("B", "ines", d)}
104
+ rowf = {"native(<=26)": _macro_f1(T[d], preds(native["ines"], d), K[d]) if K[d] <= 26 else None,
105
+ "knockout-20": _macro_f1(T[d], preds(ines_k20, d), K[d]) if K[d] > 20 else None,
106
+ "knockout-2": _macro_f1(T[d], preds(ines_k2, d), K[d]), "one-vs-rest": f1("B", "ines", d)}
107
+ R["ines_ablation"][d] = {"labels": K[d], "accuracy": row, "macro_f1": rowf}
108
+
109
+ pub = {"agnews": 0.70, "emotiondair": 0.44, "banking77": 0.61}
110
+ R["anchors"]["GLiNER2.5 Track A vs published pilot (accuracy)"] = {d: {"ours": acc("A", "gliner", d), "published": v} for d, v in pub.items()}
111
+ card = {"agnews": 0.94, "emotiondair": 0.86, "banking77": 0.64}
112
+ R["anchors"]["Julia-1 Track A vs its model card (accuracy)"] = {d: {"ours": acc("A", "julia", d), "card": v, "ours_strategy": trackA("julia", d)[1]} for d, v in card.items()}
113
+
114
+ json.dump(R, open(OUT / "results.json", "w"), indent=1)
115
+ with open(OUT / "per_dataset.csv", "w", newline="") as f:
116
+ w = csv.writer(f); w.writerow(["dataset", "labels", "in_ines1_decision_ft_mixture"] + ["%s_%s_%s" % (tr, m, k) for tr in ("A", "B") for m in MODELS for k in ("accuracy", "macro_f1")] + ["A_strategy_ines", "A_strategy_julia"])
117
+ for d in DS:
118
+ x = R["per_dataset"][d]
119
+ w.writerow([d, K[d], d in IN_MIX] + ["%.4f" % x[tr][m][k] for tr in ("A", "B") for m in MODELS for k in ("accuracy", "macro_f1")] + [x["trackA_strategy"]["ines"], x["trackA_strategy"]["julia"]])
120
+ print("missing", missing)
121
+ for g in GROUPS:
122
+ print("%-45s A %s | B %s" % (g, " ".join("%s %.3f" % (m, R["global"][g]["A"][m]) for m in MODELS), " ".join("%s %.3f" % (m, R["global"][g]["B"][m]) for m in MODELS)))
123
+ for k, v in R["bootstrap"]["all 22"].items(): print(" %-28s %+.3f [%+.3f, %+.3f]" % (k, v["difference"], *v["ci95"]))
eval/btzsc22/addon_gemma/per_dataset.csv ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ dataset,labels,in_ines1_decision_ft_mixture,A_ines_accuracy,A_ines_macro_f1,A_julia_accuracy,A_julia_macro_f1,A_laya_accuracy,A_laya_macro_f1,A_gliner_accuracy,A_gliner_macro_f1,A_gemma_accuracy,A_gemma_macro_f1,B_ines_accuracy,B_ines_macro_f1,B_julia_accuracy,B_julia_macro_f1,B_laya_accuracy,B_laya_macro_f1,B_gliner_accuracy,B_gliner_macro_f1,B_gemma_accuracy,B_gemma_macro_f1,A_strategy_ines,A_strategy_julia
2
+ agnews,4,False,0.5900,0.5843,0.9400,0.9399,0.9400,0.9393,0.7000,0.6589,0.8900,0.8797,0.4200,0.3265,0.4200,0.3995,0.8900,0.8890,0.4500,0.3986,not_run,not_run,single request,single request
3
+ emotiondair,6,False,0.4300,0.3904,0.8600,0.8557,0.3700,0.3527,0.4400,0.4070,0.4000,0.3764,0.3400,0.2948,0.2600,0.2543,0.4300,0.4279,0.3500,0.3193,not_run,not_run,single request,single request
4
+ banking77,72,True,0.3900,0.3427,0.6300,0.5704,0.0200,0.0006,0.6100,0.5693,0.8000,0.7773,0.2900,0.2235,0.0900,0.0472,0.2200,0.1880,0.4100,0.3788,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
5
+ amazonpolarity,2,True,0.7800,0.7756,0.6200,0.6162,0.8400,0.8368,0.8800,0.8800,0.9800,0.9800,0.6500,0.6179,0.5800,0.5716,0.9100,0.9098,0.8300,0.8249,not_run,not_run,single request,single request
6
+ imdb,2,False,0.7500,0.7500,0.6800,0.6799,0.7300,0.7149,0.8000,0.7947,0.9700,0.9700,0.6900,0.6808,0.5400,0.5354,0.8100,0.8091,0.7300,0.7175,not_run,not_run,single request,single request
7
+ appreviews,2,False,0.7700,0.7660,0.6500,0.6396,0.7800,0.7756,0.8900,0.8895,0.9400,0.9399,0.8100,0.8095,0.5100,0.4954,0.8900,0.8897,0.8000,0.7960,not_run,not_run,single request,single request
8
+ yelpreviews,2,False,0.7200,0.6997,0.7300,0.7278,0.9000,0.8996,0.9000,0.8994,0.9600,0.9599,0.6800,0.6435,0.5100,0.5088,0.9300,0.9298,0.7400,0.7211,not_run,not_run,single request,single request
9
+ rottentomatoes,2,False,0.5900,0.5671,0.6600,0.6264,0.6900,0.6808,0.7100,0.7064,0.9100,0.9100,0.6400,0.6347,0.4600,0.4325,0.7400,0.7400,0.6700,0.6548,not_run,not_run,single request,single request
10
+ financialphrasebank,3,False,0.5100,0.4891,0.4500,0.3711,0.8100,0.8085,0.6900,0.6821,0.8000,0.8030,0.4500,0.3886,0.2500,0.1744,0.9100,0.9097,0.5500,0.4785,not_run,not_run,single request,single request
11
+ empathetic,32,False,0.2500,0.1909,0.1700,0.1594,0.0600,0.0255,0.3300,0.2807,0.5100,0.4697,0.1800,0.1442,0.0500,0.0390,0.3300,0.2873,0.1400,0.1304,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
12
+ biasframes_intent,2,False,0.6000,0.5998,0.5700,0.5665,0.6100,0.5920,0.5800,0.5739,0.8300,0.8300,0.5900,0.5524,0.5900,0.5806,0.6600,0.6550,0.5200,0.3912,not_run,not_run,single request,single request
13
+ massive,59,True,0.3500,0.2502,0.2700,0.1802,0.0100,0.0010,0.4400,0.3827,0.7800,0.7266,0.2000,0.1528,0.0200,0.0122,0.3500,0.3255,0.3000,0.2354,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
14
+ yahootopics,10,False,0.5000,0.4915,0.3900,0.3675,0.1500,0.1369,0.4100,0.4170,0.7300,0.7218,0.1800,0.1231,0.1200,0.1127,0.3700,0.3650,0.3200,0.3175,not_run,not_run,single request,single request
15
+ trueteacher,2,False,0.5600,0.5376,0.5200,0.5152,0.5600,0.4658,0.5300,0.5242,0.7100,0.6966,0.5000,0.3333,0.5000,0.4346,0.4800,0.4207,0.5400,0.4524,not_run,not_run,single request,single request
16
+ manifesto,56,False,0.1100,0.0750,0.1000,0.0789,0.0300,0.0077,0.0200,0.0007,0.4400,0.3767,0.0500,0.0329,0.0200,0.0149,0.1100,0.0903,0.0600,0.0136,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
17
+ capsotu,21,False,0.3600,0.3121,0.4500,0.4452,0.0500,0.0050,0.3000,0.2524,0.7900,0.7913,0.2100,0.1784,0.1000,0.0799,0.2700,0.2800,0.1700,0.1837,not_run,not_run,single request,"Router(width=20, survivors=2)"
18
+ biasframes_offensive,2,False,0.5200,0.4956,0.5400,0.5119,0.6200,0.6194,0.4600,0.4375,0.8200,0.8199,0.4800,0.4800,0.4500,0.4068,0.6900,0.6885,0.5200,0.3763,not_run,not_run,single request,single request
19
+ biasframes_sex,2,False,0.3400,0.3389,0.5600,0.5556,0.7200,0.7083,0.5300,0.3967,0.9300,0.9298,0.5300,0.5192,0.5600,0.5098,0.7900,0.7900,0.5600,0.4762,not_run,not_run,single request,single request
20
+ wikitoxic_insult,2,True,0.7700,0.7694,0.5600,0.5593,0.8500,0.8482,0.7000,0.6900,0.9300,0.9300,0.6900,0.6900,0.5400,0.5308,0.8000,0.7960,0.5800,0.5000,not_run,not_run,single request,single request
21
+ wikitoxic_obscene,2,True,0.4400,0.3455,0.5600,0.5572,0.6500,0.6419,0.5200,0.3912,0.8800,0.8798,0.7300,0.7300,0.6300,0.6190,0.8300,0.8279,0.6200,0.5703,not_run,not_run,single request,single request
22
+ wikitoxic_threat,2,True,0.4700,0.3498,0.4700,0.4674,0.9100,0.9100,0.5300,0.3967,0.9700,0.9700,0.7300,0.7254,0.5500,0.5423,0.7700,0.7614,0.5200,0.3763,not_run,not_run,single request,single request
23
+ wikitoxic_toxicaggregated,2,True,0.3300,0.3049,0.5700,0.5502,0.7000,0.6921,0.6200,0.5766,0.8500,0.8500,0.7700,0.7681,0.5700,0.5572,0.8700,0.8678,0.6200,0.5766,not_run,not_run,single request,single request
eval/btzsc22/addon_gemma/results.json ADDED
@@ -0,0 +1,2044 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "protocol": "btzsc22-ext-v1",
3
+ "missing_predictions": {
4
+ "A": {
5
+ "ines": 0,
6
+ "julia": 0,
7
+ "laya": 0,
8
+ "gliner": 0,
9
+ "gemma": 0
10
+ },
11
+ "B": {
12
+ "ines": 0,
13
+ "julia": 0,
14
+ "laya": 0,
15
+ "gliner": 0,
16
+ "gemma": "not run (Track B was not run for Gemma)"
17
+ }
18
+ },
19
+ "per_dataset": {
20
+ "agnews": {
21
+ "labels": 4,
22
+ "in_ines_ft_mixture": false,
23
+ "A": {
24
+ "ines": {
25
+ "accuracy": 0.59,
26
+ "macro_f1": 0.5843451916892106
27
+ },
28
+ "julia": {
29
+ "accuracy": 0.94,
30
+ "macro_f1": 0.9399173646545361
31
+ },
32
+ "laya": {
33
+ "accuracy": 0.94,
34
+ "macro_f1": 0.939271525304066
35
+ },
36
+ "gliner": {
37
+ "accuracy": 0.7,
38
+ "macro_f1": 0.6588522320524076
39
+ },
40
+ "gemma": {
41
+ "accuracy": 0.89,
42
+ "macro_f1": 0.8797325209206743
43
+ }
44
+ },
45
+ "B": {
46
+ "ines": {
47
+ "accuracy": 0.42,
48
+ "macro_f1": 0.32652938961350175
49
+ },
50
+ "julia": {
51
+ "accuracy": 0.42,
52
+ "macro_f1": 0.399474348626891
53
+ },
54
+ "laya": {
55
+ "accuracy": 0.89,
56
+ "macro_f1": 0.8889659795261577
57
+ },
58
+ "gliner": {
59
+ "accuracy": 0.45,
60
+ "macro_f1": 0.39861358684888093
61
+ },
62
+ "gemma": null
63
+ },
64
+ "trackA_strategy": {
65
+ "ines": "single request",
66
+ "julia": "single request",
67
+ "laya": "single request",
68
+ "gliner": "single request",
69
+ "gemma": "single request"
70
+ }
71
+ },
72
+ "emotiondair": {
73
+ "labels": 6,
74
+ "in_ines_ft_mixture": false,
75
+ "A": {
76
+ "ines": {
77
+ "accuracy": 0.43,
78
+ "macro_f1": 0.39042230395613853
79
+ },
80
+ "julia": {
81
+ "accuracy": 0.86,
82
+ "macro_f1": 0.8556934538891329
83
+ },
84
+ "laya": {
85
+ "accuracy": 0.37,
86
+ "macro_f1": 0.35268000920174836
87
+ },
88
+ "gliner": {
89
+ "accuracy": 0.44,
90
+ "macro_f1": 0.4070167592505414
91
+ },
92
+ "gemma": {
93
+ "accuracy": 0.4,
94
+ "macro_f1": 0.37637917637917634
95
+ }
96
+ },
97
+ "B": {
98
+ "ines": {
99
+ "accuracy": 0.34,
100
+ "macro_f1": 0.29478472669962036
101
+ },
102
+ "julia": {
103
+ "accuracy": 0.26,
104
+ "macro_f1": 0.25425384800384804
105
+ },
106
+ "laya": {
107
+ "accuracy": 0.43,
108
+ "macro_f1": 0.4279340742423987
109
+ },
110
+ "gliner": {
111
+ "accuracy": 0.35,
112
+ "macro_f1": 0.3193197964297709
113
+ },
114
+ "gemma": null
115
+ },
116
+ "trackA_strategy": {
117
+ "ines": "single request",
118
+ "julia": "single request",
119
+ "laya": "single request",
120
+ "gliner": "single request",
121
+ "gemma": "single request"
122
+ }
123
+ },
124
+ "banking77": {
125
+ "labels": 72,
126
+ "in_ines_ft_mixture": true,
127
+ "A": {
128
+ "ines": {
129
+ "accuracy": 0.39,
130
+ "macro_f1": 0.34265873015873016
131
+ },
132
+ "julia": {
133
+ "accuracy": 0.63,
134
+ "macro_f1": 0.5703703703703703
135
+ },
136
+ "laya": {
137
+ "accuracy": 0.02,
138
+ "macro_f1": 0.0005611672278338945
139
+ },
140
+ "gliner": {
141
+ "accuracy": 0.61,
142
+ "macro_f1": 0.569268077601411
143
+ },
144
+ "gemma": {
145
+ "accuracy": 0.8,
146
+ "macro_f1": 0.7773148148148148
147
+ }
148
+ },
149
+ "B": {
150
+ "ines": {
151
+ "accuracy": 0.29,
152
+ "macro_f1": 0.2234898589065256
153
+ },
154
+ "julia": {
155
+ "accuracy": 0.09,
156
+ "macro_f1": 0.04716343327454439
157
+ },
158
+ "laya": {
159
+ "accuracy": 0.22,
160
+ "macro_f1": 0.18796296296296294
161
+ },
162
+ "gliner": {
163
+ "accuracy": 0.41,
164
+ "macro_f1": 0.3787818662818662
165
+ },
166
+ "gemma": null
167
+ },
168
+ "trackA_strategy": {
169
+ "ines": "knockout-20 (no native > 26 interface)",
170
+ "julia": "Router(width=20, survivors=2)",
171
+ "laya": "single request",
172
+ "gliner": "single request",
173
+ "gemma": "single request"
174
+ }
175
+ },
176
+ "amazonpolarity": {
177
+ "labels": 2,
178
+ "in_ines_ft_mixture": true,
179
+ "A": {
180
+ "ines": {
181
+ "accuracy": 0.78,
182
+ "macro_f1": 0.7756017951856384
183
+ },
184
+ "julia": {
185
+ "accuracy": 0.62,
186
+ "macro_f1": 0.6161616161616161
187
+ },
188
+ "laya": {
189
+ "accuracy": 0.84,
190
+ "macro_f1": 0.8368013055895553
191
+ },
192
+ "gliner": {
193
+ "accuracy": 0.88,
194
+ "macro_f1": 0.8799519807923168
195
+ },
196
+ "gemma": {
197
+ "accuracy": 0.98,
198
+ "macro_f1": 0.98
199
+ }
200
+ },
201
+ "B": {
202
+ "ines": {
203
+ "accuracy": 0.65,
204
+ "macro_f1": 0.6178622120318812
205
+ },
206
+ "julia": {
207
+ "accuracy": 0.58,
208
+ "macro_f1": 0.5716034271725826
209
+ },
210
+ "laya": {
211
+ "accuracy": 0.91,
212
+ "macro_f1": 0.9097744360902256
213
+ },
214
+ "gliner": {
215
+ "accuracy": 0.83,
216
+ "macro_f1": 0.8249407887962105
217
+ },
218
+ "gemma": null
219
+ },
220
+ "trackA_strategy": {
221
+ "ines": "single request",
222
+ "julia": "single request",
223
+ "laya": "single request",
224
+ "gliner": "single request",
225
+ "gemma": "single request"
226
+ }
227
+ },
228
+ "imdb": {
229
+ "labels": 2,
230
+ "in_ines_ft_mixture": false,
231
+ "A": {
232
+ "ines": {
233
+ "accuracy": 0.75,
234
+ "macro_f1": 0.74997499749975
235
+ },
236
+ "julia": {
237
+ "accuracy": 0.68,
238
+ "macro_f1": 0.6798719487795117
239
+ },
240
+ "laya": {
241
+ "accuracy": 0.73,
242
+ "macro_f1": 0.7149192271143491
243
+ },
244
+ "gliner": {
245
+ "accuracy": 0.8,
246
+ "macro_f1": 0.7947454844006567
247
+ },
248
+ "gemma": {
249
+ "accuracy": 0.97,
250
+ "macro_f1": 0.96999699969997
251
+ }
252
+ },
253
+ "B": {
254
+ "ines": {
255
+ "accuracy": 0.69,
256
+ "macro_f1": 0.6807743795695603
257
+ },
258
+ "julia": {
259
+ "accuracy": 0.54,
260
+ "macro_f1": 0.5353535353535354
261
+ },
262
+ "laya": {
263
+ "accuracy": 0.81,
264
+ "macro_f1": 0.8090644156366193
265
+ },
266
+ "gliner": {
267
+ "accuracy": 0.73,
268
+ "macro_f1": 0.7175436761167486
269
+ },
270
+ "gemma": null
271
+ },
272
+ "trackA_strategy": {
273
+ "ines": "single request",
274
+ "julia": "single request",
275
+ "laya": "single request",
276
+ "gliner": "single request",
277
+ "gemma": "single request"
278
+ }
279
+ },
280
+ "appreviews": {
281
+ "labels": 2,
282
+ "in_ines_ft_mixture": false,
283
+ "A": {
284
+ "ines": {
285
+ "accuracy": 0.77,
286
+ "macro_f1": 0.7660461804495982
287
+ },
288
+ "julia": {
289
+ "accuracy": 0.65,
290
+ "macro_f1": 0.6395839769333744
291
+ },
292
+ "laya": {
293
+ "accuracy": 0.78,
294
+ "macro_f1": 0.7756017951856384
295
+ },
296
+ "gliner": {
297
+ "accuracy": 0.89,
298
+ "macro_f1": 0.8894583458948849
299
+ },
300
+ "gemma": {
301
+ "accuracy": 0.94,
302
+ "macro_f1": 0.9399038461538461
303
+ }
304
+ },
305
+ "B": {
306
+ "ines": {
307
+ "accuracy": 0.81,
308
+ "macro_f1": 0.8095238095238095
309
+ },
310
+ "julia": {
311
+ "accuracy": 0.51,
312
+ "macro_f1": 0.49541756770672435
313
+ },
314
+ "laya": {
315
+ "accuracy": 0.89,
316
+ "macro_f1": 0.8897243107769424
317
+ },
318
+ "gliner": {
319
+ "accuracy": 0.8,
320
+ "macro_f1": 0.7960016319869441
321
+ },
322
+ "gemma": null
323
+ },
324
+ "trackA_strategy": {
325
+ "ines": "single request",
326
+ "julia": "single request",
327
+ "laya": "single request",
328
+ "gliner": "single request",
329
+ "gemma": "single request"
330
+ }
331
+ },
332
+ "yelpreviews": {
333
+ "labels": 2,
334
+ "in_ines_ft_mixture": false,
335
+ "A": {
336
+ "ines": {
337
+ "accuracy": 0.72,
338
+ "macro_f1": 0.6996996996996997
339
+ },
340
+ "julia": {
341
+ "accuracy": 0.73,
342
+ "macro_f1": 0.7277951406391774
343
+ },
344
+ "laya": {
345
+ "accuracy": 0.9,
346
+ "macro_f1": 0.8996386993175431
347
+ },
348
+ "gliner": {
349
+ "accuracy": 0.9,
350
+ "macro_f1": 0.8993558776167472
351
+ },
352
+ "gemma": {
353
+ "accuracy": 0.96,
354
+ "macro_f1": 0.9599358974358975
355
+ }
356
+ },
357
+ "B": {
358
+ "ines": {
359
+ "accuracy": 0.68,
360
+ "macro_f1": 0.64349376114082
361
+ },
362
+ "julia": {
363
+ "accuracy": 0.51,
364
+ "macro_f1": 0.5087719298245614
365
+ },
366
+ "laya": {
367
+ "accuracy": 0.93,
368
+ "macro_f1": 0.9298245614035088
369
+ },
370
+ "gliner": {
371
+ "accuracy": 0.74,
372
+ "macro_f1": 0.7211497211497211
373
+ },
374
+ "gemma": null
375
+ },
376
+ "trackA_strategy": {
377
+ "ines": "single request",
378
+ "julia": "single request",
379
+ "laya": "single request",
380
+ "gliner": "single request",
381
+ "gemma": "single request"
382
+ }
383
+ },
384
+ "rottentomatoes": {
385
+ "labels": 2,
386
+ "in_ines_ft_mixture": false,
387
+ "A": {
388
+ "ines": {
389
+ "accuracy": 0.59,
390
+ "macro_f1": 0.5670995670995671
391
+ },
392
+ "julia": {
393
+ "accuracy": 0.66,
394
+ "macro_f1": 0.6263736263736264
395
+ },
396
+ "laya": {
397
+ "accuracy": 0.69,
398
+ "macro_f1": 0.6807743795695603
399
+ },
400
+ "gliner": {
401
+ "accuracy": 0.71,
402
+ "macro_f1": 0.7064480210547626
403
+ },
404
+ "gemma": {
405
+ "accuracy": 0.91,
406
+ "macro_f1": 0.90999099909991
407
+ }
408
+ },
409
+ "B": {
410
+ "ines": {
411
+ "accuracy": 0.64,
412
+ "macro_f1": 0.6347402597402598
413
+ },
414
+ "julia": {
415
+ "accuracy": 0.46,
416
+ "macro_f1": 0.43253467843631777
417
+ },
418
+ "laya": {
419
+ "accuracy": 0.74,
420
+ "macro_f1": 0.74
421
+ },
422
+ "gliner": {
423
+ "accuracy": 0.67,
424
+ "macro_f1": 0.6547756041426928
425
+ },
426
+ "gemma": null
427
+ },
428
+ "trackA_strategy": {
429
+ "ines": "single request",
430
+ "julia": "single request",
431
+ "laya": "single request",
432
+ "gliner": "single request",
433
+ "gemma": "single request"
434
+ }
435
+ },
436
+ "financialphrasebank": {
437
+ "labels": 3,
438
+ "in_ines_ft_mixture": false,
439
+ "A": {
440
+ "ines": {
441
+ "accuracy": 0.51,
442
+ "macro_f1": 0.48911853494169083
443
+ },
444
+ "julia": {
445
+ "accuracy": 0.45,
446
+ "macro_f1": 0.37110209601081806
447
+ },
448
+ "laya": {
449
+ "accuracy": 0.81,
450
+ "macro_f1": 0.8084906388321863
451
+ },
452
+ "gliner": {
453
+ "accuracy": 0.69,
454
+ "macro_f1": 0.6821411017203141
455
+ },
456
+ "gemma": {
457
+ "accuracy": 0.8,
458
+ "macro_f1": 0.8029584462511292
459
+ }
460
+ },
461
+ "B": {
462
+ "ines": {
463
+ "accuracy": 0.45,
464
+ "macro_f1": 0.388558378810103
465
+ },
466
+ "julia": {
467
+ "accuracy": 0.25,
468
+ "macro_f1": 0.17441737846672992
469
+ },
470
+ "laya": {
471
+ "accuracy": 0.91,
472
+ "macro_f1": 0.9096611274030629
473
+ },
474
+ "gliner": {
475
+ "accuracy": 0.55,
476
+ "macro_f1": 0.47845237675746155
477
+ },
478
+ "gemma": null
479
+ },
480
+ "trackA_strategy": {
481
+ "ines": "single request",
482
+ "julia": "single request",
483
+ "laya": "single request",
484
+ "gliner": "single request",
485
+ "gemma": "single request"
486
+ }
487
+ },
488
+ "empathetic": {
489
+ "labels": 32,
490
+ "in_ines_ft_mixture": false,
491
+ "A": {
492
+ "ines": {
493
+ "accuracy": 0.25,
494
+ "macro_f1": 0.1909003551273288
495
+ },
496
+ "julia": {
497
+ "accuracy": 0.17,
498
+ "macro_f1": 0.15942460317460316
499
+ },
500
+ "laya": {
501
+ "accuracy": 0.06,
502
+ "macro_f1": 0.02546768707482993
503
+ },
504
+ "gliner": {
505
+ "accuracy": 0.33,
506
+ "macro_f1": 0.2807043650793651
507
+ },
508
+ "gemma": {
509
+ "accuracy": 0.51,
510
+ "macro_f1": 0.46972402597402596
511
+ }
512
+ },
513
+ "B": {
514
+ "ines": {
515
+ "accuracy": 0.18,
516
+ "macro_f1": 0.14416745535166586
517
+ },
518
+ "julia": {
519
+ "accuracy": 0.05,
520
+ "macro_f1": 0.039042207792207795
521
+ },
522
+ "laya": {
523
+ "accuracy": 0.33,
524
+ "macro_f1": 0.2872519841269841
525
+ },
526
+ "gliner": {
527
+ "accuracy": 0.14,
528
+ "macro_f1": 0.13043154761904763
529
+ },
530
+ "gemma": null
531
+ },
532
+ "trackA_strategy": {
533
+ "ines": "knockout-20 (no native > 26 interface)",
534
+ "julia": "Router(width=20, survivors=2)",
535
+ "laya": "single request",
536
+ "gliner": "single request",
537
+ "gemma": "single request"
538
+ }
539
+ },
540
+ "biasframes_intent": {
541
+ "labels": 2,
542
+ "in_ines_ft_mixture": false,
543
+ "A": {
544
+ "ines": {
545
+ "accuracy": 0.6,
546
+ "macro_f1": 0.5998399359743898
547
+ },
548
+ "julia": {
549
+ "accuracy": 0.57,
550
+ "macro_f1": 0.5664885573142454
551
+ },
552
+ "laya": {
553
+ "accuracy": 0.61,
554
+ "macro_f1": 0.5920075321686369
555
+ },
556
+ "gliner": {
557
+ "accuracy": 0.58,
558
+ "macro_f1": 0.5738636363636364
559
+ },
560
+ "gemma": {
561
+ "accuracy": 0.83,
562
+ "macro_f1": 0.8299829982998299
563
+ }
564
+ },
565
+ "B": {
566
+ "ines": {
567
+ "accuracy": 0.59,
568
+ "macro_f1": 0.5523528769516323
569
+ },
570
+ "julia": {
571
+ "accuracy": 0.59,
572
+ "macro_f1": 0.5805626598465473
573
+ },
574
+ "laya": {
575
+ "accuracy": 0.66,
576
+ "macro_f1": 0.6550324675324675
577
+ },
578
+ "gliner": {
579
+ "accuracy": 0.52,
580
+ "macro_f1": 0.39117199391172
581
+ },
582
+ "gemma": null
583
+ },
584
+ "trackA_strategy": {
585
+ "ines": "single request",
586
+ "julia": "single request",
587
+ "laya": "single request",
588
+ "gliner": "single request",
589
+ "gemma": "single request"
590
+ }
591
+ },
592
+ "massive": {
593
+ "labels": 59,
594
+ "in_ines_ft_mixture": true,
595
+ "A": {
596
+ "ines": {
597
+ "accuracy": 0.35,
598
+ "macro_f1": 0.2502433536331841
599
+ },
600
+ "julia": {
601
+ "accuracy": 0.27,
602
+ "macro_f1": 0.1801721818670971
603
+ },
604
+ "laya": {
605
+ "accuracy": 0.01,
606
+ "macro_f1": 0.0009685230024213075
607
+ },
608
+ "gliner": {
609
+ "accuracy": 0.44,
610
+ "macro_f1": 0.38272522334474407
611
+ },
612
+ "gemma": {
613
+ "accuracy": 0.78,
614
+ "macro_f1": 0.7265536723163842
615
+ }
616
+ },
617
+ "B": {
618
+ "ines": {
619
+ "accuracy": 0.2,
620
+ "macro_f1": 0.15282852740479855
621
+ },
622
+ "julia": {
623
+ "accuracy": 0.02,
624
+ "macro_f1": 0.01224105461393597
625
+ },
626
+ "laya": {
627
+ "accuracy": 0.35,
628
+ "macro_f1": 0.3254640839386602
629
+ },
630
+ "gliner": {
631
+ "accuracy": 0.3,
632
+ "macro_f1": 0.2353510895883777
633
+ },
634
+ "gemma": null
635
+ },
636
+ "trackA_strategy": {
637
+ "ines": "knockout-20 (no native > 26 interface)",
638
+ "julia": "Router(width=20, survivors=2)",
639
+ "laya": "single request",
640
+ "gliner": "single request",
641
+ "gemma": "single request"
642
+ }
643
+ },
644
+ "yahootopics": {
645
+ "labels": 10,
646
+ "in_ines_ft_mixture": false,
647
+ "A": {
648
+ "ines": {
649
+ "accuracy": 0.5,
650
+ "macro_f1": 0.4915120854091442
651
+ },
652
+ "julia": {
653
+ "accuracy": 0.39,
654
+ "macro_f1": 0.3675051008447593
655
+ },
656
+ "laya": {
657
+ "accuracy": 0.15,
658
+ "macro_f1": 0.13690074545337702
659
+ },
660
+ "gliner": {
661
+ "accuracy": 0.41,
662
+ "macro_f1": 0.4169864449276214
663
+ },
664
+ "gemma": {
665
+ "accuracy": 0.73,
666
+ "macro_f1": 0.7217792466130062
667
+ }
668
+ },
669
+ "B": {
670
+ "ines": {
671
+ "accuracy": 0.18,
672
+ "macro_f1": 0.12307189542483658
673
+ },
674
+ "julia": {
675
+ "accuracy": 0.12,
676
+ "macro_f1": 0.11272908345326083
677
+ },
678
+ "laya": {
679
+ "accuracy": 0.37,
680
+ "macro_f1": 0.36499120762278653
681
+ },
682
+ "gliner": {
683
+ "accuracy": 0.32,
684
+ "macro_f1": 0.31750090466808734
685
+ },
686
+ "gemma": null
687
+ },
688
+ "trackA_strategy": {
689
+ "ines": "single request",
690
+ "julia": "single request",
691
+ "laya": "single request",
692
+ "gliner": "single request",
693
+ "gemma": "single request"
694
+ }
695
+ },
696
+ "trueteacher": {
697
+ "labels": 2,
698
+ "in_ines_ft_mixture": false,
699
+ "A": {
700
+ "ines": {
701
+ "accuracy": 0.56,
702
+ "macro_f1": 0.537620849096259
703
+ },
704
+ "julia": {
705
+ "accuracy": 0.52,
706
+ "macro_f1": 0.5151515151515151
707
+ },
708
+ "laya": {
709
+ "accuracy": 0.56,
710
+ "macro_f1": 0.46576007770762506
711
+ },
712
+ "gliner": {
713
+ "accuracy": 0.53,
714
+ "macro_f1": 0.5242433444680635
715
+ },
716
+ "gemma": {
717
+ "accuracy": 0.71,
718
+ "macro_f1": 0.69662098545873
719
+ }
720
+ },
721
+ "B": {
722
+ "ines": {
723
+ "accuracy": 0.5,
724
+ "macro_f1": 0.3333333333333333
725
+ },
726
+ "julia": {
727
+ "accuracy": 0.5,
728
+ "macro_f1": 0.43464495703301675
729
+ },
730
+ "laya": {
731
+ "accuracy": 0.48,
732
+ "macro_f1": 0.4206773618538324
733
+ },
734
+ "gliner": {
735
+ "accuracy": 0.54,
736
+ "macro_f1": 0.45238095238095233
737
+ },
738
+ "gemma": null
739
+ },
740
+ "trackA_strategy": {
741
+ "ines": "single request",
742
+ "julia": "single request",
743
+ "laya": "single request",
744
+ "gliner": "single request",
745
+ "gemma": "single request"
746
+ }
747
+ },
748
+ "manifesto": {
749
+ "labels": 56,
750
+ "in_ines_ft_mixture": false,
751
+ "A": {
752
+ "ines": {
753
+ "accuracy": 0.11,
754
+ "macro_f1": 0.07504578754578754
755
+ },
756
+ "julia": {
757
+ "accuracy": 0.1,
758
+ "macro_f1": 0.0788690476190476
759
+ },
760
+ "laya": {
761
+ "accuracy": 0.03,
762
+ "macro_f1": 0.007711038961038961
763
+ },
764
+ "gliner": {
765
+ "accuracy": 0.02,
766
+ "macro_f1": 0.0007072135785007072
767
+ },
768
+ "gemma": {
769
+ "accuracy": 0.44,
770
+ "macro_f1": 0.37665816326530605
771
+ }
772
+ },
773
+ "B": {
774
+ "ines": {
775
+ "accuracy": 0.05,
776
+ "macro_f1": 0.032902735562310034
777
+ },
778
+ "julia": {
779
+ "accuracy": 0.02,
780
+ "macro_f1": 0.01488095238095238
781
+ },
782
+ "laya": {
783
+ "accuracy": 0.11,
784
+ "macro_f1": 0.09031507339778015
785
+ },
786
+ "gliner": {
787
+ "accuracy": 0.06,
788
+ "macro_f1": 0.013575543120473996
789
+ },
790
+ "gemma": null
791
+ },
792
+ "trackA_strategy": {
793
+ "ines": "knockout-20 (no native > 26 interface)",
794
+ "julia": "Router(width=20, survivors=2)",
795
+ "laya": "single request",
796
+ "gliner": "single request",
797
+ "gemma": "single request"
798
+ }
799
+ },
800
+ "capsotu": {
801
+ "labels": 21,
802
+ "in_ines_ft_mixture": false,
803
+ "A": {
804
+ "ines": {
805
+ "accuracy": 0.36,
806
+ "macro_f1": 0.3120540695728665
807
+ },
808
+ "julia": {
809
+ "accuracy": 0.45,
810
+ "macro_f1": 0.4451588094445238
811
+ },
812
+ "laya": {
813
+ "accuracy": 0.05,
814
+ "macro_f1": 0.005012531328320802
815
+ },
816
+ "gliner": {
817
+ "accuracy": 0.3,
818
+ "macro_f1": 0.2524050024050024
819
+ },
820
+ "gemma": {
821
+ "accuracy": 0.79,
822
+ "macro_f1": 0.7912510769653627
823
+ }
824
+ },
825
+ "B": {
826
+ "ines": {
827
+ "accuracy": 0.21,
828
+ "macro_f1": 0.17844985702128557
829
+ },
830
+ "julia": {
831
+ "accuracy": 0.1,
832
+ "macro_f1": 0.0798883656026513
833
+ },
834
+ "laya": {
835
+ "accuracy": 0.27,
836
+ "macro_f1": 0.2799847801788137
837
+ },
838
+ "gliner": {
839
+ "accuracy": 0.17,
840
+ "macro_f1": 0.18369192344347623
841
+ },
842
+ "gemma": null
843
+ },
844
+ "trackA_strategy": {
845
+ "ines": "single request",
846
+ "julia": "Router(width=20, survivors=2)",
847
+ "laya": "single request",
848
+ "gliner": "single request",
849
+ "gemma": "single request"
850
+ }
851
+ },
852
+ "biasframes_offensive": {
853
+ "labels": 2,
854
+ "in_ines_ft_mixture": false,
855
+ "A": {
856
+ "ines": {
857
+ "accuracy": 0.52,
858
+ "macro_f1": 0.49558638083228246
859
+ },
860
+ "julia": {
861
+ "accuracy": 0.54,
862
+ "macro_f1": 0.5118845500848896
863
+ },
864
+ "laya": {
865
+ "accuracy": 0.62,
866
+ "macro_f1": 0.6193910256410255
867
+ },
868
+ "gliner": {
869
+ "accuracy": 0.46,
870
+ "macro_f1": 0.4375
871
+ },
872
+ "gemma": {
873
+ "accuracy": 0.82,
874
+ "macro_f1": 0.8199279711884754
875
+ }
876
+ },
877
+ "B": {
878
+ "ines": {
879
+ "accuracy": 0.48,
880
+ "macro_f1": 0.48
881
+ },
882
+ "julia": {
883
+ "accuracy": 0.45,
884
+ "macro_f1": 0.40675223816201056
885
+ },
886
+ "laya": {
887
+ "accuracy": 0.69,
888
+ "macro_f1": 0.6884735202492211
889
+ },
890
+ "gliner": {
891
+ "accuracy": 0.52,
892
+ "macro_f1": 0.37629937629937626
893
+ },
894
+ "gemma": null
895
+ },
896
+ "trackA_strategy": {
897
+ "ines": "single request",
898
+ "julia": "single request",
899
+ "laya": "single request",
900
+ "gliner": "single request",
901
+ "gemma": "single request"
902
+ }
903
+ },
904
+ "biasframes_sex": {
905
+ "labels": 2,
906
+ "in_ines_ft_mixture": false,
907
+ "A": {
908
+ "ines": {
909
+ "accuracy": 0.34,
910
+ "macro_f1": 0.3389423076923077
911
+ },
912
+ "julia": {
913
+ "accuracy": 0.56,
914
+ "macro_f1": 0.5555555555555556
915
+ },
916
+ "laya": {
917
+ "accuracy": 0.72,
918
+ "macro_f1": 0.7083333333333334
919
+ },
920
+ "gliner": {
921
+ "accuracy": 0.53,
922
+ "macro_f1": 0.39673982800667434
923
+ },
924
+ "gemma": {
925
+ "accuracy": 0.93,
926
+ "macro_f1": 0.9298245614035088
927
+ }
928
+ },
929
+ "B": {
930
+ "ines": {
931
+ "accuracy": 0.53,
932
+ "macro_f1": 0.5191815856777494
933
+ },
934
+ "julia": {
935
+ "accuracy": 0.56,
936
+ "macro_f1": 0.5098039215686274
937
+ },
938
+ "laya": {
939
+ "accuracy": 0.79,
940
+ "macro_f1": 0.78997899789979
941
+ },
942
+ "gliner": {
943
+ "accuracy": 0.56,
944
+ "macro_f1": 0.47619047619047616
945
+ },
946
+ "gemma": null
947
+ },
948
+ "trackA_strategy": {
949
+ "ines": "single request",
950
+ "julia": "single request",
951
+ "laya": "single request",
952
+ "gliner": "single request",
953
+ "gemma": "single request"
954
+ }
955
+ },
956
+ "wikitoxic_insult": {
957
+ "labels": 2,
958
+ "in_ines_ft_mixture": true,
959
+ "A": {
960
+ "ines": {
961
+ "accuracy": 0.77,
962
+ "macro_f1": 0.7694235588972431
963
+ },
964
+ "julia": {
965
+ "accuracy": 0.56,
966
+ "macro_f1": 0.5592948717948718
967
+ },
968
+ "laya": {
969
+ "accuracy": 0.85,
970
+ "macro_f1": 0.8481627695110842
971
+ },
972
+ "gliner": {
973
+ "accuracy": 0.7,
974
+ "macro_f1": 0.6899545266639107
975
+ },
976
+ "gemma": {
977
+ "accuracy": 0.93,
978
+ "macro_f1": 0.9299929992999301
979
+ }
980
+ },
981
+ "B": {
982
+ "ines": {
983
+ "accuracy": 0.69,
984
+ "macro_f1": 0.6899689968996899
985
+ },
986
+ "julia": {
987
+ "accuracy": 0.54,
988
+ "macro_f1": 0.5308037535699714
989
+ },
990
+ "laya": {
991
+ "accuracy": 0.8,
992
+ "macro_f1": 0.7960016319869441
993
+ },
994
+ "gliner": {
995
+ "accuracy": 0.58,
996
+ "macro_f1": 0.5
997
+ },
998
+ "gemma": null
999
+ },
1000
+ "trackA_strategy": {
1001
+ "ines": "single request",
1002
+ "julia": "single request",
1003
+ "laya": "single request",
1004
+ "gliner": "single request",
1005
+ "gemma": "single request"
1006
+ }
1007
+ },
1008
+ "wikitoxic_obscene": {
1009
+ "labels": 2,
1010
+ "in_ines_ft_mixture": true,
1011
+ "A": {
1012
+ "ines": {
1013
+ "accuracy": 0.44,
1014
+ "macro_f1": 0.34548854604955587
1015
+ },
1016
+ "julia": {
1017
+ "accuracy": 0.56,
1018
+ "macro_f1": 0.5571658615136876
1019
+ },
1020
+ "laya": {
1021
+ "accuracy": 0.65,
1022
+ "macro_f1": 0.6419437340153453
1023
+ },
1024
+ "gliner": {
1025
+ "accuracy": 0.52,
1026
+ "macro_f1": 0.39117199391172
1027
+ },
1028
+ "gemma": {
1029
+ "accuracy": 0.88,
1030
+ "macro_f1": 0.8798076923076923
1031
+ }
1032
+ },
1033
+ "B": {
1034
+ "ines": {
1035
+ "accuracy": 0.73,
1036
+ "macro_f1": 0.72997299729973
1037
+ },
1038
+ "julia": {
1039
+ "accuracy": 0.63,
1040
+ "macro_f1": 0.6189887756152817
1041
+ },
1042
+ "laya": {
1043
+ "accuracy": 0.83,
1044
+ "macro_f1": 0.8279178054458953
1045
+ },
1046
+ "gliner": {
1047
+ "accuracy": 0.62,
1048
+ "macro_f1": 0.5703301673450927
1049
+ },
1050
+ "gemma": null
1051
+ },
1052
+ "trackA_strategy": {
1053
+ "ines": "single request",
1054
+ "julia": "single request",
1055
+ "laya": "single request",
1056
+ "gliner": "single request",
1057
+ "gemma": "single request"
1058
+ }
1059
+ },
1060
+ "wikitoxic_threat": {
1061
+ "labels": 2,
1062
+ "in_ines_ft_mixture": true,
1063
+ "A": {
1064
+ "ines": {
1065
+ "accuracy": 0.47,
1066
+ "macro_f1": 0.3497730339835603
1067
+ },
1068
+ "julia": {
1069
+ "accuracy": 0.47,
1070
+ "macro_f1": 0.467390212038991
1071
+ },
1072
+ "laya": {
1073
+ "accuracy": 0.91,
1074
+ "macro_f1": 0.90999099909991
1075
+ },
1076
+ "gliner": {
1077
+ "accuracy": 0.53,
1078
+ "macro_f1": 0.39673982800667434
1079
+ },
1080
+ "gemma": {
1081
+ "accuracy": 0.97,
1082
+ "macro_f1": 0.96999699969997
1083
+ }
1084
+ },
1085
+ "B": {
1086
+ "ines": {
1087
+ "accuracy": 0.73,
1088
+ "macro_f1": 0.7253585596582239
1089
+ },
1090
+ "julia": {
1091
+ "accuracy": 0.55,
1092
+ "macro_f1": 0.54226426609704
1093
+ },
1094
+ "laya": {
1095
+ "accuracy": 0.77,
1096
+ "macro_f1": 0.7613860358958398
1097
+ },
1098
+ "gliner": {
1099
+ "accuracy": 0.52,
1100
+ "macro_f1": 0.37629937629937626
1101
+ },
1102
+ "gemma": null
1103
+ },
1104
+ "trackA_strategy": {
1105
+ "ines": "single request",
1106
+ "julia": "single request",
1107
+ "laya": "single request",
1108
+ "gliner": "single request",
1109
+ "gemma": "single request"
1110
+ }
1111
+ },
1112
+ "wikitoxic_toxicaggregated": {
1113
+ "labels": 2,
1114
+ "in_ines_ft_mixture": true,
1115
+ "A": {
1116
+ "ines": {
1117
+ "accuracy": 0.33,
1118
+ "macro_f1": 0.30490714804440294
1119
+ },
1120
+ "julia": {
1121
+ "accuracy": 0.57,
1122
+ "macro_f1": 0.5501621508525996
1123
+ },
1124
+ "laya": {
1125
+ "accuracy": 0.7,
1126
+ "macro_f1": 0.6921182266009853
1127
+ },
1128
+ "gliner": {
1129
+ "accuracy": 0.62,
1130
+ "macro_f1": 0.5766488413547237
1131
+ },
1132
+ "gemma": {
1133
+ "accuracy": 0.85,
1134
+ "macro_f1": 0.84998499849985
1135
+ }
1136
+ },
1137
+ "B": {
1138
+ "ines": {
1139
+ "accuracy": 0.77,
1140
+ "macro_f1": 0.7681217864704104
1141
+ },
1142
+ "julia": {
1143
+ "accuracy": 0.57,
1144
+ "macro_f1": 0.557203171661003
1145
+ },
1146
+ "laya": {
1147
+ "accuracy": 0.87,
1148
+ "macro_f1": 0.8677652324280338
1149
+ },
1150
+ "gliner": {
1151
+ "accuracy": 0.62,
1152
+ "macro_f1": 0.5766488413547237
1153
+ },
1154
+ "gemma": null
1155
+ },
1156
+ "trackA_strategy": {
1157
+ "ines": "single request",
1158
+ "julia": "single request",
1159
+ "laya": "single request",
1160
+ "gliner": "single request",
1161
+ "gemma": "single request"
1162
+ }
1163
+ }
1164
+ },
1165
+ "global": {
1166
+ "all 22": {
1167
+ "A": {
1168
+ "ines": 0.4739229278426516,
1169
+ "julia": 0.5245951186849342,
1170
+ "laya": 0.5301139532382007,
1171
+ "gliner": 0.536710369477031,
1172
+ "gemma": 0.7994690041839767
1173
+ },
1174
+ "B": {
1175
+ "ines": 0.45679397195871585,
1176
+ "julia": 0.3572179797391928,
1177
+ "laya": 0.6294614568454058,
1178
+ "gliner": 0.44952051094233986,
1179
+ "gemma": null
1180
+ }
1181
+ },
1182
+ "2 labels (13)": {
1183
+ "A": {
1184
+ "ines": 0.561538769269558,
1185
+ "julia": 0.5825291987072047,
1186
+ "laya": 0.7219571619118916,
1187
+ "gliner": 0.627447823733444,
1188
+ "gemma": 0.8973820729652007
1189
+ },
1190
+ "B": {
1191
+ "ines": 0.6295911198690077,
1192
+ "julia": 0.5172849909267092,
1193
+ "laya": 0.7758169828614861,
1194
+ "gliner": 0.5718255850749258,
1195
+ "gemma": null
1196
+ }
1197
+ },
1198
+ "3-20 labels (4)": {
1199
+ "A": {
1200
+ "ines": 0.48884952899904605,
1201
+ "julia": 0.6335545038498116,
1202
+ "laya": 0.5593357296978444,
1203
+ "gliner": 0.5412491344877212,
1204
+ "gemma": 0.6952123475409966
1205
+ },
1206
+ "B": {
1207
+ "ines": 0.28323609763701546,
1208
+ "julia": 0.23521866463768246,
1209
+ "laya": 0.6478880971986014,
1210
+ "gliner": 0.3784716661760502,
1211
+ "gemma": null
1212
+ }
1213
+ },
1214
+ "21-72 labels (5)": {
1215
+ "A": {
1216
+ "ines": 0.23418045920757943,
1217
+ "julia": 0.28679900249512835,
1218
+ "laya": 0.00794418951888898,
1219
+ "gliner": 0.29716197640180464,
1220
+ "gemma": 0.6283003506671788
1221
+ },
1222
+ "B": {
1223
+ "ines": 0.14636768684931714,
1224
+ "julia": 0.03864320273285837,
1225
+ "laya": 0.23419577692104018,
1226
+ "gliner": 0.18836639401064836,
1227
+ "gemma": null
1228
+ }
1229
+ },
1230
+ "absent from Ines-1 decision FT mixture (15)": {
1231
+ "A": {
1232
+ "ines": 0.4858805497724014,
1233
+ "julia": 0.5360250230979545,
1234
+ "laya": 0.5154640164128853,
1235
+ "gliner": 0.5280778437879452,
1236
+ "gemma": 0.7649777943405899
1237
+ },
1238
+ "B": {
1239
+ "ines": 0.4094576296280325,
1240
+ "julia": 0.3319018448171922,
1241
+ "laya": 0.6114586574566909,
1242
+ "gliner": 0.4284732740710553,
1243
+ "gemma": null
1244
+ }
1245
+ },
1246
+ "families in Ines-1 decision FT mixture (7)": {
1247
+ "A": {
1248
+ "ines": 0.44829945227890217,
1249
+ "julia": 0.500102466371319,
1250
+ "laya": 0.5615066750067336,
1251
+ "gliner": 0.5552086388107859,
1252
+ "gemma": 0.8733787395626631
1253
+ },
1254
+ "B": {
1255
+ "ines": 0.5582289912387514,
1256
+ "julia": 0.411466840286337,
1257
+ "laya": 0.6680388841069373,
1258
+ "gliner": 0.4946217328093781,
1259
+ "gemma": null
1260
+ }
1261
+ }
1262
+ },
1263
+ "bootstrap": {
1264
+ "all 22": {
1265
+ "A:julia - A:ines": {
1266
+ "difference": 0.050672190842282465,
1267
+ "ci95": [
1268
+ -0.026197789339992814,
1269
+ 0.12574349835362167
1270
+ ]
1271
+ },
1272
+ "A:laya - A:ines": {
1273
+ "difference": 0.05619102539554903,
1274
+ "ci95": [
1275
+ -0.03957741468351192,
1276
+ 0.16332715413710902
1277
+ ]
1278
+ },
1279
+ "A:gliner - A:ines": {
1280
+ "difference": 0.06278744163437923,
1281
+ "ci95": [
1282
+ 0.014208626140706566,
1283
+ 0.10706083324209556
1284
+ ]
1285
+ },
1286
+ "A:gemma - A:ines": {
1287
+ "difference": 0.32554607634132515,
1288
+ "ci95": [
1289
+ 0.2531753321872848,
1290
+ 0.39198685847503467
1291
+ ]
1292
+ },
1293
+ "B:julia - B:ines": {
1294
+ "difference": -0.099575992219523,
1295
+ "ci95": [
1296
+ -0.14637550498902815,
1297
+ -0.05017137416536893
1298
+ ]
1299
+ },
1300
+ "B:laya - B:ines": {
1301
+ "difference": 0.17266748488669,
1302
+ "ci95": [
1303
+ 0.11284488099527132,
1304
+ 0.23744000375671462
1305
+ ]
1306
+ },
1307
+ "B:gliner - B:ines": {
1308
+ "difference": -0.007273461016375937,
1309
+ "ci95": [
1310
+ -0.06982555475121606,
1311
+ 0.05100715361862577
1312
+ ]
1313
+ },
1314
+ "B:ines - A:ines": {
1315
+ "difference": -0.017128955883935842,
1316
+ "ci95": [
1317
+ -0.09913170297625284,
1318
+ 0.07045906519324419
1319
+ ]
1320
+ },
1321
+ "B:julia - A:julia": {
1322
+ "difference": -0.16737713894574127,
1323
+ "ci95": [
1324
+ -0.2550239493204071,
1325
+ -0.08802480974576297
1326
+ ]
1327
+ },
1328
+ "B:laya - A:laya": {
1329
+ "difference": 0.09934750360720516,
1330
+ "ci95": [
1331
+ 0.0454206858231911,
1332
+ 0.14433272375703962
1333
+ ]
1334
+ },
1335
+ "B:gliner - A:gliner": {
1336
+ "difference": -0.087189858534691,
1337
+ "ci95": [
1338
+ -0.13204415683257073,
1339
+ -0.0406803117967699
1340
+ ]
1341
+ }
1342
+ },
1343
+ "2 labels (13)": {
1344
+ "A:julia - A:ines": {
1345
+ "difference": 0.020990429437646712,
1346
+ "ci95": [
1347
+ -0.06061145608979183,
1348
+ 0.10740818043124331
1349
+ ]
1350
+ },
1351
+ "A:laya - A:ines": {
1352
+ "difference": 0.16041839264233365,
1353
+ "ci95": [
1354
+ 0.059458802921318936,
1355
+ 0.27206775011759543
1356
+ ]
1357
+ },
1358
+ "A:gliner - A:ines": {
1359
+ "difference": 0.0659090544638859,
1360
+ "ci95": [
1361
+ 0.004675393764386895,
1362
+ 0.12582906809147268
1363
+ ]
1364
+ },
1365
+ "A:gemma - A:ines": {
1366
+ "difference": 0.33584330369564275,
1367
+ "ci95": [
1368
+ 0.25214648887513785,
1369
+ 0.4349238101458538
1370
+ ]
1371
+ },
1372
+ "B:julia - B:ines": {
1373
+ "difference": -0.11230612894229851,
1374
+ "ci95": [
1375
+ -0.18080307502401038,
1376
+ -0.042581067377537536
1377
+ ]
1378
+ },
1379
+ "B:laya - B:ines": {
1380
+ "difference": 0.14622586299247847,
1381
+ "ci95": [
1382
+ 0.09520841847771228,
1383
+ 0.20567775456268422
1384
+ ]
1385
+ },
1386
+ "B:gliner - B:ines": {
1387
+ "difference": -0.057765534794081974,
1388
+ "ci95": [
1389
+ -0.1428996478267053,
1390
+ 0.029876283921421498
1391
+ ]
1392
+ },
1393
+ "B:ines - A:ines": {
1394
+ "difference": 0.06805235059944967,
1395
+ "ci95": [
1396
+ -0.04906521051519053,
1397
+ 0.19755199873107607
1398
+ ]
1399
+ },
1400
+ "B:julia - A:julia": {
1401
+ "difference": -0.06524420778049556,
1402
+ "ci95": [
1403
+ -0.12548066475124145,
1404
+ -0.012510733064128867
1405
+ ]
1406
+ },
1407
+ "B:laya - A:laya": {
1408
+ "difference": 0.05385982094959447,
1409
+ "ci95": [
1410
+ -0.0009695761275055773,
1411
+ 0.1065780484194153
1412
+ ]
1413
+ },
1414
+ "B:gliner - A:gliner": {
1415
+ "difference": -0.0556222386585182,
1416
+ "ci95": [
1417
+ -0.11618959302529945,
1418
+ 0.007571405517821818
1419
+ ]
1420
+ }
1421
+ },
1422
+ "3-20 labels (4)": {
1423
+ "A:julia - A:ines": {
1424
+ "difference": 0.14470497485076556,
1425
+ "ci95": [
1426
+ -0.12895039644581718,
1427
+ 0.42568245069405136
1428
+ ]
1429
+ },
1430
+ "A:laya - A:ines": {
1431
+ "difference": 0.07048620069879838,
1432
+ "ci95": [
1433
+ -0.229964039114142,
1434
+ 0.3531563180395061
1435
+ ]
1436
+ },
1437
+ "A:gliner - A:ines": {
1438
+ "difference": 0.052399605488675075,
1439
+ "ci95": [
1440
+ -0.05581876515491288,
1441
+ 0.1600753585487263
1442
+ ]
1443
+ },
1444
+ "A:gemma - A:ines": {
1445
+ "difference": 0.2063628185419505,
1446
+ "ci95": [
1447
+ 0.06159339629875036,
1448
+ 0.3329951554044995
1449
+ ]
1450
+ },
1451
+ "B:julia - B:ines": {
1452
+ "difference": -0.048017432999332976,
1453
+ "ci95": [
1454
+ -0.16164512046595278,
1455
+ 0.05754769356829411
1456
+ ]
1457
+ },
1458
+ "B:laya - B:ines": {
1459
+ "difference": 0.364651999561586,
1460
+ "ci95": [
1461
+ 0.18329205393405798,
1462
+ 0.5540391760322673
1463
+ ]
1464
+ },
1465
+ "B:gliner - B:ines": {
1466
+ "difference": 0.09523556853903474,
1467
+ "ci95": [
1468
+ 0.01900107622902177,
1469
+ 0.17430955225269407
1470
+ ]
1471
+ },
1472
+ "B:ines - A:ines": {
1473
+ "difference": -0.2056134313620306,
1474
+ "ci95": [
1475
+ -0.3164880910714138,
1476
+ -0.09244502441135767
1477
+ ]
1478
+ },
1479
+ "B:julia - A:julia": {
1480
+ "difference": -0.3983358392121291,
1481
+ "ci95": [
1482
+ -0.5756158997267588,
1483
+ -0.20827382157801752
1484
+ ]
1485
+ },
1486
+ "B:laya - A:laya": {
1487
+ "difference": 0.08855236750075707,
1488
+ "ci95": [
1489
+ -0.008752159824462086,
1490
+ 0.1939655463264431
1491
+ ]
1492
+ },
1493
+ "B:gliner - A:gliner": {
1494
+ "difference": -0.16277746831167095,
1495
+ "ci95": [
1496
+ -0.25637377988986315,
1497
+ -0.07162803351218687
1498
+ ]
1499
+ }
1500
+ },
1501
+ "21-72 labels (5)": {
1502
+ "A:julia - A:ines": {
1503
+ "difference": 0.05261854328754896,
1504
+ "ci95": [
1505
+ -0.04301551177445001,
1506
+ 0.15765204279382006
1507
+ ]
1508
+ },
1509
+ "A:laya - A:ines": {
1510
+ "difference": -0.22623626968869046,
1511
+ "ci95": [
1512
+ -0.30032025667320744,
1513
+ -0.12187817221877785
1514
+ ]
1515
+ },
1516
+ "A:gliner - A:ines": {
1517
+ "difference": 0.06298151719422522,
1518
+ "ci95": [
1519
+ -0.039063579916225874,
1520
+ 0.1689149673736492
1521
+ ]
1522
+ },
1523
+ "A:gemma - A:ines": {
1524
+ "difference": 0.39411989145959925,
1525
+ "ci95": [
1526
+ 0.3019916275555586,
1527
+ 0.473337812412976
1528
+ ]
1529
+ },
1530
+ "B:julia - B:ines": {
1531
+ "difference": -0.10772448411645877,
1532
+ "ci95": [
1533
+ -0.15259487999276594,
1534
+ -0.04920310152464395
1535
+ ]
1536
+ },
1537
+ "B:laya - B:ines": {
1538
+ "difference": 0.0878280900717231,
1539
+ "ci95": [
1540
+ 0.011100563313016499,
1541
+ 0.14839732102914302
1542
+ ]
1543
+ },
1544
+ "B:gliner - B:ines": {
1545
+ "difference": 0.041998707161331236,
1546
+ "ci95": [
1547
+ -0.019529909671919345,
1548
+ 0.11057761360469363
1549
+ ]
1550
+ },
1551
+ "B:ines - A:ines": {
1552
+ "difference": -0.08781277235826232,
1553
+ "ci95": [
1554
+ -0.12624266008952723,
1555
+ -0.04849955912024958
1556
+ ]
1557
+ },
1558
+ "B:julia - A:julia": {
1559
+ "difference": -0.24815579976227004,
1560
+ "ci95": [
1561
+ -0.3988206525705068,
1562
+ -0.09861247338677506
1563
+ ]
1564
+ },
1565
+ "B:laya - A:laya": {
1566
+ "difference": 0.22625158740215126,
1567
+ "ci95": [
1568
+ 0.12816497504348637,
1569
+ 0.286499258070288
1570
+ ]
1571
+ },
1572
+ "B:gliner - A:gliner": {
1573
+ "difference": -0.1087955823911563,
1574
+ "ci95": [
1575
+ -0.1700071752364474,
1576
+ -0.032048604061386245
1577
+ ]
1578
+ }
1579
+ },
1580
+ "absent from Ines-1 decision FT mixture (15)": {
1581
+ "A:julia - A:ines": {
1582
+ "difference": 0.050144473325553045,
1583
+ "ci95": [
1584
+ -0.028885054771255515,
1585
+ 0.1500566532935792
1586
+ ]
1587
+ },
1588
+ "A:laya - A:ines": {
1589
+ "difference": 0.029583466640483884,
1590
+ "ci95": [
1591
+ -0.08075161791194013,
1592
+ 0.14421513826606108
1593
+ ]
1594
+ },
1595
+ "A:gliner - A:ines": {
1596
+ "difference": 0.04219729401554383,
1597
+ "ci95": [
1598
+ -0.006600902674686908,
1599
+ 0.09437438464034366
1600
+ ]
1601
+ },
1602
+ "A:gemma - A:ines": {
1603
+ "difference": 0.2790972445681884,
1604
+ "ci95": [
1605
+ 0.211381616754686,
1606
+ 0.3521485454260135
1607
+ ]
1608
+ },
1609
+ "B:julia - B:ines": {
1610
+ "difference": -0.0775557848108404,
1611
+ "ci95": [
1612
+ -0.14064220126958463,
1613
+ -0.013517268561773333
1614
+ ]
1615
+ },
1616
+ "B:laya - B:ines": {
1617
+ "difference": 0.2020010278286585,
1618
+ "ci95": [
1619
+ 0.1269579464717378,
1620
+ 0.2900629696566373
1621
+ ]
1622
+ },
1623
+ "B:gliner - B:ines": {
1624
+ "difference": 0.019015644443022794,
1625
+ "ci95": [
1626
+ -0.03174972794352044,
1627
+ 0.07223435697834166
1628
+ ]
1629
+ },
1630
+ "B:ines - A:ines": {
1631
+ "difference": -0.0764229201443689,
1632
+ "ci95": [
1633
+ -0.15017289217507543,
1634
+ -0.009184764537925973
1635
+ ]
1636
+ },
1637
+ "B:julia - A:julia": {
1638
+ "difference": -0.20412317828076224,
1639
+ "ci95": [
1640
+ -0.2998785466263466,
1641
+ -0.11592881332069362
1642
+ ]
1643
+ },
1644
+ "B:laya - A:laya": {
1645
+ "difference": 0.09599464104380576,
1646
+ "ci95": [
1647
+ 0.0459047217540702,
1648
+ 0.1442013728576654
1649
+ ]
1650
+ },
1651
+ "B:gliner - A:gliner": {
1652
+ "difference": -0.09960456971688988,
1653
+ "ci95": [
1654
+ -0.14649500828750583,
1655
+ -0.04960181354566028
1656
+ ]
1657
+ }
1658
+ },
1659
+ "families in Ines-1 decision FT mixture (7)": {
1660
+ "A:julia - A:ines": {
1661
+ "difference": 0.051803014092416944,
1662
+ "ci95": [
1663
+ -0.10019425716510703,
1664
+ 0.1778425425557031
1665
+ ]
1666
+ },
1667
+ "A:laya - A:ines": {
1668
+ "difference": 0.11320722272783149,
1669
+ "ci95": [
1670
+ -0.12479087223953271,
1671
+ 0.3295312555301822
1672
+ ]
1673
+ },
1674
+ "A:gliner - A:ines": {
1675
+ "difference": 0.10690918653188367,
1676
+ "ci95": [
1677
+ 0.019280206335977032,
1678
+ 0.19106643449909544
1679
+ ]
1680
+ },
1681
+ "A:gemma - A:ines": {
1682
+ "difference": 0.42507928728376093,
1683
+ "ci95": [
1684
+ 0.28040693225491636,
1685
+ 0.5352594526522105
1686
+ ]
1687
+ },
1688
+ "B:julia - B:ines": {
1689
+ "difference": -0.14676215095241435,
1690
+ "ci95": [
1691
+ -0.1986636251903599,
1692
+ -0.08740265140015356
1693
+ ]
1694
+ },
1695
+ "B:laya - B:ines": {
1696
+ "difference": 0.10980989286818602,
1697
+ "ci95": [
1698
+ 0.03243625252632493,
1699
+ 0.19414880936010162
1700
+ ]
1701
+ },
1702
+ "B:gliner - B:ines": {
1703
+ "difference": -0.06360725842937322,
1704
+ "ci95": [
1705
+ -0.21886027228587646,
1706
+ 0.08094216222892091
1707
+ ]
1708
+ },
1709
+ "B:ines - A:ines": {
1710
+ "difference": 0.10992953895984924,
1711
+ "ci95": [
1712
+ -0.07562649281899102,
1713
+ 0.2927199100383737
1714
+ ]
1715
+ },
1716
+ "B:julia - A:julia": {
1717
+ "difference": -0.08863562608498207,
1718
+ "ci95": [
1719
+ -0.24495270619294315,
1720
+ 0.033542744774361914
1721
+ ]
1722
+ },
1723
+ "B:laya - A:laya": {
1724
+ "difference": 0.10653220910020378,
1725
+ "ci95": [
1726
+ -0.015590406007179902,
1727
+ 0.20797413933948058
1728
+ ]
1729
+ },
1730
+ "B:gliner - A:gliner": {
1731
+ "difference": -0.06058690600140765,
1732
+ "ci95": [
1733
+ -0.14995189706107254,
1734
+ 0.032545729267887305
1735
+ ]
1736
+ }
1737
+ }
1738
+ },
1739
+ "calibration": {
1740
+ "native (Track A, full native distributions, <=20 labels, 17 datasets)": {
1741
+ "ines": {
1742
+ "brier": 0.5014338202978408,
1743
+ "nll": 0.8000206904458703,
1744
+ "ece": 0.1270706365126021,
1745
+ "coverage_at_error_budget": 0.08941176470588236,
1746
+ "mean_confidence": 0.6118652016888647,
1747
+ "datasets": 17
1748
+ },
1749
+ "julia": {
1750
+ "brier": 0.645280763707045,
1751
+ "nll": 1.6618435294062641,
1752
+ "ece": 0.2983721069531266,
1753
+ "coverage_at_error_budget": 0.1288235294117647,
1754
+ "mean_confidence": 0.8941619627630315,
1755
+ "datasets": 17
1756
+ },
1757
+ "laya": {
1758
+ "brier": 0.4006447652764705,
1759
+ "nll": 0.6925248427789752,
1760
+ "ece": 0.09332229411764704,
1761
+ "coverage_at_error_budget": 0.25588235294117645,
1762
+ "mean_confidence": 0.6874152352941176,
1763
+ "datasets": 17
1764
+ },
1765
+ "gliner": {
1766
+ "brier": 0.48542838104085484,
1767
+ "nll": 0.8599173300405782,
1768
+ "ece": 0.1585389382487745,
1769
+ "coverage_at_error_budget": 0.19705882352941173,
1770
+ "mean_confidence": 0.7310719386879542,
1771
+ "datasets": 17
1772
+ },
1773
+ "gemma": {
1774
+ "brier": 0.24999771502464618,
1775
+ "nll": 0.6128898920857043,
1776
+ "ece": 0.1156891273014901,
1777
+ "coverage_at_error_budget": 0.693529411764706,
1778
+ "mean_confidence": 0.9580137741914275,
1779
+ "datasets": 17
1780
+ }
1781
+ },
1782
+ "derived (Track B softmax of log-odds, same 17 datasets)": {
1783
+ "ines": {
1784
+ "brier": 0.5266089613872135,
1785
+ "nll": 0.8626628971730973,
1786
+ "ece": 0.11876053221277025,
1787
+ "coverage_at_error_budget": 0.1335294117647059,
1788
+ "mean_confidence": 0.4924301710578631,
1789
+ "datasets": 17
1790
+ },
1791
+ "julia": {
1792
+ "brier": 0.7685502933200778,
1793
+ "nll": 1.5719702226511314,
1794
+ "ece": 0.31343832319933446,
1795
+ "coverage_at_error_budget": 0.014117647058823528,
1796
+ "mean_confidence": 0.7770917395785171,
1797
+ "datasets": 17
1798
+ },
1799
+ "laya": {
1800
+ "brier": 0.35941673886508035,
1801
+ "nll": 0.6807608663005745,
1802
+ "ece": 0.12656927310814167,
1803
+ "coverage_at_error_budget": 0.38176470588235295,
1804
+ "mean_confidence": 0.7694262437653563,
1805
+ "datasets": 17
1806
+ },
1807
+ "gliner": {
1808
+ "brier": 0.5872442915941616,
1809
+ "nll": 1.0308428769681912,
1810
+ "ece": 0.2029575594285007,
1811
+ "coverage_at_error_budget": 0.03705882352941177,
1812
+ "mean_confidence": 0.7019368747032854,
1813
+ "datasets": 17
1814
+ }
1815
+ },
1816
+ "derived (Track B softmax of log-odds, all 22 datasets)": {
1817
+ "ines": {
1818
+ "brier": 0.6231953131399431,
1819
+ "nll": 1.4522999819886815,
1820
+ "ece": 0.12354853986128157,
1821
+ "coverage_at_error_budget": 0.10590909090909091,
1822
+ "mean_confidence": 0.39202771650296064,
1823
+ "datasets": 22
1824
+ },
1825
+ "julia": {
1826
+ "brier": 0.8554517402951244,
1827
+ "nll": 2.468274766821923,
1828
+ "ece": 0.31679964418123413,
1829
+ "coverage_at_error_budget": 0.011363636363636364,
1830
+ "mean_confidence": 0.687741939307485,
1831
+ "datasets": 22
1832
+ },
1833
+ "laya": {
1834
+ "brier": 0.47884708479764443,
1835
+ "nll": 1.203588850522752,
1836
+ "ece": 0.12330875423821963,
1837
+ "coverage_at_error_budget": 0.29681818181818187,
1838
+ "mean_confidence": 0.6690071844208508,
1839
+ "datasets": 22
1840
+ },
1841
+ "gliner": {
1842
+ "brier": 0.6643794367941797,
1843
+ "nll": 1.5673671850720794,
1844
+ "ece": 0.18732672413963056,
1845
+ "coverage_at_error_budget": 0.031363636363636364,
1846
+ "mean_confidence": 0.5655688289154754,
1847
+ "datasets": 22
1848
+ }
1849
+ }
1850
+ },
1851
+ "ines_ablation": {
1852
+ "agnews": {
1853
+ "labels": 4,
1854
+ "accuracy": {
1855
+ "native(<=26)": 0.59,
1856
+ "knockout-20": null,
1857
+ "knockout-2": 0.64,
1858
+ "one-vs-rest": 0.42
1859
+ },
1860
+ "macro_f1": {
1861
+ "native(<=26)": 0.5843451916892106,
1862
+ "knockout-20": null,
1863
+ "knockout-2": 0.639691081949832,
1864
+ "one-vs-rest": 0.32652938961350175
1865
+ }
1866
+ },
1867
+ "emotiondair": {
1868
+ "labels": 6,
1869
+ "accuracy": {
1870
+ "native(<=26)": 0.43,
1871
+ "knockout-20": null,
1872
+ "knockout-2": 0.43,
1873
+ "one-vs-rest": 0.34
1874
+ },
1875
+ "macro_f1": {
1876
+ "native(<=26)": 0.39042230395613853,
1877
+ "knockout-20": null,
1878
+ "knockout-2": 0.39217171717171717,
1879
+ "one-vs-rest": 0.29478472669962036
1880
+ }
1881
+ },
1882
+ "banking77": {
1883
+ "labels": 72,
1884
+ "accuracy": {
1885
+ "native(<=26)": null,
1886
+ "knockout-20": 0.39,
1887
+ "knockout-2": 0.45,
1888
+ "one-vs-rest": 0.29
1889
+ },
1890
+ "macro_f1": {
1891
+ "native(<=26)": null,
1892
+ "knockout-20": 0.34265873015873016,
1893
+ "knockout-2": 0.4010912698412698,
1894
+ "one-vs-rest": 0.2234898589065256
1895
+ }
1896
+ },
1897
+ "financialphrasebank": {
1898
+ "labels": 3,
1899
+ "accuracy": {
1900
+ "native(<=26)": 0.51,
1901
+ "knockout-20": null,
1902
+ "knockout-2": 0.44,
1903
+ "one-vs-rest": 0.45
1904
+ },
1905
+ "macro_f1": {
1906
+ "native(<=26)": 0.48911853494169083,
1907
+ "knockout-20": null,
1908
+ "knockout-2": 0.3744570258331727,
1909
+ "one-vs-rest": 0.388558378810103
1910
+ }
1911
+ },
1912
+ "empathetic": {
1913
+ "labels": 32,
1914
+ "accuracy": {
1915
+ "native(<=26)": null,
1916
+ "knockout-20": 0.25,
1917
+ "knockout-2": 0.16,
1918
+ "one-vs-rest": 0.18
1919
+ },
1920
+ "macro_f1": {
1921
+ "native(<=26)": null,
1922
+ "knockout-20": 0.1909003551273288,
1923
+ "knockout-2": 0.10230728593731689,
1924
+ "one-vs-rest": 0.14416745535166586
1925
+ }
1926
+ },
1927
+ "massive": {
1928
+ "labels": 59,
1929
+ "accuracy": {
1930
+ "native(<=26)": null,
1931
+ "knockout-20": 0.35,
1932
+ "knockout-2": 0.41,
1933
+ "one-vs-rest": 0.2
1934
+ },
1935
+ "macro_f1": {
1936
+ "native(<=26)": null,
1937
+ "knockout-20": 0.2502433536331841,
1938
+ "knockout-2": 0.3015873015873015,
1939
+ "one-vs-rest": 0.15282852740479855
1940
+ }
1941
+ },
1942
+ "yahootopics": {
1943
+ "labels": 10,
1944
+ "accuracy": {
1945
+ "native(<=26)": 0.5,
1946
+ "knockout-20": null,
1947
+ "knockout-2": 0.57,
1948
+ "one-vs-rest": 0.18
1949
+ },
1950
+ "macro_f1": {
1951
+ "native(<=26)": 0.4915120854091442,
1952
+ "knockout-20": null,
1953
+ "knockout-2": 0.5584408865987813,
1954
+ "one-vs-rest": 0.12307189542483658
1955
+ }
1956
+ },
1957
+ "manifesto": {
1958
+ "labels": 56,
1959
+ "accuracy": {
1960
+ "native(<=26)": null,
1961
+ "knockout-20": 0.11,
1962
+ "knockout-2": 0.18,
1963
+ "one-vs-rest": 0.05
1964
+ },
1965
+ "macro_f1": {
1966
+ "native(<=26)": null,
1967
+ "knockout-20": 0.07504578754578754,
1968
+ "knockout-2": 0.14917929292929294,
1969
+ "one-vs-rest": 0.032902735562310034
1970
+ }
1971
+ },
1972
+ "capsotu": {
1973
+ "labels": 21,
1974
+ "accuracy": {
1975
+ "native(<=26)": 0.36,
1976
+ "knockout-20": 0.37,
1977
+ "knockout-2": 0.39,
1978
+ "one-vs-rest": 0.21
1979
+ },
1980
+ "macro_f1": {
1981
+ "native(<=26)": 0.3120540695728665,
1982
+ "knockout-20": 0.3317682178584434,
1983
+ "knockout-2": 0.3739835503810659,
1984
+ "one-vs-rest": 0.17844985702128557
1985
+ }
1986
+ }
1987
+ },
1988
+ "anchors": {
1989
+ "GLiNER2.5 Track A vs published pilot (accuracy)": {
1990
+ "agnews": {
1991
+ "ours": 0.7,
1992
+ "published": 0.7
1993
+ },
1994
+ "emotiondair": {
1995
+ "ours": 0.44,
1996
+ "published": 0.44
1997
+ },
1998
+ "banking77": {
1999
+ "ours": 0.61,
2000
+ "published": 0.61
2001
+ }
2002
+ },
2003
+ "Julia-1 Track A vs its model card (accuracy)": {
2004
+ "agnews": {
2005
+ "ours": 0.94,
2006
+ "card": 0.94,
2007
+ "ours_strategy": "single request"
2008
+ },
2009
+ "emotiondair": {
2010
+ "ours": 0.86,
2011
+ "card": 0.86,
2012
+ "ours_strategy": "single request"
2013
+ },
2014
+ "banking77": {
2015
+ "ours": 0.63,
2016
+ "card": 0.64,
2017
+ "ours_strategy": "Router(width=20, survivors=2)"
2018
+ }
2019
+ }
2020
+ },
2021
+ "ovr_ties": {
2022
+ "ines": {
2023
+ "examples_with_tie_at_max": 268,
2024
+ "n": 2200
2025
+ },
2026
+ "julia": {
2027
+ "examples_with_tie_at_max": 4,
2028
+ "n": 2200
2029
+ },
2030
+ "laya": {
2031
+ "examples_with_tie_at_max": 11,
2032
+ "n": 2200
2033
+ },
2034
+ "gliner": {
2035
+ "examples_with_tie_at_max": 0,
2036
+ "n": 2200
2037
+ }
2038
+ },
2039
+ "addon": {
2040
+ "date": "2026-10-06",
2041
+ "model": "RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic (DiffusionGemma-26B-A4B-it, FP8) via razorback16/openjev:0.4.0, model openjev-0.1, think 0",
2042
+ "scope": "Track A (native) only; added after the btzsc22-ext-v1 freeze; the four original models are recomputed by the same code and match results.json exactly"
2043
+ }
2044
+ }
eval/btzsc22/addon_gemma/trackA_native.gemma.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
eval/speed.md CHANGED
@@ -34,3 +34,36 @@
34
  ~180k prompt tokens/s here, close to its isolated prefill ceiling (~230k).
35
  - Not a general speed ranking: different serving stacks (1 vs 4 processes, graphs vs none), different prompt
36
  lengths for the same request, one GPU.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
  ~180k prompt tokens/s here, close to its isolated prefill ceiling (~230k).
35
  - Not a general speed ranking: different serving stacks (1 vs 4 processes, graphs vs none), different prompt
36
  lengths for the same request, one GPU.
37
+
38
+ ## Run 3 (2026-10-06): four models, one GPU, with memory
39
+
40
+ Same load generator (`scripts/bench_jev.py`, closed loop, 16 client processes, HTTP/1.1 without keep-alive), same two
41
+ request sets and levels as above, **one free H200 per model in turn** (GPU 2 of the node, nothing else on it). The
42
+ only change to the load generator: the request's `model` field is taken from `BENCH_MODEL` (OpenJev rejects requests
43
+ without a known model; the other servers ignore it). GPU memory was sampled every 0.5 s with `nvidia-smi` (whole GPU).
44
+ Raw output: `speed_runs/2026-10-06_part1_gemma_laya.txt`, `speed_runs/2026-10-06_part2_gliner_ines.txt` (the first
45
+ GLiNER and Ines-1 start attempts failed on environment issues before any request and were rerun in part 2).
46
+
47
+ | model (params total / active) | server | short c=1 p50 / p95 | short max q/s (c) | typed c=1 p50 / p95 | typed max q/s (c) | GPU memory idle / peak |
48
+ |---|---|---:|---:|---:|---:|---:|
49
+ | **Ines-1** (1.59B / 405M, bf16) | `scripts/serve_jev.py`, 1 process, CUDA graphs, dynamic batching | **9.4 / 9.7 ms** | **1,896** (128) | 30.6 / **40.5 ms** | 544 (32) | 9.7 / 21.9 GB |
50
+ | laya-multilingual arch. (322M, bf16) | production Jev server, 4 processes, micro-batching | 13.0 / 25.9 ms | 1,848 (256) | **20.6** / 89.2 ms | **840** (128) | 8.0 / 24.1 GB |
51
+ | GLiNER2.5-multi arch. (287M) | GLiNER Jev server, 4 workers, **no batching** | 48.8 / 49.3 ms | 68 (64) | 123.1 / 129.3 ms | 49 (32) | 6.7 / 26.5 GB |
52
+ | DiffusionGemma-26B-A4B-it FP8 (~26B / ~4B) | OpenJev 0.4.0 (vLLM), max 64 sequences | 91.7 / 94.0 ms | 79 (32) | 158.8 / 170.2 ms | 143 (32) | weights 25.8 GiB; **120.6 / 135.1 GB reserved** |
53
+
54
+ Latency at the best-throughput level (p50 / p95): Ines-1 118 / 197 ms (short), 282 / 386 ms (typed); laya 212 / 483 ms,
55
+ 604 / 913 ms; GLiNER 1,729 / 3,805 ms, 2,070 / 6,871 ms; Gemma 718 / 1,439 ms, 1,008 / 1,545 ms.
56
+
57
+ Reading:
58
+ - **Ines-1 has the lowest single-request latency on short requests and matches the 322M encoder's peak short
59
+ throughput** with one process. On ~300-token questions the encoder is still faster (840 vs 544 questions/s,
60
+ 20.6 vs 30.6 ms median), but Ines-1's tail is tighter (p95 40.5 vs 89.2 ms), as in runs 1–2.
61
+ - **Gemma is 5–10× slower at c=1 and ~4–24× lower in peak throughput** than Ines-1 or laya, and returned HTTP 503 on
62
+ 12 of the ~12,000 short requests sent at c ≥ 32 (none on typed cases). Its memory figure is vLLM's preallocation (`gpu_memory_utilization`
63
+ 0.80: 25.8 GiB of weights + 83 GiB of KV cache), not a minimum; it needs at least the ~26 GiB of weights.
64
+ - **GLiNER's numbers measure its server, not the architecture's ceiling**: that server answers one request per
65
+ worker with no batching, so throughput stays flat as concurrency rises. Input tokens per question are counted
66
+ differently by each server (`usage`): Ines-1 61.5 / 316.2, laya 35.0 / 311.6, GLiNER 8.5 / 37.0, Gemma 67.0 / 136.7.
67
+ - Peak memory includes activations at the highest concurrency; idle is after load and warm-up. The encoder rows use
68
+ domain fine-tunes of those architectures (speed is the architecture's), as in runs 1–2.
69
+ - Not a general speed ranking: different serving stacks, prompt templates and process counts, one GPU.
eval/speed_runs/2026-10-06_part1_gemma_laya.txt ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ===== gemma 11:23:26
2
+ VRAM_IDLE_MiB 120557
3
+ == short (2 questions)
4
+ c=1 n=3000 11 req/s 22 questions/s p50 91.7 ms p90 93.1 p95 94.0 p99 95.9 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 3000}
5
+ c=8 n=3000 32 req/s 64 questions/s p50 266.5 ms p90 292.5 p95 300.6 p99 357.7 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 3000}
6
+ c=32 n=2992 40 req/s 79 questions/s p50 718.2 ms p90 1260.2 p95 1438.8 p99 1856.4 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2985, 503: 7}
7
+ c=64 n=2992 38 req/s 76 questions/s p50 1506.4 ms p90 2498.3 p95 2818.2 p99 3528.7 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2991, 503: 1}
8
+ c=128 n=2992 36 req/s 71 questions/s p50 3376.2 ms p90 4356.4 p95 4636.1 p99 5255.2 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2988, 503: 4}
9
+ c=256 n=2992 36 req/s 72 questions/s p50 6448.7 ms p90 7529.1 p95 7880.5 p99 8527.4 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2992}
10
+ == typed-decisions ES (5 questions per case)
11
+ c=1 n=800 6 req/s 32 questions/s p50 158.8 ms p90 166.4 p95 170.2 p99 177.3 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
12
+ c=8 n=800 21 req/s 107 questions/s p50 382.0 ms p90 410.2 p95 429.2 p99 522.9 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
13
+ c=32 n=800 29 req/s 143 questions/s p50 1007.7 ms p90 1382.2 p95 1544.9 p99 1773.0 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
14
+ c=128 n=800 25 req/s 127 questions/s p50 4236.0 ms p90 4894.6 p95 5139.9 p99 5471.4 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
15
+ VRAM_PEAK_MiB 135149
16
+ (APIServer pid=560) INFO 10-06 11:07:52 [api_utils.py:286] non-default args: {'model_tag': 'RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic', 'enable_auto_tool_choice': True, 'tool_call_parser': 'gemma4', 'host': '127.0.0.1', 'model': 'RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic', 'max_model_len': 65536, 'max_logprobs': 32, 'served_model_name': ['dgemma'], 'override_generation_config': {'max_new_tokens': None}, 'attention_backend': 'TRITON_ATTN', 'reasoning_parser': 'gemma4', 'gpu_memory_utilization': 0.8, 'enable_prefix_caching': True, 'limit_mm_per_prompt': {'image': 8, 'video': 0}, 'max_num_seqs': 64, 'async_scheduling': True, 'diffusion_config': {'canvas_length': 64}}
17
+ (EngineCore pid=1209) INFO 10-06 11:08:27 [default_loader.py:430] Loading weights took 16.33 seconds
18
+ (EngineCore pid=1209) INFO 10-06 11:08:28 [model_runner.py:408] Model loading took 25.83 GiB memory and 20.123655 seconds
19
+ (EngineCore pid=1209) INFO 10-06 11:08:28 [utils.py:320] Using LBNHC KV cache layout.
20
+ (EngineCore pid=1209) INFO 10-06 11:08:50 [gpu_worker.py:674] Available KV cache memory: 83.17 GiB
21
+ (EngineCore pid=1209) INFO 10-06 11:08:50 [gpu_worker.py:689] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.8000 is equivalent to --gpu-memory-utilization=0.7968 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.8032. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0.
22
+ (EngineCore pid=1209) INFO 10-06 11:08:50 [kv_cache_utils.py:2432] GPU KV cache size: 1,191,787 tokens, Maximum concurrency for 65,536 tokens per request: 18.19x
23
+ (EngineCore pid=1209) INFO 10-06 11:09:07 [gpu_worker.py:860] CUDA graph pool memory: 0.13 GiB (actual), 0.44 GiB (estimated), difference: 0.31 GiB (231.6%).
24
+ ===== laya 11:39:11
25
+ VRAM_IDLE_MiB 7986
26
+ == short (2 questions)
27
+ c=1 n=3000 61 req/s 122 questions/s p50 13.0 ms p90 25.0 p95 25.9 p99 29.0 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 3000}
28
+ c=8 n=3000 284 req/s 567 questions/s p50 28.0 ms p90 33.1 p95 34.9 p99 107.4 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 3000}
29
+ c=32 n=2992 598 req/s 1197 questions/s p50 37.8 ms p90 107.6 p95 142.1 p99 213.6 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
30
+ c=64 n=2992 765 req/s 1529 questions/s p50 50.2 ms p90 182.6 p95 220.8 p99 261.2 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
31
+ c=128 n=2992 873 req/s 1747 questions/s p50 115.0 ms p90 219.9 p95 239.6 p99 272.7 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
32
+ c=256 n=2992 924 req/s 1848 questions/s p50 211.6 ms p90 344.3 p95 483.3 p99 566.3 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
33
+ == typed-decisions ES (5 questions per case)
34
+ c=1 n=800 30 req/s 152 questions/s p50 20.6 ms p90 85.2 p95 89.2 p99 101.3 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
35
+ c=8 n=800 107 req/s 533 questions/s p50 33.1 ms p90 174.0 p95 208.0 p99 270.4 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
36
+ c=32 n=800 150 req/s 750 questions/s p50 178.9 ms p90 351.6 p95 454.6 p99 652.9 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
37
+ c=128 n=800 168 req/s 840 questions/s p50 603.9 ms p90 873.1 p95 912.8 p99 1285.5 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
38
+ VRAM_PEAK_MiB 24076
39
+ NO ARRANCA http://localhost:30043
40
+ Error response from daemon: No such container: banco-gliner
eval/speed_runs/2026-10-06_part2_gliner_ines.txt ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ===== gliner 11:57:45
2
+ VRAM_IDLE_MiB 6650
3
+ == short (2 questions)
4
+ c=1 n=3000 20 req/s 40 questions/s p50 48.8 ms p90 49.1 p95 49.3 p99 51.2 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 3000}
5
+ c=8 n=3000 21 req/s 43 questions/s p50 419.9 ms p90 487.7 p95 510.9 p99 555.9 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 3000}
6
+ c=32 n=2992 21 req/s 42 questions/s p50 1672.7 ms p90 2033.1 p95 2096.2 p99 2262.7 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
7
+ c=64 n=2992 34 req/s 68 questions/s p50 1729.4 ms p90 3542.9 p95 3804.8 p99 4448.3 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
8
+ c=128 n=2992 33 req/s 65 questions/s p50 2407.2 ms p90 9051.0 p95 9471.4 p99 9858.6 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
9
+ c=256 n=2992 26 req/s 53 questions/s p50 5912.3 ms p90 17059.3 p95 17745.1 p99 17961.1 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
10
+ == typed-decisions ES (5 questions per case)
11
+ c=1 n=800 8 req/s 40 questions/s p50 123.1 ms p90 125.5 p95 129.3 p99 198.9 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
12
+ c=8 n=800 8 req/s 41 questions/s p50 997.9 ms p90 1161.4 p95 1297.9 p99 1472.9 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
13
+ c=32 n=800 10 req/s 49 questions/s p50 2069.6 ms p90 6033.7 p95 6870.5 p99 7986.3 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
14
+ c=128 n=800 9 req/s 45 questions/s p50 11465.9 ms p90 17167.3 p95 19262.3 p99 20541.8 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
15
+ VRAM_PEAK_MiB 26544
16
+ ===== ines 12:16:37
17
+ VRAM_IDLE_MiB 9747
18
+ == short (2 questions)
19
+ c=1 n=3000 105 req/s 210 questions/s p50 9.4 ms p90 9.7 p95 9.7 p99 9.8 input tok/question 61.5 rows/forward 2.0 in forward 70 % {200: 3000}
20
+ c=8 n=3000 443 req/s 887 questions/s p50 17.8 ms p90 18.0 p95 18.1 p99 18.4 input tok/question 61.5 rows/forward 8.0 in forward 98 % {200: 3000}
21
+ c=32 n=2992 664 req/s 1328 questions/s p50 45.2 ms p90 47.2 p95 52.9 p99 143.1 input tok/question 61.5 rows/forward 30.2 in forward 95 % {200: 2992}
22
+ c=64 n=2992 891 req/s 1782 questions/s p50 66.2 ms p90 69.4 p95 83.9 p99 142.0 input tok/question 61.5 rows/forward 41.6 in forward 96 % {200: 2992}
23
+ c=128 n=2992 948 req/s 1896 questions/s p50 117.8 ms p90 153.1 p95 197.1 p99 198.6 input tok/question 61.5 rows/forward 49.9 in forward 95 % {200: 2992}
24
+ c=256 n=2992 908 req/s 1816 questions/s p50 224.6 ms p90 307.5 p95 343.7 p99 353.6 input tok/question 61.5 rows/forward 53.2 in forward 86 % {200: 2992}
25
+ == typed-decisions ES (5 questions per case)
26
+ c=1 n=800 32 req/s 162 questions/s p50 30.6 ms p90 38.4 p95 40.5 p99 48.4 input tok/question 316.2 rows/forward 5.0 in forward 55 % {200: 800}
27
+ c=8 n=800 93 req/s 465 questions/s p50 85.4 ms p90 110.2 p95 116.2 p99 124.0 input tok/question 316.2 rows/forward 17.6 in forward 99 % {200: 800}
28
+ c=32 n=800 109 req/s 544 questions/s p50 282.3 ms p90 370.3 p95 385.5 p99 457.8 input tok/question 316.2 rows/forward 48.4 in forward 99 % {200: 800}
29
+ c=128 n=800 106 req/s 528 questions/s p50 1022.8 ms p90 1304.1 p95 1306.1 p99 1308.2 input tok/question 316.2 rows/forward 54.0 in forward 98 % {200: 800}
30
+ VRAM_PEAK_MiB 21871
31
+ FIN 12:18:17