Comparison tables: BTZSC-22 with DiffusionGemma-26B-A4B add-on (Track A), speed/concurrency/GPU-memory run 3 (Ines-1, laya, GLiNER2.5-multi, Gemma)
Browse files- README.md +33 -11
- SHA256SUMS +11 -2
- eval/btzsc22/SHA256SUMS +6 -0
- eval/btzsc22/addon_gemma/README.md +39 -0
- eval/btzsc22/addon_gemma/code/run_gemma.py +36 -0
- eval/btzsc22/addon_gemma/code/score_addon_gemma.py +123 -0
- eval/btzsc22/addon_gemma/per_dataset.csv +23 -0
- eval/btzsc22/addon_gemma/results.json +2044 -0
- eval/btzsc22/addon_gemma/trackA_native.gemma.jsonl +0 -0
- eval/speed.md +33 -0
- eval/speed_runs/2026-10-06_part1_gemma_laya.txt +40 -0
- eval/speed_runs/2026-10-06_part2_gliner_ines.txt +31 -0
README.md
CHANGED
|
@@ -36,7 +36,7 @@ datasets:
|
|
| 36 |
> `tokenizer/`; the documentation (`README.md` and the other `.md` files, `paper/`, `assets/`, `tasksource_license_audit.csv`); the
|
| 37 |
> training records (`training/*.json`); the model's own evaluation outputs in `eval/`; `SHA256SUMS`.
|
| 38 |
> - **MIT** ([LICENSE-CODE](LICENSE-CODE)): the code — `mini_v41/`, `mini_v41_jev/`, `scripts/`, `examples/`,
|
| 39 |
-
> `training/code/`, `eval/btzsc22/code/` — and the environment and build files: `Dockerfile`, `.dockerignore`, `requirements.txt`,
|
| 40 |
> `requirements.lock`, `.gitattributes`, `.gitignore`.
|
| 41 |
> - **Third-party terms, not relicensed:** the gold labels and teacher distributions of the public test sets that
|
| 42 |
> `eval/items/`, `eval/reference_rows.json` and `eval/btzsc22/` (BTZSC label texts and targets) include keep their
|
|
@@ -199,6 +199,21 @@ learning the teacher's quirks.
|
|
| 199 |
We evaluated Ines-1 on 22 public BTZSC classification datasets (100 examples per dataset), using an extension of the
|
| 200 |
jev-benchmarks protocol. Fifteen of the 22 datasets are absent from Ines-1's decision fine-tuning mixture.
|
| 201 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 202 |
Under the model-specific classification track, mean macro-F1 across datasets was 0.474 for Ines-1, 0.525 for Julia-1,
|
| 203 |
0.530 for laya-multilingual, and 0.537 for GLiNER2.5. Using a paired hierarchical bootstrap with datasets as the unit of
|
| 204 |
inference, the difference relative to Ines-1 was resolved for GLiNER2.5 (+0.063, 95% CI +0.014 to +0.107), but not for
|
|
@@ -244,19 +259,26 @@ Cost: 21,423 domain cases, 23,608 questions per epoch, 28.5M prompt tokens, 2.25
|
|
| 244 |
"Engram removed at inference" = every Engram module returns its input (gate 0) on this checkpoint, which was trained
|
| 245 |
with Engram: it measures this model's reliance on Engram, not Engram's contribution to training.
|
| 246 |
|
| 247 |
-
### Speed
|
|
|
|
|
|
|
|
|
|
| 248 |
|
| 249 |
-
|
| 250 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 251 |
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
|
| 255 |
-
| same, 144M / 322M encoders, 4 processes each | 12.7 / 21–26 ms | ~1,700–1,850 |
|
| 256 |
-
| typed-decisions case (5 questions, ~315 tokens each) — Ines-1 | 29.6 / 38.3 ms | ~540 |
|
| 257 |
-
| same, 144M / 322M encoders | 15–16 / 87–88 ms | ~760–1,110 |
|
| 258 |
|
| 259 |
-
|
|
|
|
|
|
|
|
|
|
| 260 |
|
| 261 |
## Training
|
| 262 |
|
|
|
|
| 36 |
> `tokenizer/`; the documentation (`README.md` and the other `.md` files, `paper/`, `assets/`, `tasksource_license_audit.csv`); the
|
| 37 |
> training records (`training/*.json`); the model's own evaluation outputs in `eval/`; `SHA256SUMS`.
|
| 38 |
> - **MIT** ([LICENSE-CODE](LICENSE-CODE)): the code — `mini_v41/`, `mini_v41_jev/`, `scripts/`, `examples/`,
|
| 39 |
+
> `training/code/`, `eval/btzsc22/code/`, `eval/btzsc22/addon_gemma/code/` — and the environment and build files: `Dockerfile`, `.dockerignore`, `requirements.txt`,
|
| 40 |
> `requirements.lock`, `.gitattributes`, `.gitignore`.
|
| 41 |
> - **Third-party terms, not relicensed:** the gold labels and teacher distributions of the public test sets that
|
| 42 |
> `eval/items/`, `eval/reference_rows.json` and `eval/btzsc22/` (BTZSC label texts and targets) include keep their
|
|
|
|
| 199 |
We evaluated Ines-1 on 22 public BTZSC classification datasets (100 examples per dataset), using an extension of the
|
| 200 |
jev-benchmarks protocol. Fifteen of the 22 datasets are absent from Ines-1's decision fine-tuning mixture.
|
| 201 |
|
| 202 |
+
| mean macro-F1 (Track A, native interface) | Ines-1 | Julia-1 | laya-multilingual | GLiNER2.5-multi | DiffusionGemma-26B-A4B¹ |
|
| 203 |
+
|---|---:|---:|---:|---:|---:|
|
| 204 |
+
| parameters (total / active per token) | 1.59B / 405M | 144M | 322M | 287M | ~26B / ~4B |
|
| 205 |
+
| **all 22 datasets** | 0.474 | 0.525 | 0.530 | 0.537 | **0.799** |
|
| 206 |
+
| 2 labels (13) | 0.562 | 0.583 | 0.722 | 0.627 | **0.897** |
|
| 207 |
+
| 3–20 labels (4) | 0.489 | 0.634 | 0.559 | 0.541 | **0.695** |
|
| 208 |
+
| 21–72 labels (5) | 0.234 | 0.287 | 0.008 | 0.297 | **0.628** |
|
| 209 |
+
| absent from Ines-1's decision FT mix (15) | 0.486 | 0.536 | 0.515 | 0.528 | **0.765** |
|
| 210 |
+
| difference vs Ines-1, all 22 (95 % CI) | — | +0.051 (−0.026, +0.126) | +0.056 (−0.040, +0.163) | +0.063 (+0.014, +0.107) | +0.326 (+0.253, +0.392) |
|
| 211 |
+
| native ECE, ≤ 20 labels (17 datasets) | 0.127 | 0.298 | 0.093 | 0.159 | 0.116 |
|
| 212 |
+
|
| 213 |
+
¹ Added on 2026-10-06, after the protocol freeze, on Track A only (FP8 weights served by OpenJev 0.4.0, `think: 0`,
|
| 214 |
+
same request): [`eval/btzsc22/addon_gemma/`](eval/btzsc22/addon_gemma/). The scorer reproduces the other four columns
|
| 215 |
+
exactly. It is a ~16× larger generalist model; speed and memory for the same models are under *Speed* below.
|
| 216 |
+
|
| 217 |
Under the model-specific classification track, mean macro-F1 across datasets was 0.474 for Ines-1, 0.525 for Julia-1,
|
| 218 |
0.530 for laya-multilingual, and 0.537 for GLiNER2.5. Using a paired hierarchical bootstrap with datasets as the unit of
|
| 219 |
inference, the difference relative to Ines-1 was resolved for GLiNER2.5 (+0.063, 95% CI +0.014 to +0.107), but not for
|
|
|
|
| 259 |
"Engram removed at inference" = every Engram module returns its input (gate 0) on this checkpoint, which was trained
|
| 260 |
with Engram: it measures this model's reliance on Engram, not Engram's contribution to training.
|
| 261 |
|
| 262 |
+
### Speed and memory
|
| 263 |
+
|
| 264 |
+
[eval/speed.md](eval/speed.md) has the full protocol and all runs. One H200 per model in turn (nothing else on it), HTTP
|
| 265 |
+
closed loop, same requests and load generator for every model (run 3, 2026-10-06):
|
| 266 |
|
| 267 |
+
| model | short request¹ c=1 p50 / p95 | short: max questions/s | typed case² c=1 p50 / p95 | typed: max questions/s | GPU memory idle / peak |
|
| 268 |
+
|---|---:|---:|---:|---:|---:|
|
| 269 |
+
| **Ines-1** (1 process, CUDA graphs) | **9.4 / 9.7 ms** | **1,896** | 30.6 / **40.5 ms** | 544 | 9.7 / 21.9 GB |
|
| 270 |
+
| laya-multilingual architecture (4 processes) | 13.0 / 25.9 ms | 1,848 | **20.6** / 89.2 ms | **840** | 8.0 / 24.1 GB |
|
| 271 |
+
| GLiNER2.5-multi architecture (4 workers, no batching) | 48.8 / 49.3 ms | 68 | 123.1 / 129.3 ms | 49 | 6.7 / 26.5 GB |
|
| 272 |
+
| DiffusionGemma-26B-A4B FP8 (OpenJev / vLLM) | 91.7 / 94.0 ms | 79 | 158.8 / 170.2 ms | 143 | 25.8 GiB weights; 120.6 / 135.1 GB reserved³ |
|
| 273 |
|
| 274 |
+
¹ 2 questions, ~60 prompt tokens each (Ines-1 template). ² typed-decisions case, 5 questions, ~300 prompt tokens each
|
| 275 |
+
(Ines-1 template; each server counts tokens differently, see eval/speed.md).
|
| 276 |
+
³ vLLM preallocates KV cache (`gpu_memory_utilization` 0.80); the floor is the weights.
|
|
|
|
|
|
|
|
|
|
| 277 |
|
| 278 |
+
Ines-1 has the lowest single-request latency and matches the encoder's peak short-request throughput (within the
|
| 279 |
+
±10 % run-to-run variation); on ~300-token questions the
|
| 280 |
+
322M encoder is faster (1.5× the throughput, lower median) while Ines-1 keeps the tighter tail. The GLiNER row measures
|
| 281 |
+
a server without batching, not the architecture's ceiling. Different serving stacks: not a general speed ranking.
|
| 282 |
|
| 283 |
## Training
|
| 284 |
|
SHA256SUMS
CHANGED
|
@@ -8,7 +8,7 @@ cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
|
|
| 8 |
6a7ddcad077839a44697c005ea7e6f455caecaa8054412913108d6e21a1c9dbe LICENSE-CODE
|
| 9 |
121b31536257bd7147d9c9b054f6fb4e2bb68734024415c081ac0b50c7cf48f1 LICENSING_NOTES.md
|
| 10 |
1faada60db1ccfc073ae063fc474d9f75cba0c9d2683d03ad66b2074efeb166a NOTICE
|
| 11 |
-
|
| 12 |
6ec8a2bb55f9331d77b9adbeae7fd71eeb5614a988cd21eee18c7bde775df5ba TASKSOURCE_LICENSE_AUDIT.md
|
| 13 |
1a4df8f10bd30b01e6acc19ef2d935faee5485f0c25452ff64ba6c8598fabd97 THIRD_PARTY_DATA_NOTICE.md
|
| 14 |
9708385cf095b543e009b0dd6013aa9985ef68a5b05c5ebef4c14c6a7f97cfa2 assets/architecture.png
|
|
@@ -16,6 +16,13 @@ cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE
|
|
| 16 |
757003b29ec1d935f236c06ea1e39cf9362c43de358f8b3df74326f8bfdbd40d config.json
|
| 17 |
8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b eval/btzsc22/PROTOCOL.md
|
| 18 |
5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 eval/btzsc22/REPORT.md
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
95a4d4c459940195f03c6b937d750cd8e678527c4e2acaee1f7f2f3a9adc512c eval/btzsc22/btzsc_test_files.json
|
| 20 |
bb7dbba94cc130a09c77226e51cb7b72e5228685fbf8c2f78522290f485eff1e eval/btzsc22/code/README.md
|
| 21 |
afac77fa429f2f37c45e6cf46d0201551e9c98ddf9c01e7ba0cf1ac1feca0ef7 eval/btzsc22/code/btzsc22.yaml
|
|
@@ -57,8 +64,10 @@ be0892eb9d9e45f2173e59052d503b26d809253b01bc44a8a70c0e7c48add480 eval/items/v3_
|
|
| 57 |
d0b5cb709b104e8fafcdba8f8c676896ce048c2a009f93c1b19ad3e949fcb128 eval/reference_rows.json
|
| 58 |
508a3e33bf77cd4ab9f559d5dcd2ec3e473c30b3b1d54c90a7de5e6490a435a2 eval/report.json
|
| 59 |
d5f3f5d214ac20264176a552c7779ddbbee79f97524b64f8c6d1804e1f305ba2 eval/report.md
|
| 60 |
-
|
| 61 |
d0b6fcd59367eb58c02eb3b21e7d6c50caa6eccf5d822e8e7ecec22f0e3a275b eval/speed_raw.log
|
|
|
|
|
|
|
| 62 |
22974cc87fd92b5c21168e350ff3db492209a98293ec28a7f5d70396f13c970d examples/quickstart.py
|
| 63 |
6c4fc03dc8ec311323338f6f4f6c0a8a3514f417b57c99f70827d797fe2ead28 mini_v41/__init__.py
|
| 64 |
7c57633d0d69fe8a709e0629cbe36b07590d94d0333039c5378fa66df4a74f05 mini_v41/attention.py
|
|
|
|
| 8 |
6a7ddcad077839a44697c005ea7e6f455caecaa8054412913108d6e21a1c9dbe LICENSE-CODE
|
| 9 |
121b31536257bd7147d9c9b054f6fb4e2bb68734024415c081ac0b50c7cf48f1 LICENSING_NOTES.md
|
| 10 |
1faada60db1ccfc073ae063fc474d9f75cba0c9d2683d03ad66b2074efeb166a NOTICE
|
| 11 |
+
aed9b754a211decd074a0ddccd654852fc3797a04376bc80b882ea47cc200f58 README.md
|
| 12 |
6ec8a2bb55f9331d77b9adbeae7fd71eeb5614a988cd21eee18c7bde775df5ba TASKSOURCE_LICENSE_AUDIT.md
|
| 13 |
1a4df8f10bd30b01e6acc19ef2d935faee5485f0c25452ff64ba6c8598fabd97 THIRD_PARTY_DATA_NOTICE.md
|
| 14 |
9708385cf095b543e009b0dd6013aa9985ef68a5b05c5ebef4c14c6a7f97cfa2 assets/architecture.png
|
|
|
|
| 16 |
757003b29ec1d935f236c06ea1e39cf9362c43de358f8b3df74326f8bfdbd40d config.json
|
| 17 |
8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b eval/btzsc22/PROTOCOL.md
|
| 18 |
5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 eval/btzsc22/REPORT.md
|
| 19 |
+
775c02f424c71a5b944ccaf19ba466ad1444644e585cf4785e9fe39d1a0e6140 eval/btzsc22/SHA256SUMS
|
| 20 |
+
cc8ce27ca26efcdc9985323e75fdbad20b602d69b1602c2883a1c80d37dbe15e eval/btzsc22/addon_gemma/README.md
|
| 21 |
+
80bc6a7b10e75160f88265611a7be4fc07d625980a55697cee3ff3a60dfc5632 eval/btzsc22/addon_gemma/code/run_gemma.py
|
| 22 |
+
d4a2cb890861b2076e79934b5a1dd0f184c5b4418dec4916c8288def58a279e9 eval/btzsc22/addon_gemma/code/score_addon_gemma.py
|
| 23 |
+
7107cfe4a47f306829ee4e30fa50ddeecc9f18f107c5eb300657877b3ce9aded eval/btzsc22/addon_gemma/per_dataset.csv
|
| 24 |
+
2eca0b816e2ba974b31fd25043ab7ea38b39306f5ed97b16091a6905098a4a2f eval/btzsc22/addon_gemma/results.json
|
| 25 |
+
5e853fc07f7ab18b695be6b356f8d2c1e6ae23a61e3d79fed0f3be310fe3e26f eval/btzsc22/addon_gemma/trackA_native.gemma.jsonl
|
| 26 |
95a4d4c459940195f03c6b937d750cd8e678527c4e2acaee1f7f2f3a9adc512c eval/btzsc22/btzsc_test_files.json
|
| 27 |
bb7dbba94cc130a09c77226e51cb7b72e5228685fbf8c2f78522290f485eff1e eval/btzsc22/code/README.md
|
| 28 |
afac77fa429f2f37c45e6cf46d0201551e9c98ddf9c01e7ba0cf1ac1feca0ef7 eval/btzsc22/code/btzsc22.yaml
|
|
|
|
| 64 |
d0b5cb709b104e8fafcdba8f8c676896ce048c2a009f93c1b19ad3e949fcb128 eval/reference_rows.json
|
| 65 |
508a3e33bf77cd4ab9f559d5dcd2ec3e473c30b3b1d54c90a7de5e6490a435a2 eval/report.json
|
| 66 |
d5f3f5d214ac20264176a552c7779ddbbee79f97524b64f8c6d1804e1f305ba2 eval/report.md
|
| 67 |
+
2845d32914cc266a9b8a14f20b9ea7b37ad6febf2ae52a084c46ba51af140165 eval/speed.md
|
| 68 |
d0b6fcd59367eb58c02eb3b21e7d6c50caa6eccf5d822e8e7ecec22f0e3a275b eval/speed_raw.log
|
| 69 |
+
da6b8271d68d79a13d4e79e5c5e90e6722d91754bf15e797670a43a6120e99d0 eval/speed_runs/2026-10-06_part1_gemma_laya.txt
|
| 70 |
+
34b9bd9bf06e21f67ad73fc13e8d0deb49efa95a9f181d820b6eced9aa133d6d eval/speed_runs/2026-10-06_part2_gliner_ines.txt
|
| 71 |
22974cc87fd92b5c21168e350ff3db492209a98293ec28a7f5d70396f13c970d examples/quickstart.py
|
| 72 |
6c4fc03dc8ec311323338f6f4f6c0a8a3514f417b57c99f70827d797fe2ead28 mini_v41/__init__.py
|
| 73 |
7c57633d0d69fe8a709e0629cbe36b07590d94d0333039c5378fa66df4a74f05 mini_v41/attention.py
|
eval/btzsc22/SHA256SUMS
CHANGED
|
@@ -22,3 +22,9 @@ d0ca40988b8ad697223f7ccfc18ac35840c1390a5d46fabc219dab0c02069ed3 ./predictions/
|
|
| 22 |
8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b ./PROTOCOL.md
|
| 23 |
5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 ./REPORT.md
|
| 24 |
707f23cd6b4ae30a38573d4f29ce852a82878b23ee062c61a8025d4bae1360ff ./results.json
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
8f296d2c1dbe85d285bb45cd80d841e733f84168ff130daf825933917684993b ./PROTOCOL.md
|
| 23 |
5cc7f2cd08668abe6ac5dd5455e61bc2315e556e26c45906150426810fd49455 ./REPORT.md
|
| 24 |
707f23cd6b4ae30a38573d4f29ce852a82878b23ee062c61a8025d4bae1360ff ./results.json
|
| 25 |
+
cc8ce27ca26efcdc9985323e75fdbad20b602d69b1602c2883a1c80d37dbe15e ./addon_gemma/README.md
|
| 26 |
+
80bc6a7b10e75160f88265611a7be4fc07d625980a55697cee3ff3a60dfc5632 ./addon_gemma/code/run_gemma.py
|
| 27 |
+
d4a2cb890861b2076e79934b5a1dd0f184c5b4418dec4916c8288def58a279e9 ./addon_gemma/code/score_addon_gemma.py
|
| 28 |
+
7107cfe4a47f306829ee4e30fa50ddeecc9f18f107c5eb300657877b3ce9aded ./addon_gemma/per_dataset.csv
|
| 29 |
+
2eca0b816e2ba974b31fd25043ab7ea38b39306f5ed97b16091a6905098a4a2f ./addon_gemma/results.json
|
| 30 |
+
5e853fc07f7ab18b695be6b356f8d2c1e6ae23a61e3d79fed0f3be310fe3e26f ./addon_gemma/trackA_native.gemma.jsonl
|
eval/btzsc22/addon_gemma/README.md
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# BTZSC-22 add-on: DiffusionGemma-26B-A4B-it (2026-10-06)
|
| 2 |
+
|
| 3 |
+
**Added after the `btzsc22-ext-v1` freeze.** The protocol, the manifest (2,200 examples) and the scorer are unchanged; this
|
| 4 |
+
directory adds one more model on **Track A (native) only**. Track B (one-vs-rest) was not run for it.
|
| 5 |
+
|
| 6 |
+
| | |
|
| 7 |
+
|---|---|
|
| 8 |
+
| model | `RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic` (FP8 of Google's DiffusionGemma-26B-A4B-it) |
|
| 9 |
+
| serving | `razorback16/openjev:0.4.0` (vLLM structured reads, `/v1/systemone`), model `openjev-0.1`, `think: 0` (the server default), canvas 64, 1× H200 |
|
| 10 |
+
| request | identical to the other models: `state = {"text": <text>}`, one Choice question `"Which single label best describes the input text?"`, options `label_000…` with the label texts as descriptions |
|
| 11 |
+
| interface limit | 255 options, so all 22 datasets are a **single request** (no knockout) |
|
| 12 |
+
| failures | 0 of 2,200 |
|
| 13 |
+
|
| 14 |
+
## Files
|
| 15 |
+
|
| 16 |
+
- `trackA_native.gemma.jsonl`: one prediction per example (full native probability vector). `latency_seconds` was
|
| 17 |
+
recorded under 16 concurrent clients and is **not** a single-request latency (see `../../speed.md`).
|
| 18 |
+
- `code/run_gemma.py`: the client used.
|
| 19 |
+
- `code/score_addon_gemma.py`: `../code/score_final.py` with Gemma added to the model list. The Track A logic,
|
| 20 |
+
the paired hierarchical bootstrap and the calibration code are copied verbatim. Rerunning it reproduces the four
|
| 21 |
+
original models' numbers in `../results.json` exactly (0.474 / 0.525 / 0.530 / 0.537).
|
| 22 |
+
- `results.json`, `per_dataset.csv`: the output. Gemma's Track B fields are `null` / `not_run`.
|
| 23 |
+
|
| 24 |
+
## Result (Track A, equal-weight mean macro-F1 over datasets)
|
| 25 |
+
|
| 26 |
+
| group | Ines-1 | Julia-1 | laya-multilingual | GLiNER2.5-multi | DiffusionGemma-26B-A4B |
|
| 27 |
+
|---|---:|---:|---:|---:|---:|
|
| 28 |
+
| all 22 | 0.474 | 0.525 | 0.530 | 0.537 | **0.799** |
|
| 29 |
+
| 2 labels (13) | 0.562 | 0.583 | 0.722 | 0.627 | **0.897** |
|
| 30 |
+
| 3–20 labels (4) | 0.489 | 0.634 | 0.559 | 0.541 | **0.695** |
|
| 31 |
+
| 21–72 labels (5) | 0.234 | 0.287 | 0.008 | 0.297 | **0.628** |
|
| 32 |
+
| absent from Ines-1's decision FT mix (15) | 0.486 | 0.536 | 0.515 | 0.528 | **0.765** |
|
| 33 |
+
|
| 34 |
+
Gemma − Ines-1, all 22: **+0.326** (95 % CI +0.253 to +0.392, paired hierarchical bootstrap).
|
| 35 |
+
Native calibration (17 datasets with ≤ 20 labels): ECE 0.116 (Ines-1 0.127), Brier 0.250 (Ines-1 0.501),
|
| 36 |
+
coverage at 5 % error budget 0.694 (Ines-1 0.089).
|
| 37 |
+
|
| 38 |
+
DiffusionGemma-26B-A4B has ~26B total parameters (~4B active per token), about 16× Ines-1's total and 10× its active
|
| 39 |
+
parameters. Speed and memory for the same models are in `../../speed.md` (run 3).
|
eval/btzsc22/addon_gemma/code/run_gemma.py
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""BTZSC-22 backend for DiffusionGemma-26B-A4B-it (FP8) served by OpenJev 0.4.0 (/v1/systemone).
|
| 2 |
+
Same request as run_model.py (state={"text": text}, one Choice, options label_000...), think 0 (OpenJev default),
|
| 3 |
+
native mode only (OpenJev allows 255 options, every dataset fits). argv: url manifest out.jsonl
|
| 4 |
+
"""
|
| 5 |
+
import json, os, sys, time, urllib.request
|
| 6 |
+
from concurrent.futures import ThreadPoolExecutor
|
| 7 |
+
URL, MAN, OUT = sys.argv[1:4]
|
| 8 |
+
Q = "Which single label best describes the input text?"
|
| 9 |
+
|
| 10 |
+
def call(text, labels):
|
| 11 |
+
crit = {"label_%03d" % i: l for i, l in enumerate(labels)}
|
| 12 |
+
body = {"model": "openjev-0.1", "state": {"text": text}, "questions": {"label": {"type": "choice", "instructions": Q, "criteria": crit}}}
|
| 13 |
+
req = urllib.request.Request(URL, json.dumps(body).encode(), {"Content-Type": "application/json"})
|
| 14 |
+
with urllib.request.urlopen(req, timeout=300) as r:
|
| 15 |
+
a = json.loads(r.read())["answers"]["label"]
|
| 16 |
+
return [float(a["probabilities"][k]) for k in crit]
|
| 17 |
+
|
| 18 |
+
def one(e):
|
| 19 |
+
t = time.perf_counter()
|
| 20 |
+
try:
|
| 21 |
+
p = call(e["text"], e["labels"]); pred = max(range(len(p)), key=p.__getitem__); err = None
|
| 22 |
+
except Exception as x: # failures stay in the denominator
|
| 23 |
+
p, pred, err = None, -1, repr(x)[:300]
|
| 24 |
+
return {"backend": "gemma", "mode": "native", "group": 0, "dataset": e["dataset"], "example_id": e["example_id"],
|
| 25 |
+
"target_index": e["target_index"], "predicted_index": pred, "n_labels": len(e["labels"]), "probabilities": p,
|
| 26 |
+
"calls": 1, "latency_seconds": time.perf_counter() - t, "error": err}
|
| 27 |
+
|
| 28 |
+
done = set()
|
| 29 |
+
if os.path.exists(OUT):
|
| 30 |
+
done = {json.loads(l)["example_id"] for l in open(OUT)}
|
| 31 |
+
ex = [e for e in map(json.loads, open(MAN)) if e["example_id"] not in done]
|
| 32 |
+
call(ex[0]["text"], ex[0]["labels"][:2]) # warm-up
|
| 33 |
+
with open(OUT, "a") as f, ThreadPoolExecutor(16) as pool:
|
| 34 |
+
for r in pool.map(one, ex):
|
| 35 |
+
f.write(json.dumps(r) + "\n"); f.flush()
|
| 36 |
+
print("done gemma native")
|
eval/btzsc22/addon_gemma/code/score_addon_gemma.py
ADDED
|
@@ -0,0 +1,123 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""ADD-ON (6-oct, after the freeze): score_final.py + DiffusionGemma-26B-A4B-it (OpenJev) on Track A only. Track A logic, bootstrap and calibration copied verbatim; Track B not run for Gemma.
|
| 2 |
+
Original: btzsc22-ext-v1 analysis (frozen before reading Track B). argv: run_dir(v1) exploratory_dir out_dir
|
| 3 |
+
Run with the gliner venv (uses jev_benchmarks.metrics unchanged)."""
|
| 4 |
+
import csv, json, sys
|
| 5 |
+
from pathlib import Path
|
| 6 |
+
import numpy as np
|
| 7 |
+
from jev_benchmarks.metrics import score_predictions, _macro_f1
|
| 8 |
+
from jev_benchmarks.models import Prediction
|
| 9 |
+
|
| 10 |
+
V1, EXP, OUT = Path(sys.argv[1]), Path(sys.argv[2]), Path(sys.argv[3]); OUT.mkdir(parents=True, exist_ok=True)
|
| 11 |
+
MODELS = ["ines", "julia", "laya", "gliner", "gemma"]
|
| 12 |
+
GEMMA = Path(__file__).resolve().parent / "runs/btzsc22/addon_gemma/gemma.native.jsonl"
|
| 13 |
+
NAME = {"ines": "Ines-1", "julia": "Julia-1", "laya": "laya-multilingual", "gliner": "GLiNER2.5", "gemma": "DiffusionGemma-26B-A4B-it"}
|
| 14 |
+
IN_MIX = {"amazonpolarity", "banking77", "massive", "wikitoxic_insult", "wikitoxic_obscene", "wikitoxic_threat", "wikitoxic_toxicaggregated"}
|
| 15 |
+
man = [json.loads(l) for l in open(EXP / "manifest.jsonl")]
|
| 16 |
+
DS = list(dict.fromkeys(e["dataset"] for e in man)); K = {e["dataset"]: len(e["labels"]) for e in man}
|
| 17 |
+
IDS = {d: [e["example_id"] for e in man if e["dataset"] == d] for d in DS}
|
| 18 |
+
T = {d: np.array([e["target_index"] for e in man if e["dataset"] == d]) for d in DS}
|
| 19 |
+
LAB = {e["example_id"]: tuple(e["labels"]) for e in man}
|
| 20 |
+
rd = lambda p: {json.loads(l)["example_id"]: json.loads(l) for l in open(p)} if Path(p).exists() else {}
|
| 21 |
+
native = {m: rd(V1 / ("%s.native.jsonl" % m)) for m in MODELS if m != "gemma"}; native["gemma"] = rd(GEMMA)
|
| 22 |
+
ovr = {m: rd(V1 / ("%s.ovr.jsonl" % m)) for m in MODELS}
|
| 23 |
+
ines_k20, ines_k2 = rd(EXP / "ines.k20.jsonl"), rd(EXP / "ines.k2.jsonl")
|
| 24 |
+
|
| 25 |
+
def preds(src, d):
|
| 26 |
+
return np.array([(src[i]["predicted_index"] if i in src and not src[i].get("error") else -1) for i in IDS[d]])
|
| 27 |
+
def trackA(m, d):
|
| 28 |
+
if m == "ines" and K[d] > 26:
|
| 29 |
+
return preds(ines_k20, d), "knockout-20 (no native > 26 interface)"
|
| 30 |
+
r = native[m]; s = "single request"
|
| 31 |
+
if m == "julia" and K[d] > 20: s = "Router(width=20, survivors=2)"
|
| 32 |
+
return preds(r, d), s
|
| 33 |
+
P = {"A": {m: {d: trackA(m, d)[0] for d in DS} for m in MODELS}, "B": {m: {d: preds(ovr[m], d) for d in DS} for m in MODELS}}
|
| 34 |
+
missing = {tr: {m: int(sum((P[tr][m][d] < 0).sum() for d in DS)) for m in MODELS} for tr in P}
|
| 35 |
+
f1 = lambda tr, m, d, idx=None: _macro_f1(T[d] if idx is None else T[d][idx], P[tr][m][d] if idx is None else P[tr][m][d][idx], K[d])
|
| 36 |
+
acc = lambda tr, m, d: float(np.mean(P[tr][m][d] == T[d]))
|
| 37 |
+
|
| 38 |
+
GROUPS = {"all 22": DS, "2 labels (13)": [d for d in DS if K[d] == 2], "3-20 labels (4)": [d for d in DS if 2 < K[d] <= 20],
|
| 39 |
+
"21-72 labels (5)": [d for d in DS if K[d] > 20], "absent from Ines-1 decision FT mixture (15)": [d for d in DS if d not in IN_MIX],
|
| 40 |
+
"families in Ines-1 decision FT mixture (7)": [d for d in DS if d in IN_MIX]}
|
| 41 |
+
rng = np.random.default_rng(20260917)
|
| 42 |
+
STRATA = {d: [np.where(T[d] == c)[0] for c in np.unique(T[d])] for d in DS}
|
| 43 |
+
def hboot(sel, pairs, B=2000):
|
| 44 |
+
"""paired hierarchical bootstrap: resample datasets, then examples stratified by target; equal-weight mean macro-F1"""
|
| 45 |
+
out = {p: [] for p in pairs}
|
| 46 |
+
for _ in range(B):
|
| 47 |
+
draw = rng.integers(0, len(sel), len(sel))
|
| 48 |
+
acc_ = {p: [] for p in pairs}
|
| 49 |
+
for j in draw:
|
| 50 |
+
d = sel[j]; idx = np.concatenate([rng.choice(c, len(c)) for c in STRATA[d]])
|
| 51 |
+
cache = {}
|
| 52 |
+
for p in pairs:
|
| 53 |
+
(ta, ma), (tb, mb) = p
|
| 54 |
+
for key in ((ta, ma), (tb, mb)):
|
| 55 |
+
if key not in cache: cache[key] = f1(key[0], key[1], d, idx)
|
| 56 |
+
acc_[p].append(cache[(tb, mb)] - cache[(ta, ma)])
|
| 57 |
+
for p in pairs: out[p].append(np.mean(acc_[p]))
|
| 58 |
+
res = {}
|
| 59 |
+
for p, v in out.items():
|
| 60 |
+
(ta, ma), (tb, mb) = p
|
| 61 |
+
obs = float(np.mean([f1(tb, mb, d) - f1(ta, ma, d) for d in sel]))
|
| 62 |
+
res["%s:%s - %s:%s" % (tb, mb, ta, ma)] = {"difference": obs, "ci95": [float(np.quantile(v, .025)), float(np.quantile(v, .975))]}
|
| 63 |
+
return res
|
| 64 |
+
|
| 65 |
+
R = {"protocol": "btzsc22-ext-v1", "missing_predictions": missing, "per_dataset": {}, "global": {}, "bootstrap": {}, "calibration": {},
|
| 66 |
+
"ines_ablation": {}, "anchors": {}, "ovr_ties": {}}
|
| 67 |
+
for d in DS:
|
| 68 |
+
R["per_dataset"][d] = {"labels": K[d], "in_ines_ft_mixture": d in IN_MIX,
|
| 69 |
+
**{tr: {m: {"accuracy": acc(tr, m, d), "macro_f1": f1(tr, m, d)} for m in MODELS} for tr in ("A", "B")},
|
| 70 |
+
"trackA_strategy": {m: trackA(m, d)[1] for m in MODELS}}
|
| 71 |
+
for g, sel in GROUPS.items():
|
| 72 |
+
R["global"][g] = {tr: {m: float(np.mean([f1(tr, m, d) for d in sel])) for m in MODELS} for tr in ("A", "B")}
|
| 73 |
+
cross = [(("A", "ines"), ("A", m)) for m in MODELS[1:]] + [(("B", "ines"), ("B", m)) for m in MODELS[1:]] + [(("A", m), ("B", m)) for m in MODELS]
|
| 74 |
+
for g, sel in GROUPS.items():
|
| 75 |
+
R["bootstrap"][g] = hboot(sel, cross, B=2000 if g == "all 22" else 1000)
|
| 76 |
+
|
| 77 |
+
def calib(srcs, field, sel):
|
| 78 |
+
out = {}
|
| 79 |
+
for m in MODELS:
|
| 80 |
+
per = []
|
| 81 |
+
for d in sel:
|
| 82 |
+
rows = []
|
| 83 |
+
for i in IDS[d]:
|
| 84 |
+
r = srcs[m].get(i)
|
| 85 |
+
if not r or r.get("error") or not r.get(field): break
|
| 86 |
+
rows.append(Prediction(experiment_id="x", backend=m, model_requested=m, model_resolved=m, dataset=d, example_id=i,
|
| 87 |
+
target_index=int(T[d][IDS[d].index(i)]), predicted_index=r["predicted_index"], labels=LAB[i],
|
| 88 |
+
probabilities=tuple(r[field]), latency_seconds=0.0))
|
| 89 |
+
else:
|
| 90 |
+
per.append(score_predictions(rows))
|
| 91 |
+
if per:
|
| 92 |
+
out[m] = {k: float(np.mean([s[k] for s in per])) for k in ("brier", "nll", "ece", "coverage_at_error_budget", "mean_confidence")} | {"datasets": len(per)}
|
| 93 |
+
return out
|
| 94 |
+
native_full = [d for d in DS if K[d] <= 20]
|
| 95 |
+
R["calibration"]["native (Track A, full native distributions, <=20 labels, 17 datasets)"] = calib(native, "probabilities", native_full)
|
| 96 |
+
R["calibration"]["derived (Track B softmax of log-odds, same 17 datasets)"] = calib(ovr, "derived_probabilities", native_full)
|
| 97 |
+
R["calibration"]["derived (Track B softmax of log-odds, all 22 datasets)"] = calib(ovr, "derived_probabilities", DS)
|
| 98 |
+
R["ovr_ties"] = {m: {"examples_with_tie_at_max": sum(1 for r in ovr[m].values() if r.get("ties_at_max", 1) > 1), "n": len(ovr[m])} for m in MODELS}
|
| 99 |
+
|
| 100 |
+
multi = [d for d in DS if K[d] > 2]
|
| 101 |
+
for d in multi:
|
| 102 |
+
row = {"native(<=26)": acc("A", "ines", d) if K[d] <= 26 else None, "knockout-20": float(np.mean(preds(ines_k20, d) == T[d])) if K[d] > 20 else None,
|
| 103 |
+
"knockout-2": float(np.mean(preds(ines_k2, d) == T[d])), "one-vs-rest": acc("B", "ines", d)}
|
| 104 |
+
rowf = {"native(<=26)": _macro_f1(T[d], preds(native["ines"], d), K[d]) if K[d] <= 26 else None,
|
| 105 |
+
"knockout-20": _macro_f1(T[d], preds(ines_k20, d), K[d]) if K[d] > 20 else None,
|
| 106 |
+
"knockout-2": _macro_f1(T[d], preds(ines_k2, d), K[d]), "one-vs-rest": f1("B", "ines", d)}
|
| 107 |
+
R["ines_ablation"][d] = {"labels": K[d], "accuracy": row, "macro_f1": rowf}
|
| 108 |
+
|
| 109 |
+
pub = {"agnews": 0.70, "emotiondair": 0.44, "banking77": 0.61}
|
| 110 |
+
R["anchors"]["GLiNER2.5 Track A vs published pilot (accuracy)"] = {d: {"ours": acc("A", "gliner", d), "published": v} for d, v in pub.items()}
|
| 111 |
+
card = {"agnews": 0.94, "emotiondair": 0.86, "banking77": 0.64}
|
| 112 |
+
R["anchors"]["Julia-1 Track A vs its model card (accuracy)"] = {d: {"ours": acc("A", "julia", d), "card": v, "ours_strategy": trackA("julia", d)[1]} for d, v in card.items()}
|
| 113 |
+
|
| 114 |
+
json.dump(R, open(OUT / "results.json", "w"), indent=1)
|
| 115 |
+
with open(OUT / "per_dataset.csv", "w", newline="") as f:
|
| 116 |
+
w = csv.writer(f); w.writerow(["dataset", "labels", "in_ines1_decision_ft_mixture"] + ["%s_%s_%s" % (tr, m, k) for tr in ("A", "B") for m in MODELS for k in ("accuracy", "macro_f1")] + ["A_strategy_ines", "A_strategy_julia"])
|
| 117 |
+
for d in DS:
|
| 118 |
+
x = R["per_dataset"][d]
|
| 119 |
+
w.writerow([d, K[d], d in IN_MIX] + ["%.4f" % x[tr][m][k] for tr in ("A", "B") for m in MODELS for k in ("accuracy", "macro_f1")] + [x["trackA_strategy"]["ines"], x["trackA_strategy"]["julia"]])
|
| 120 |
+
print("missing", missing)
|
| 121 |
+
for g in GROUPS:
|
| 122 |
+
print("%-45s A %s | B %s" % (g, " ".join("%s %.3f" % (m, R["global"][g]["A"][m]) for m in MODELS), " ".join("%s %.3f" % (m, R["global"][g]["B"][m]) for m in MODELS)))
|
| 123 |
+
for k, v in R["bootstrap"]["all 22"].items(): print(" %-28s %+.3f [%+.3f, %+.3f]" % (k, v["difference"], *v["ci95"]))
|
eval/btzsc22/addon_gemma/per_dataset.csv
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
dataset,labels,in_ines1_decision_ft_mixture,A_ines_accuracy,A_ines_macro_f1,A_julia_accuracy,A_julia_macro_f1,A_laya_accuracy,A_laya_macro_f1,A_gliner_accuracy,A_gliner_macro_f1,A_gemma_accuracy,A_gemma_macro_f1,B_ines_accuracy,B_ines_macro_f1,B_julia_accuracy,B_julia_macro_f1,B_laya_accuracy,B_laya_macro_f1,B_gliner_accuracy,B_gliner_macro_f1,B_gemma_accuracy,B_gemma_macro_f1,A_strategy_ines,A_strategy_julia
|
| 2 |
+
agnews,4,False,0.5900,0.5843,0.9400,0.9399,0.9400,0.9393,0.7000,0.6589,0.8900,0.8797,0.4200,0.3265,0.4200,0.3995,0.8900,0.8890,0.4500,0.3986,not_run,not_run,single request,single request
|
| 3 |
+
emotiondair,6,False,0.4300,0.3904,0.8600,0.8557,0.3700,0.3527,0.4400,0.4070,0.4000,0.3764,0.3400,0.2948,0.2600,0.2543,0.4300,0.4279,0.3500,0.3193,not_run,not_run,single request,single request
|
| 4 |
+
banking77,72,True,0.3900,0.3427,0.6300,0.5704,0.0200,0.0006,0.6100,0.5693,0.8000,0.7773,0.2900,0.2235,0.0900,0.0472,0.2200,0.1880,0.4100,0.3788,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
|
| 5 |
+
amazonpolarity,2,True,0.7800,0.7756,0.6200,0.6162,0.8400,0.8368,0.8800,0.8800,0.9800,0.9800,0.6500,0.6179,0.5800,0.5716,0.9100,0.9098,0.8300,0.8249,not_run,not_run,single request,single request
|
| 6 |
+
imdb,2,False,0.7500,0.7500,0.6800,0.6799,0.7300,0.7149,0.8000,0.7947,0.9700,0.9700,0.6900,0.6808,0.5400,0.5354,0.8100,0.8091,0.7300,0.7175,not_run,not_run,single request,single request
|
| 7 |
+
appreviews,2,False,0.7700,0.7660,0.6500,0.6396,0.7800,0.7756,0.8900,0.8895,0.9400,0.9399,0.8100,0.8095,0.5100,0.4954,0.8900,0.8897,0.8000,0.7960,not_run,not_run,single request,single request
|
| 8 |
+
yelpreviews,2,False,0.7200,0.6997,0.7300,0.7278,0.9000,0.8996,0.9000,0.8994,0.9600,0.9599,0.6800,0.6435,0.5100,0.5088,0.9300,0.9298,0.7400,0.7211,not_run,not_run,single request,single request
|
| 9 |
+
rottentomatoes,2,False,0.5900,0.5671,0.6600,0.6264,0.6900,0.6808,0.7100,0.7064,0.9100,0.9100,0.6400,0.6347,0.4600,0.4325,0.7400,0.7400,0.6700,0.6548,not_run,not_run,single request,single request
|
| 10 |
+
financialphrasebank,3,False,0.5100,0.4891,0.4500,0.3711,0.8100,0.8085,0.6900,0.6821,0.8000,0.8030,0.4500,0.3886,0.2500,0.1744,0.9100,0.9097,0.5500,0.4785,not_run,not_run,single request,single request
|
| 11 |
+
empathetic,32,False,0.2500,0.1909,0.1700,0.1594,0.0600,0.0255,0.3300,0.2807,0.5100,0.4697,0.1800,0.1442,0.0500,0.0390,0.3300,0.2873,0.1400,0.1304,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
|
| 12 |
+
biasframes_intent,2,False,0.6000,0.5998,0.5700,0.5665,0.6100,0.5920,0.5800,0.5739,0.8300,0.8300,0.5900,0.5524,0.5900,0.5806,0.6600,0.6550,0.5200,0.3912,not_run,not_run,single request,single request
|
| 13 |
+
massive,59,True,0.3500,0.2502,0.2700,0.1802,0.0100,0.0010,0.4400,0.3827,0.7800,0.7266,0.2000,0.1528,0.0200,0.0122,0.3500,0.3255,0.3000,0.2354,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
|
| 14 |
+
yahootopics,10,False,0.5000,0.4915,0.3900,0.3675,0.1500,0.1369,0.4100,0.4170,0.7300,0.7218,0.1800,0.1231,0.1200,0.1127,0.3700,0.3650,0.3200,0.3175,not_run,not_run,single request,single request
|
| 15 |
+
trueteacher,2,False,0.5600,0.5376,0.5200,0.5152,0.5600,0.4658,0.5300,0.5242,0.7100,0.6966,0.5000,0.3333,0.5000,0.4346,0.4800,0.4207,0.5400,0.4524,not_run,not_run,single request,single request
|
| 16 |
+
manifesto,56,False,0.1100,0.0750,0.1000,0.0789,0.0300,0.0077,0.0200,0.0007,0.4400,0.3767,0.0500,0.0329,0.0200,0.0149,0.1100,0.0903,0.0600,0.0136,not_run,not_run,knockout-20 (no native > 26 interface),"Router(width=20, survivors=2)"
|
| 17 |
+
capsotu,21,False,0.3600,0.3121,0.4500,0.4452,0.0500,0.0050,0.3000,0.2524,0.7900,0.7913,0.2100,0.1784,0.1000,0.0799,0.2700,0.2800,0.1700,0.1837,not_run,not_run,single request,"Router(width=20, survivors=2)"
|
| 18 |
+
biasframes_offensive,2,False,0.5200,0.4956,0.5400,0.5119,0.6200,0.6194,0.4600,0.4375,0.8200,0.8199,0.4800,0.4800,0.4500,0.4068,0.6900,0.6885,0.5200,0.3763,not_run,not_run,single request,single request
|
| 19 |
+
biasframes_sex,2,False,0.3400,0.3389,0.5600,0.5556,0.7200,0.7083,0.5300,0.3967,0.9300,0.9298,0.5300,0.5192,0.5600,0.5098,0.7900,0.7900,0.5600,0.4762,not_run,not_run,single request,single request
|
| 20 |
+
wikitoxic_insult,2,True,0.7700,0.7694,0.5600,0.5593,0.8500,0.8482,0.7000,0.6900,0.9300,0.9300,0.6900,0.6900,0.5400,0.5308,0.8000,0.7960,0.5800,0.5000,not_run,not_run,single request,single request
|
| 21 |
+
wikitoxic_obscene,2,True,0.4400,0.3455,0.5600,0.5572,0.6500,0.6419,0.5200,0.3912,0.8800,0.8798,0.7300,0.7300,0.6300,0.6190,0.8300,0.8279,0.6200,0.5703,not_run,not_run,single request,single request
|
| 22 |
+
wikitoxic_threat,2,True,0.4700,0.3498,0.4700,0.4674,0.9100,0.9100,0.5300,0.3967,0.9700,0.9700,0.7300,0.7254,0.5500,0.5423,0.7700,0.7614,0.5200,0.3763,not_run,not_run,single request,single request
|
| 23 |
+
wikitoxic_toxicaggregated,2,True,0.3300,0.3049,0.5700,0.5502,0.7000,0.6921,0.6200,0.5766,0.8500,0.8500,0.7700,0.7681,0.5700,0.5572,0.8700,0.8678,0.6200,0.5766,not_run,not_run,single request,single request
|
eval/btzsc22/addon_gemma/results.json
ADDED
|
@@ -0,0 +1,2044 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"protocol": "btzsc22-ext-v1",
|
| 3 |
+
"missing_predictions": {
|
| 4 |
+
"A": {
|
| 5 |
+
"ines": 0,
|
| 6 |
+
"julia": 0,
|
| 7 |
+
"laya": 0,
|
| 8 |
+
"gliner": 0,
|
| 9 |
+
"gemma": 0
|
| 10 |
+
},
|
| 11 |
+
"B": {
|
| 12 |
+
"ines": 0,
|
| 13 |
+
"julia": 0,
|
| 14 |
+
"laya": 0,
|
| 15 |
+
"gliner": 0,
|
| 16 |
+
"gemma": "not run (Track B was not run for Gemma)"
|
| 17 |
+
}
|
| 18 |
+
},
|
| 19 |
+
"per_dataset": {
|
| 20 |
+
"agnews": {
|
| 21 |
+
"labels": 4,
|
| 22 |
+
"in_ines_ft_mixture": false,
|
| 23 |
+
"A": {
|
| 24 |
+
"ines": {
|
| 25 |
+
"accuracy": 0.59,
|
| 26 |
+
"macro_f1": 0.5843451916892106
|
| 27 |
+
},
|
| 28 |
+
"julia": {
|
| 29 |
+
"accuracy": 0.94,
|
| 30 |
+
"macro_f1": 0.9399173646545361
|
| 31 |
+
},
|
| 32 |
+
"laya": {
|
| 33 |
+
"accuracy": 0.94,
|
| 34 |
+
"macro_f1": 0.939271525304066
|
| 35 |
+
},
|
| 36 |
+
"gliner": {
|
| 37 |
+
"accuracy": 0.7,
|
| 38 |
+
"macro_f1": 0.6588522320524076
|
| 39 |
+
},
|
| 40 |
+
"gemma": {
|
| 41 |
+
"accuracy": 0.89,
|
| 42 |
+
"macro_f1": 0.8797325209206743
|
| 43 |
+
}
|
| 44 |
+
},
|
| 45 |
+
"B": {
|
| 46 |
+
"ines": {
|
| 47 |
+
"accuracy": 0.42,
|
| 48 |
+
"macro_f1": 0.32652938961350175
|
| 49 |
+
},
|
| 50 |
+
"julia": {
|
| 51 |
+
"accuracy": 0.42,
|
| 52 |
+
"macro_f1": 0.399474348626891
|
| 53 |
+
},
|
| 54 |
+
"laya": {
|
| 55 |
+
"accuracy": 0.89,
|
| 56 |
+
"macro_f1": 0.8889659795261577
|
| 57 |
+
},
|
| 58 |
+
"gliner": {
|
| 59 |
+
"accuracy": 0.45,
|
| 60 |
+
"macro_f1": 0.39861358684888093
|
| 61 |
+
},
|
| 62 |
+
"gemma": null
|
| 63 |
+
},
|
| 64 |
+
"trackA_strategy": {
|
| 65 |
+
"ines": "single request",
|
| 66 |
+
"julia": "single request",
|
| 67 |
+
"laya": "single request",
|
| 68 |
+
"gliner": "single request",
|
| 69 |
+
"gemma": "single request"
|
| 70 |
+
}
|
| 71 |
+
},
|
| 72 |
+
"emotiondair": {
|
| 73 |
+
"labels": 6,
|
| 74 |
+
"in_ines_ft_mixture": false,
|
| 75 |
+
"A": {
|
| 76 |
+
"ines": {
|
| 77 |
+
"accuracy": 0.43,
|
| 78 |
+
"macro_f1": 0.39042230395613853
|
| 79 |
+
},
|
| 80 |
+
"julia": {
|
| 81 |
+
"accuracy": 0.86,
|
| 82 |
+
"macro_f1": 0.8556934538891329
|
| 83 |
+
},
|
| 84 |
+
"laya": {
|
| 85 |
+
"accuracy": 0.37,
|
| 86 |
+
"macro_f1": 0.35268000920174836
|
| 87 |
+
},
|
| 88 |
+
"gliner": {
|
| 89 |
+
"accuracy": 0.44,
|
| 90 |
+
"macro_f1": 0.4070167592505414
|
| 91 |
+
},
|
| 92 |
+
"gemma": {
|
| 93 |
+
"accuracy": 0.4,
|
| 94 |
+
"macro_f1": 0.37637917637917634
|
| 95 |
+
}
|
| 96 |
+
},
|
| 97 |
+
"B": {
|
| 98 |
+
"ines": {
|
| 99 |
+
"accuracy": 0.34,
|
| 100 |
+
"macro_f1": 0.29478472669962036
|
| 101 |
+
},
|
| 102 |
+
"julia": {
|
| 103 |
+
"accuracy": 0.26,
|
| 104 |
+
"macro_f1": 0.25425384800384804
|
| 105 |
+
},
|
| 106 |
+
"laya": {
|
| 107 |
+
"accuracy": 0.43,
|
| 108 |
+
"macro_f1": 0.4279340742423987
|
| 109 |
+
},
|
| 110 |
+
"gliner": {
|
| 111 |
+
"accuracy": 0.35,
|
| 112 |
+
"macro_f1": 0.3193197964297709
|
| 113 |
+
},
|
| 114 |
+
"gemma": null
|
| 115 |
+
},
|
| 116 |
+
"trackA_strategy": {
|
| 117 |
+
"ines": "single request",
|
| 118 |
+
"julia": "single request",
|
| 119 |
+
"laya": "single request",
|
| 120 |
+
"gliner": "single request",
|
| 121 |
+
"gemma": "single request"
|
| 122 |
+
}
|
| 123 |
+
},
|
| 124 |
+
"banking77": {
|
| 125 |
+
"labels": 72,
|
| 126 |
+
"in_ines_ft_mixture": true,
|
| 127 |
+
"A": {
|
| 128 |
+
"ines": {
|
| 129 |
+
"accuracy": 0.39,
|
| 130 |
+
"macro_f1": 0.34265873015873016
|
| 131 |
+
},
|
| 132 |
+
"julia": {
|
| 133 |
+
"accuracy": 0.63,
|
| 134 |
+
"macro_f1": 0.5703703703703703
|
| 135 |
+
},
|
| 136 |
+
"laya": {
|
| 137 |
+
"accuracy": 0.02,
|
| 138 |
+
"macro_f1": 0.0005611672278338945
|
| 139 |
+
},
|
| 140 |
+
"gliner": {
|
| 141 |
+
"accuracy": 0.61,
|
| 142 |
+
"macro_f1": 0.569268077601411
|
| 143 |
+
},
|
| 144 |
+
"gemma": {
|
| 145 |
+
"accuracy": 0.8,
|
| 146 |
+
"macro_f1": 0.7773148148148148
|
| 147 |
+
}
|
| 148 |
+
},
|
| 149 |
+
"B": {
|
| 150 |
+
"ines": {
|
| 151 |
+
"accuracy": 0.29,
|
| 152 |
+
"macro_f1": 0.2234898589065256
|
| 153 |
+
},
|
| 154 |
+
"julia": {
|
| 155 |
+
"accuracy": 0.09,
|
| 156 |
+
"macro_f1": 0.04716343327454439
|
| 157 |
+
},
|
| 158 |
+
"laya": {
|
| 159 |
+
"accuracy": 0.22,
|
| 160 |
+
"macro_f1": 0.18796296296296294
|
| 161 |
+
},
|
| 162 |
+
"gliner": {
|
| 163 |
+
"accuracy": 0.41,
|
| 164 |
+
"macro_f1": 0.3787818662818662
|
| 165 |
+
},
|
| 166 |
+
"gemma": null
|
| 167 |
+
},
|
| 168 |
+
"trackA_strategy": {
|
| 169 |
+
"ines": "knockout-20 (no native > 26 interface)",
|
| 170 |
+
"julia": "Router(width=20, survivors=2)",
|
| 171 |
+
"laya": "single request",
|
| 172 |
+
"gliner": "single request",
|
| 173 |
+
"gemma": "single request"
|
| 174 |
+
}
|
| 175 |
+
},
|
| 176 |
+
"amazonpolarity": {
|
| 177 |
+
"labels": 2,
|
| 178 |
+
"in_ines_ft_mixture": true,
|
| 179 |
+
"A": {
|
| 180 |
+
"ines": {
|
| 181 |
+
"accuracy": 0.78,
|
| 182 |
+
"macro_f1": 0.7756017951856384
|
| 183 |
+
},
|
| 184 |
+
"julia": {
|
| 185 |
+
"accuracy": 0.62,
|
| 186 |
+
"macro_f1": 0.6161616161616161
|
| 187 |
+
},
|
| 188 |
+
"laya": {
|
| 189 |
+
"accuracy": 0.84,
|
| 190 |
+
"macro_f1": 0.8368013055895553
|
| 191 |
+
},
|
| 192 |
+
"gliner": {
|
| 193 |
+
"accuracy": 0.88,
|
| 194 |
+
"macro_f1": 0.8799519807923168
|
| 195 |
+
},
|
| 196 |
+
"gemma": {
|
| 197 |
+
"accuracy": 0.98,
|
| 198 |
+
"macro_f1": 0.98
|
| 199 |
+
}
|
| 200 |
+
},
|
| 201 |
+
"B": {
|
| 202 |
+
"ines": {
|
| 203 |
+
"accuracy": 0.65,
|
| 204 |
+
"macro_f1": 0.6178622120318812
|
| 205 |
+
},
|
| 206 |
+
"julia": {
|
| 207 |
+
"accuracy": 0.58,
|
| 208 |
+
"macro_f1": 0.5716034271725826
|
| 209 |
+
},
|
| 210 |
+
"laya": {
|
| 211 |
+
"accuracy": 0.91,
|
| 212 |
+
"macro_f1": 0.9097744360902256
|
| 213 |
+
},
|
| 214 |
+
"gliner": {
|
| 215 |
+
"accuracy": 0.83,
|
| 216 |
+
"macro_f1": 0.8249407887962105
|
| 217 |
+
},
|
| 218 |
+
"gemma": null
|
| 219 |
+
},
|
| 220 |
+
"trackA_strategy": {
|
| 221 |
+
"ines": "single request",
|
| 222 |
+
"julia": "single request",
|
| 223 |
+
"laya": "single request",
|
| 224 |
+
"gliner": "single request",
|
| 225 |
+
"gemma": "single request"
|
| 226 |
+
}
|
| 227 |
+
},
|
| 228 |
+
"imdb": {
|
| 229 |
+
"labels": 2,
|
| 230 |
+
"in_ines_ft_mixture": false,
|
| 231 |
+
"A": {
|
| 232 |
+
"ines": {
|
| 233 |
+
"accuracy": 0.75,
|
| 234 |
+
"macro_f1": 0.74997499749975
|
| 235 |
+
},
|
| 236 |
+
"julia": {
|
| 237 |
+
"accuracy": 0.68,
|
| 238 |
+
"macro_f1": 0.6798719487795117
|
| 239 |
+
},
|
| 240 |
+
"laya": {
|
| 241 |
+
"accuracy": 0.73,
|
| 242 |
+
"macro_f1": 0.7149192271143491
|
| 243 |
+
},
|
| 244 |
+
"gliner": {
|
| 245 |
+
"accuracy": 0.8,
|
| 246 |
+
"macro_f1": 0.7947454844006567
|
| 247 |
+
},
|
| 248 |
+
"gemma": {
|
| 249 |
+
"accuracy": 0.97,
|
| 250 |
+
"macro_f1": 0.96999699969997
|
| 251 |
+
}
|
| 252 |
+
},
|
| 253 |
+
"B": {
|
| 254 |
+
"ines": {
|
| 255 |
+
"accuracy": 0.69,
|
| 256 |
+
"macro_f1": 0.6807743795695603
|
| 257 |
+
},
|
| 258 |
+
"julia": {
|
| 259 |
+
"accuracy": 0.54,
|
| 260 |
+
"macro_f1": 0.5353535353535354
|
| 261 |
+
},
|
| 262 |
+
"laya": {
|
| 263 |
+
"accuracy": 0.81,
|
| 264 |
+
"macro_f1": 0.8090644156366193
|
| 265 |
+
},
|
| 266 |
+
"gliner": {
|
| 267 |
+
"accuracy": 0.73,
|
| 268 |
+
"macro_f1": 0.7175436761167486
|
| 269 |
+
},
|
| 270 |
+
"gemma": null
|
| 271 |
+
},
|
| 272 |
+
"trackA_strategy": {
|
| 273 |
+
"ines": "single request",
|
| 274 |
+
"julia": "single request",
|
| 275 |
+
"laya": "single request",
|
| 276 |
+
"gliner": "single request",
|
| 277 |
+
"gemma": "single request"
|
| 278 |
+
}
|
| 279 |
+
},
|
| 280 |
+
"appreviews": {
|
| 281 |
+
"labels": 2,
|
| 282 |
+
"in_ines_ft_mixture": false,
|
| 283 |
+
"A": {
|
| 284 |
+
"ines": {
|
| 285 |
+
"accuracy": 0.77,
|
| 286 |
+
"macro_f1": 0.7660461804495982
|
| 287 |
+
},
|
| 288 |
+
"julia": {
|
| 289 |
+
"accuracy": 0.65,
|
| 290 |
+
"macro_f1": 0.6395839769333744
|
| 291 |
+
},
|
| 292 |
+
"laya": {
|
| 293 |
+
"accuracy": 0.78,
|
| 294 |
+
"macro_f1": 0.7756017951856384
|
| 295 |
+
},
|
| 296 |
+
"gliner": {
|
| 297 |
+
"accuracy": 0.89,
|
| 298 |
+
"macro_f1": 0.8894583458948849
|
| 299 |
+
},
|
| 300 |
+
"gemma": {
|
| 301 |
+
"accuracy": 0.94,
|
| 302 |
+
"macro_f1": 0.9399038461538461
|
| 303 |
+
}
|
| 304 |
+
},
|
| 305 |
+
"B": {
|
| 306 |
+
"ines": {
|
| 307 |
+
"accuracy": 0.81,
|
| 308 |
+
"macro_f1": 0.8095238095238095
|
| 309 |
+
},
|
| 310 |
+
"julia": {
|
| 311 |
+
"accuracy": 0.51,
|
| 312 |
+
"macro_f1": 0.49541756770672435
|
| 313 |
+
},
|
| 314 |
+
"laya": {
|
| 315 |
+
"accuracy": 0.89,
|
| 316 |
+
"macro_f1": 0.8897243107769424
|
| 317 |
+
},
|
| 318 |
+
"gliner": {
|
| 319 |
+
"accuracy": 0.8,
|
| 320 |
+
"macro_f1": 0.7960016319869441
|
| 321 |
+
},
|
| 322 |
+
"gemma": null
|
| 323 |
+
},
|
| 324 |
+
"trackA_strategy": {
|
| 325 |
+
"ines": "single request",
|
| 326 |
+
"julia": "single request",
|
| 327 |
+
"laya": "single request",
|
| 328 |
+
"gliner": "single request",
|
| 329 |
+
"gemma": "single request"
|
| 330 |
+
}
|
| 331 |
+
},
|
| 332 |
+
"yelpreviews": {
|
| 333 |
+
"labels": 2,
|
| 334 |
+
"in_ines_ft_mixture": false,
|
| 335 |
+
"A": {
|
| 336 |
+
"ines": {
|
| 337 |
+
"accuracy": 0.72,
|
| 338 |
+
"macro_f1": 0.6996996996996997
|
| 339 |
+
},
|
| 340 |
+
"julia": {
|
| 341 |
+
"accuracy": 0.73,
|
| 342 |
+
"macro_f1": 0.7277951406391774
|
| 343 |
+
},
|
| 344 |
+
"laya": {
|
| 345 |
+
"accuracy": 0.9,
|
| 346 |
+
"macro_f1": 0.8996386993175431
|
| 347 |
+
},
|
| 348 |
+
"gliner": {
|
| 349 |
+
"accuracy": 0.9,
|
| 350 |
+
"macro_f1": 0.8993558776167472
|
| 351 |
+
},
|
| 352 |
+
"gemma": {
|
| 353 |
+
"accuracy": 0.96,
|
| 354 |
+
"macro_f1": 0.9599358974358975
|
| 355 |
+
}
|
| 356 |
+
},
|
| 357 |
+
"B": {
|
| 358 |
+
"ines": {
|
| 359 |
+
"accuracy": 0.68,
|
| 360 |
+
"macro_f1": 0.64349376114082
|
| 361 |
+
},
|
| 362 |
+
"julia": {
|
| 363 |
+
"accuracy": 0.51,
|
| 364 |
+
"macro_f1": 0.5087719298245614
|
| 365 |
+
},
|
| 366 |
+
"laya": {
|
| 367 |
+
"accuracy": 0.93,
|
| 368 |
+
"macro_f1": 0.9298245614035088
|
| 369 |
+
},
|
| 370 |
+
"gliner": {
|
| 371 |
+
"accuracy": 0.74,
|
| 372 |
+
"macro_f1": 0.7211497211497211
|
| 373 |
+
},
|
| 374 |
+
"gemma": null
|
| 375 |
+
},
|
| 376 |
+
"trackA_strategy": {
|
| 377 |
+
"ines": "single request",
|
| 378 |
+
"julia": "single request",
|
| 379 |
+
"laya": "single request",
|
| 380 |
+
"gliner": "single request",
|
| 381 |
+
"gemma": "single request"
|
| 382 |
+
}
|
| 383 |
+
},
|
| 384 |
+
"rottentomatoes": {
|
| 385 |
+
"labels": 2,
|
| 386 |
+
"in_ines_ft_mixture": false,
|
| 387 |
+
"A": {
|
| 388 |
+
"ines": {
|
| 389 |
+
"accuracy": 0.59,
|
| 390 |
+
"macro_f1": 0.5670995670995671
|
| 391 |
+
},
|
| 392 |
+
"julia": {
|
| 393 |
+
"accuracy": 0.66,
|
| 394 |
+
"macro_f1": 0.6263736263736264
|
| 395 |
+
},
|
| 396 |
+
"laya": {
|
| 397 |
+
"accuracy": 0.69,
|
| 398 |
+
"macro_f1": 0.6807743795695603
|
| 399 |
+
},
|
| 400 |
+
"gliner": {
|
| 401 |
+
"accuracy": 0.71,
|
| 402 |
+
"macro_f1": 0.7064480210547626
|
| 403 |
+
},
|
| 404 |
+
"gemma": {
|
| 405 |
+
"accuracy": 0.91,
|
| 406 |
+
"macro_f1": 0.90999099909991
|
| 407 |
+
}
|
| 408 |
+
},
|
| 409 |
+
"B": {
|
| 410 |
+
"ines": {
|
| 411 |
+
"accuracy": 0.64,
|
| 412 |
+
"macro_f1": 0.6347402597402598
|
| 413 |
+
},
|
| 414 |
+
"julia": {
|
| 415 |
+
"accuracy": 0.46,
|
| 416 |
+
"macro_f1": 0.43253467843631777
|
| 417 |
+
},
|
| 418 |
+
"laya": {
|
| 419 |
+
"accuracy": 0.74,
|
| 420 |
+
"macro_f1": 0.74
|
| 421 |
+
},
|
| 422 |
+
"gliner": {
|
| 423 |
+
"accuracy": 0.67,
|
| 424 |
+
"macro_f1": 0.6547756041426928
|
| 425 |
+
},
|
| 426 |
+
"gemma": null
|
| 427 |
+
},
|
| 428 |
+
"trackA_strategy": {
|
| 429 |
+
"ines": "single request",
|
| 430 |
+
"julia": "single request",
|
| 431 |
+
"laya": "single request",
|
| 432 |
+
"gliner": "single request",
|
| 433 |
+
"gemma": "single request"
|
| 434 |
+
}
|
| 435 |
+
},
|
| 436 |
+
"financialphrasebank": {
|
| 437 |
+
"labels": 3,
|
| 438 |
+
"in_ines_ft_mixture": false,
|
| 439 |
+
"A": {
|
| 440 |
+
"ines": {
|
| 441 |
+
"accuracy": 0.51,
|
| 442 |
+
"macro_f1": 0.48911853494169083
|
| 443 |
+
},
|
| 444 |
+
"julia": {
|
| 445 |
+
"accuracy": 0.45,
|
| 446 |
+
"macro_f1": 0.37110209601081806
|
| 447 |
+
},
|
| 448 |
+
"laya": {
|
| 449 |
+
"accuracy": 0.81,
|
| 450 |
+
"macro_f1": 0.8084906388321863
|
| 451 |
+
},
|
| 452 |
+
"gliner": {
|
| 453 |
+
"accuracy": 0.69,
|
| 454 |
+
"macro_f1": 0.6821411017203141
|
| 455 |
+
},
|
| 456 |
+
"gemma": {
|
| 457 |
+
"accuracy": 0.8,
|
| 458 |
+
"macro_f1": 0.8029584462511292
|
| 459 |
+
}
|
| 460 |
+
},
|
| 461 |
+
"B": {
|
| 462 |
+
"ines": {
|
| 463 |
+
"accuracy": 0.45,
|
| 464 |
+
"macro_f1": 0.388558378810103
|
| 465 |
+
},
|
| 466 |
+
"julia": {
|
| 467 |
+
"accuracy": 0.25,
|
| 468 |
+
"macro_f1": 0.17441737846672992
|
| 469 |
+
},
|
| 470 |
+
"laya": {
|
| 471 |
+
"accuracy": 0.91,
|
| 472 |
+
"macro_f1": 0.9096611274030629
|
| 473 |
+
},
|
| 474 |
+
"gliner": {
|
| 475 |
+
"accuracy": 0.55,
|
| 476 |
+
"macro_f1": 0.47845237675746155
|
| 477 |
+
},
|
| 478 |
+
"gemma": null
|
| 479 |
+
},
|
| 480 |
+
"trackA_strategy": {
|
| 481 |
+
"ines": "single request",
|
| 482 |
+
"julia": "single request",
|
| 483 |
+
"laya": "single request",
|
| 484 |
+
"gliner": "single request",
|
| 485 |
+
"gemma": "single request"
|
| 486 |
+
}
|
| 487 |
+
},
|
| 488 |
+
"empathetic": {
|
| 489 |
+
"labels": 32,
|
| 490 |
+
"in_ines_ft_mixture": false,
|
| 491 |
+
"A": {
|
| 492 |
+
"ines": {
|
| 493 |
+
"accuracy": 0.25,
|
| 494 |
+
"macro_f1": 0.1909003551273288
|
| 495 |
+
},
|
| 496 |
+
"julia": {
|
| 497 |
+
"accuracy": 0.17,
|
| 498 |
+
"macro_f1": 0.15942460317460316
|
| 499 |
+
},
|
| 500 |
+
"laya": {
|
| 501 |
+
"accuracy": 0.06,
|
| 502 |
+
"macro_f1": 0.02546768707482993
|
| 503 |
+
},
|
| 504 |
+
"gliner": {
|
| 505 |
+
"accuracy": 0.33,
|
| 506 |
+
"macro_f1": 0.2807043650793651
|
| 507 |
+
},
|
| 508 |
+
"gemma": {
|
| 509 |
+
"accuracy": 0.51,
|
| 510 |
+
"macro_f1": 0.46972402597402596
|
| 511 |
+
}
|
| 512 |
+
},
|
| 513 |
+
"B": {
|
| 514 |
+
"ines": {
|
| 515 |
+
"accuracy": 0.18,
|
| 516 |
+
"macro_f1": 0.14416745535166586
|
| 517 |
+
},
|
| 518 |
+
"julia": {
|
| 519 |
+
"accuracy": 0.05,
|
| 520 |
+
"macro_f1": 0.039042207792207795
|
| 521 |
+
},
|
| 522 |
+
"laya": {
|
| 523 |
+
"accuracy": 0.33,
|
| 524 |
+
"macro_f1": 0.2872519841269841
|
| 525 |
+
},
|
| 526 |
+
"gliner": {
|
| 527 |
+
"accuracy": 0.14,
|
| 528 |
+
"macro_f1": 0.13043154761904763
|
| 529 |
+
},
|
| 530 |
+
"gemma": null
|
| 531 |
+
},
|
| 532 |
+
"trackA_strategy": {
|
| 533 |
+
"ines": "knockout-20 (no native > 26 interface)",
|
| 534 |
+
"julia": "Router(width=20, survivors=2)",
|
| 535 |
+
"laya": "single request",
|
| 536 |
+
"gliner": "single request",
|
| 537 |
+
"gemma": "single request"
|
| 538 |
+
}
|
| 539 |
+
},
|
| 540 |
+
"biasframes_intent": {
|
| 541 |
+
"labels": 2,
|
| 542 |
+
"in_ines_ft_mixture": false,
|
| 543 |
+
"A": {
|
| 544 |
+
"ines": {
|
| 545 |
+
"accuracy": 0.6,
|
| 546 |
+
"macro_f1": 0.5998399359743898
|
| 547 |
+
},
|
| 548 |
+
"julia": {
|
| 549 |
+
"accuracy": 0.57,
|
| 550 |
+
"macro_f1": 0.5664885573142454
|
| 551 |
+
},
|
| 552 |
+
"laya": {
|
| 553 |
+
"accuracy": 0.61,
|
| 554 |
+
"macro_f1": 0.5920075321686369
|
| 555 |
+
},
|
| 556 |
+
"gliner": {
|
| 557 |
+
"accuracy": 0.58,
|
| 558 |
+
"macro_f1": 0.5738636363636364
|
| 559 |
+
},
|
| 560 |
+
"gemma": {
|
| 561 |
+
"accuracy": 0.83,
|
| 562 |
+
"macro_f1": 0.8299829982998299
|
| 563 |
+
}
|
| 564 |
+
},
|
| 565 |
+
"B": {
|
| 566 |
+
"ines": {
|
| 567 |
+
"accuracy": 0.59,
|
| 568 |
+
"macro_f1": 0.5523528769516323
|
| 569 |
+
},
|
| 570 |
+
"julia": {
|
| 571 |
+
"accuracy": 0.59,
|
| 572 |
+
"macro_f1": 0.5805626598465473
|
| 573 |
+
},
|
| 574 |
+
"laya": {
|
| 575 |
+
"accuracy": 0.66,
|
| 576 |
+
"macro_f1": 0.6550324675324675
|
| 577 |
+
},
|
| 578 |
+
"gliner": {
|
| 579 |
+
"accuracy": 0.52,
|
| 580 |
+
"macro_f1": 0.39117199391172
|
| 581 |
+
},
|
| 582 |
+
"gemma": null
|
| 583 |
+
},
|
| 584 |
+
"trackA_strategy": {
|
| 585 |
+
"ines": "single request",
|
| 586 |
+
"julia": "single request",
|
| 587 |
+
"laya": "single request",
|
| 588 |
+
"gliner": "single request",
|
| 589 |
+
"gemma": "single request"
|
| 590 |
+
}
|
| 591 |
+
},
|
| 592 |
+
"massive": {
|
| 593 |
+
"labels": 59,
|
| 594 |
+
"in_ines_ft_mixture": true,
|
| 595 |
+
"A": {
|
| 596 |
+
"ines": {
|
| 597 |
+
"accuracy": 0.35,
|
| 598 |
+
"macro_f1": 0.2502433536331841
|
| 599 |
+
},
|
| 600 |
+
"julia": {
|
| 601 |
+
"accuracy": 0.27,
|
| 602 |
+
"macro_f1": 0.1801721818670971
|
| 603 |
+
},
|
| 604 |
+
"laya": {
|
| 605 |
+
"accuracy": 0.01,
|
| 606 |
+
"macro_f1": 0.0009685230024213075
|
| 607 |
+
},
|
| 608 |
+
"gliner": {
|
| 609 |
+
"accuracy": 0.44,
|
| 610 |
+
"macro_f1": 0.38272522334474407
|
| 611 |
+
},
|
| 612 |
+
"gemma": {
|
| 613 |
+
"accuracy": 0.78,
|
| 614 |
+
"macro_f1": 0.7265536723163842
|
| 615 |
+
}
|
| 616 |
+
},
|
| 617 |
+
"B": {
|
| 618 |
+
"ines": {
|
| 619 |
+
"accuracy": 0.2,
|
| 620 |
+
"macro_f1": 0.15282852740479855
|
| 621 |
+
},
|
| 622 |
+
"julia": {
|
| 623 |
+
"accuracy": 0.02,
|
| 624 |
+
"macro_f1": 0.01224105461393597
|
| 625 |
+
},
|
| 626 |
+
"laya": {
|
| 627 |
+
"accuracy": 0.35,
|
| 628 |
+
"macro_f1": 0.3254640839386602
|
| 629 |
+
},
|
| 630 |
+
"gliner": {
|
| 631 |
+
"accuracy": 0.3,
|
| 632 |
+
"macro_f1": 0.2353510895883777
|
| 633 |
+
},
|
| 634 |
+
"gemma": null
|
| 635 |
+
},
|
| 636 |
+
"trackA_strategy": {
|
| 637 |
+
"ines": "knockout-20 (no native > 26 interface)",
|
| 638 |
+
"julia": "Router(width=20, survivors=2)",
|
| 639 |
+
"laya": "single request",
|
| 640 |
+
"gliner": "single request",
|
| 641 |
+
"gemma": "single request"
|
| 642 |
+
}
|
| 643 |
+
},
|
| 644 |
+
"yahootopics": {
|
| 645 |
+
"labels": 10,
|
| 646 |
+
"in_ines_ft_mixture": false,
|
| 647 |
+
"A": {
|
| 648 |
+
"ines": {
|
| 649 |
+
"accuracy": 0.5,
|
| 650 |
+
"macro_f1": 0.4915120854091442
|
| 651 |
+
},
|
| 652 |
+
"julia": {
|
| 653 |
+
"accuracy": 0.39,
|
| 654 |
+
"macro_f1": 0.3675051008447593
|
| 655 |
+
},
|
| 656 |
+
"laya": {
|
| 657 |
+
"accuracy": 0.15,
|
| 658 |
+
"macro_f1": 0.13690074545337702
|
| 659 |
+
},
|
| 660 |
+
"gliner": {
|
| 661 |
+
"accuracy": 0.41,
|
| 662 |
+
"macro_f1": 0.4169864449276214
|
| 663 |
+
},
|
| 664 |
+
"gemma": {
|
| 665 |
+
"accuracy": 0.73,
|
| 666 |
+
"macro_f1": 0.7217792466130062
|
| 667 |
+
}
|
| 668 |
+
},
|
| 669 |
+
"B": {
|
| 670 |
+
"ines": {
|
| 671 |
+
"accuracy": 0.18,
|
| 672 |
+
"macro_f1": 0.12307189542483658
|
| 673 |
+
},
|
| 674 |
+
"julia": {
|
| 675 |
+
"accuracy": 0.12,
|
| 676 |
+
"macro_f1": 0.11272908345326083
|
| 677 |
+
},
|
| 678 |
+
"laya": {
|
| 679 |
+
"accuracy": 0.37,
|
| 680 |
+
"macro_f1": 0.36499120762278653
|
| 681 |
+
},
|
| 682 |
+
"gliner": {
|
| 683 |
+
"accuracy": 0.32,
|
| 684 |
+
"macro_f1": 0.31750090466808734
|
| 685 |
+
},
|
| 686 |
+
"gemma": null
|
| 687 |
+
},
|
| 688 |
+
"trackA_strategy": {
|
| 689 |
+
"ines": "single request",
|
| 690 |
+
"julia": "single request",
|
| 691 |
+
"laya": "single request",
|
| 692 |
+
"gliner": "single request",
|
| 693 |
+
"gemma": "single request"
|
| 694 |
+
}
|
| 695 |
+
},
|
| 696 |
+
"trueteacher": {
|
| 697 |
+
"labels": 2,
|
| 698 |
+
"in_ines_ft_mixture": false,
|
| 699 |
+
"A": {
|
| 700 |
+
"ines": {
|
| 701 |
+
"accuracy": 0.56,
|
| 702 |
+
"macro_f1": 0.537620849096259
|
| 703 |
+
},
|
| 704 |
+
"julia": {
|
| 705 |
+
"accuracy": 0.52,
|
| 706 |
+
"macro_f1": 0.5151515151515151
|
| 707 |
+
},
|
| 708 |
+
"laya": {
|
| 709 |
+
"accuracy": 0.56,
|
| 710 |
+
"macro_f1": 0.46576007770762506
|
| 711 |
+
},
|
| 712 |
+
"gliner": {
|
| 713 |
+
"accuracy": 0.53,
|
| 714 |
+
"macro_f1": 0.5242433444680635
|
| 715 |
+
},
|
| 716 |
+
"gemma": {
|
| 717 |
+
"accuracy": 0.71,
|
| 718 |
+
"macro_f1": 0.69662098545873
|
| 719 |
+
}
|
| 720 |
+
},
|
| 721 |
+
"B": {
|
| 722 |
+
"ines": {
|
| 723 |
+
"accuracy": 0.5,
|
| 724 |
+
"macro_f1": 0.3333333333333333
|
| 725 |
+
},
|
| 726 |
+
"julia": {
|
| 727 |
+
"accuracy": 0.5,
|
| 728 |
+
"macro_f1": 0.43464495703301675
|
| 729 |
+
},
|
| 730 |
+
"laya": {
|
| 731 |
+
"accuracy": 0.48,
|
| 732 |
+
"macro_f1": 0.4206773618538324
|
| 733 |
+
},
|
| 734 |
+
"gliner": {
|
| 735 |
+
"accuracy": 0.54,
|
| 736 |
+
"macro_f1": 0.45238095238095233
|
| 737 |
+
},
|
| 738 |
+
"gemma": null
|
| 739 |
+
},
|
| 740 |
+
"trackA_strategy": {
|
| 741 |
+
"ines": "single request",
|
| 742 |
+
"julia": "single request",
|
| 743 |
+
"laya": "single request",
|
| 744 |
+
"gliner": "single request",
|
| 745 |
+
"gemma": "single request"
|
| 746 |
+
}
|
| 747 |
+
},
|
| 748 |
+
"manifesto": {
|
| 749 |
+
"labels": 56,
|
| 750 |
+
"in_ines_ft_mixture": false,
|
| 751 |
+
"A": {
|
| 752 |
+
"ines": {
|
| 753 |
+
"accuracy": 0.11,
|
| 754 |
+
"macro_f1": 0.07504578754578754
|
| 755 |
+
},
|
| 756 |
+
"julia": {
|
| 757 |
+
"accuracy": 0.1,
|
| 758 |
+
"macro_f1": 0.0788690476190476
|
| 759 |
+
},
|
| 760 |
+
"laya": {
|
| 761 |
+
"accuracy": 0.03,
|
| 762 |
+
"macro_f1": 0.007711038961038961
|
| 763 |
+
},
|
| 764 |
+
"gliner": {
|
| 765 |
+
"accuracy": 0.02,
|
| 766 |
+
"macro_f1": 0.0007072135785007072
|
| 767 |
+
},
|
| 768 |
+
"gemma": {
|
| 769 |
+
"accuracy": 0.44,
|
| 770 |
+
"macro_f1": 0.37665816326530605
|
| 771 |
+
}
|
| 772 |
+
},
|
| 773 |
+
"B": {
|
| 774 |
+
"ines": {
|
| 775 |
+
"accuracy": 0.05,
|
| 776 |
+
"macro_f1": 0.032902735562310034
|
| 777 |
+
},
|
| 778 |
+
"julia": {
|
| 779 |
+
"accuracy": 0.02,
|
| 780 |
+
"macro_f1": 0.01488095238095238
|
| 781 |
+
},
|
| 782 |
+
"laya": {
|
| 783 |
+
"accuracy": 0.11,
|
| 784 |
+
"macro_f1": 0.09031507339778015
|
| 785 |
+
},
|
| 786 |
+
"gliner": {
|
| 787 |
+
"accuracy": 0.06,
|
| 788 |
+
"macro_f1": 0.013575543120473996
|
| 789 |
+
},
|
| 790 |
+
"gemma": null
|
| 791 |
+
},
|
| 792 |
+
"trackA_strategy": {
|
| 793 |
+
"ines": "knockout-20 (no native > 26 interface)",
|
| 794 |
+
"julia": "Router(width=20, survivors=2)",
|
| 795 |
+
"laya": "single request",
|
| 796 |
+
"gliner": "single request",
|
| 797 |
+
"gemma": "single request"
|
| 798 |
+
}
|
| 799 |
+
},
|
| 800 |
+
"capsotu": {
|
| 801 |
+
"labels": 21,
|
| 802 |
+
"in_ines_ft_mixture": false,
|
| 803 |
+
"A": {
|
| 804 |
+
"ines": {
|
| 805 |
+
"accuracy": 0.36,
|
| 806 |
+
"macro_f1": 0.3120540695728665
|
| 807 |
+
},
|
| 808 |
+
"julia": {
|
| 809 |
+
"accuracy": 0.45,
|
| 810 |
+
"macro_f1": 0.4451588094445238
|
| 811 |
+
},
|
| 812 |
+
"laya": {
|
| 813 |
+
"accuracy": 0.05,
|
| 814 |
+
"macro_f1": 0.005012531328320802
|
| 815 |
+
},
|
| 816 |
+
"gliner": {
|
| 817 |
+
"accuracy": 0.3,
|
| 818 |
+
"macro_f1": 0.2524050024050024
|
| 819 |
+
},
|
| 820 |
+
"gemma": {
|
| 821 |
+
"accuracy": 0.79,
|
| 822 |
+
"macro_f1": 0.7912510769653627
|
| 823 |
+
}
|
| 824 |
+
},
|
| 825 |
+
"B": {
|
| 826 |
+
"ines": {
|
| 827 |
+
"accuracy": 0.21,
|
| 828 |
+
"macro_f1": 0.17844985702128557
|
| 829 |
+
},
|
| 830 |
+
"julia": {
|
| 831 |
+
"accuracy": 0.1,
|
| 832 |
+
"macro_f1": 0.0798883656026513
|
| 833 |
+
},
|
| 834 |
+
"laya": {
|
| 835 |
+
"accuracy": 0.27,
|
| 836 |
+
"macro_f1": 0.2799847801788137
|
| 837 |
+
},
|
| 838 |
+
"gliner": {
|
| 839 |
+
"accuracy": 0.17,
|
| 840 |
+
"macro_f1": 0.18369192344347623
|
| 841 |
+
},
|
| 842 |
+
"gemma": null
|
| 843 |
+
},
|
| 844 |
+
"trackA_strategy": {
|
| 845 |
+
"ines": "single request",
|
| 846 |
+
"julia": "Router(width=20, survivors=2)",
|
| 847 |
+
"laya": "single request",
|
| 848 |
+
"gliner": "single request",
|
| 849 |
+
"gemma": "single request"
|
| 850 |
+
}
|
| 851 |
+
},
|
| 852 |
+
"biasframes_offensive": {
|
| 853 |
+
"labels": 2,
|
| 854 |
+
"in_ines_ft_mixture": false,
|
| 855 |
+
"A": {
|
| 856 |
+
"ines": {
|
| 857 |
+
"accuracy": 0.52,
|
| 858 |
+
"macro_f1": 0.49558638083228246
|
| 859 |
+
},
|
| 860 |
+
"julia": {
|
| 861 |
+
"accuracy": 0.54,
|
| 862 |
+
"macro_f1": 0.5118845500848896
|
| 863 |
+
},
|
| 864 |
+
"laya": {
|
| 865 |
+
"accuracy": 0.62,
|
| 866 |
+
"macro_f1": 0.6193910256410255
|
| 867 |
+
},
|
| 868 |
+
"gliner": {
|
| 869 |
+
"accuracy": 0.46,
|
| 870 |
+
"macro_f1": 0.4375
|
| 871 |
+
},
|
| 872 |
+
"gemma": {
|
| 873 |
+
"accuracy": 0.82,
|
| 874 |
+
"macro_f1": 0.8199279711884754
|
| 875 |
+
}
|
| 876 |
+
},
|
| 877 |
+
"B": {
|
| 878 |
+
"ines": {
|
| 879 |
+
"accuracy": 0.48,
|
| 880 |
+
"macro_f1": 0.48
|
| 881 |
+
},
|
| 882 |
+
"julia": {
|
| 883 |
+
"accuracy": 0.45,
|
| 884 |
+
"macro_f1": 0.40675223816201056
|
| 885 |
+
},
|
| 886 |
+
"laya": {
|
| 887 |
+
"accuracy": 0.69,
|
| 888 |
+
"macro_f1": 0.6884735202492211
|
| 889 |
+
},
|
| 890 |
+
"gliner": {
|
| 891 |
+
"accuracy": 0.52,
|
| 892 |
+
"macro_f1": 0.37629937629937626
|
| 893 |
+
},
|
| 894 |
+
"gemma": null
|
| 895 |
+
},
|
| 896 |
+
"trackA_strategy": {
|
| 897 |
+
"ines": "single request",
|
| 898 |
+
"julia": "single request",
|
| 899 |
+
"laya": "single request",
|
| 900 |
+
"gliner": "single request",
|
| 901 |
+
"gemma": "single request"
|
| 902 |
+
}
|
| 903 |
+
},
|
| 904 |
+
"biasframes_sex": {
|
| 905 |
+
"labels": 2,
|
| 906 |
+
"in_ines_ft_mixture": false,
|
| 907 |
+
"A": {
|
| 908 |
+
"ines": {
|
| 909 |
+
"accuracy": 0.34,
|
| 910 |
+
"macro_f1": 0.3389423076923077
|
| 911 |
+
},
|
| 912 |
+
"julia": {
|
| 913 |
+
"accuracy": 0.56,
|
| 914 |
+
"macro_f1": 0.5555555555555556
|
| 915 |
+
},
|
| 916 |
+
"laya": {
|
| 917 |
+
"accuracy": 0.72,
|
| 918 |
+
"macro_f1": 0.7083333333333334
|
| 919 |
+
},
|
| 920 |
+
"gliner": {
|
| 921 |
+
"accuracy": 0.53,
|
| 922 |
+
"macro_f1": 0.39673982800667434
|
| 923 |
+
},
|
| 924 |
+
"gemma": {
|
| 925 |
+
"accuracy": 0.93,
|
| 926 |
+
"macro_f1": 0.9298245614035088
|
| 927 |
+
}
|
| 928 |
+
},
|
| 929 |
+
"B": {
|
| 930 |
+
"ines": {
|
| 931 |
+
"accuracy": 0.53,
|
| 932 |
+
"macro_f1": 0.5191815856777494
|
| 933 |
+
},
|
| 934 |
+
"julia": {
|
| 935 |
+
"accuracy": 0.56,
|
| 936 |
+
"macro_f1": 0.5098039215686274
|
| 937 |
+
},
|
| 938 |
+
"laya": {
|
| 939 |
+
"accuracy": 0.79,
|
| 940 |
+
"macro_f1": 0.78997899789979
|
| 941 |
+
},
|
| 942 |
+
"gliner": {
|
| 943 |
+
"accuracy": 0.56,
|
| 944 |
+
"macro_f1": 0.47619047619047616
|
| 945 |
+
},
|
| 946 |
+
"gemma": null
|
| 947 |
+
},
|
| 948 |
+
"trackA_strategy": {
|
| 949 |
+
"ines": "single request",
|
| 950 |
+
"julia": "single request",
|
| 951 |
+
"laya": "single request",
|
| 952 |
+
"gliner": "single request",
|
| 953 |
+
"gemma": "single request"
|
| 954 |
+
}
|
| 955 |
+
},
|
| 956 |
+
"wikitoxic_insult": {
|
| 957 |
+
"labels": 2,
|
| 958 |
+
"in_ines_ft_mixture": true,
|
| 959 |
+
"A": {
|
| 960 |
+
"ines": {
|
| 961 |
+
"accuracy": 0.77,
|
| 962 |
+
"macro_f1": 0.7694235588972431
|
| 963 |
+
},
|
| 964 |
+
"julia": {
|
| 965 |
+
"accuracy": 0.56,
|
| 966 |
+
"macro_f1": 0.5592948717948718
|
| 967 |
+
},
|
| 968 |
+
"laya": {
|
| 969 |
+
"accuracy": 0.85,
|
| 970 |
+
"macro_f1": 0.8481627695110842
|
| 971 |
+
},
|
| 972 |
+
"gliner": {
|
| 973 |
+
"accuracy": 0.7,
|
| 974 |
+
"macro_f1": 0.6899545266639107
|
| 975 |
+
},
|
| 976 |
+
"gemma": {
|
| 977 |
+
"accuracy": 0.93,
|
| 978 |
+
"macro_f1": 0.9299929992999301
|
| 979 |
+
}
|
| 980 |
+
},
|
| 981 |
+
"B": {
|
| 982 |
+
"ines": {
|
| 983 |
+
"accuracy": 0.69,
|
| 984 |
+
"macro_f1": 0.6899689968996899
|
| 985 |
+
},
|
| 986 |
+
"julia": {
|
| 987 |
+
"accuracy": 0.54,
|
| 988 |
+
"macro_f1": 0.5308037535699714
|
| 989 |
+
},
|
| 990 |
+
"laya": {
|
| 991 |
+
"accuracy": 0.8,
|
| 992 |
+
"macro_f1": 0.7960016319869441
|
| 993 |
+
},
|
| 994 |
+
"gliner": {
|
| 995 |
+
"accuracy": 0.58,
|
| 996 |
+
"macro_f1": 0.5
|
| 997 |
+
},
|
| 998 |
+
"gemma": null
|
| 999 |
+
},
|
| 1000 |
+
"trackA_strategy": {
|
| 1001 |
+
"ines": "single request",
|
| 1002 |
+
"julia": "single request",
|
| 1003 |
+
"laya": "single request",
|
| 1004 |
+
"gliner": "single request",
|
| 1005 |
+
"gemma": "single request"
|
| 1006 |
+
}
|
| 1007 |
+
},
|
| 1008 |
+
"wikitoxic_obscene": {
|
| 1009 |
+
"labels": 2,
|
| 1010 |
+
"in_ines_ft_mixture": true,
|
| 1011 |
+
"A": {
|
| 1012 |
+
"ines": {
|
| 1013 |
+
"accuracy": 0.44,
|
| 1014 |
+
"macro_f1": 0.34548854604955587
|
| 1015 |
+
},
|
| 1016 |
+
"julia": {
|
| 1017 |
+
"accuracy": 0.56,
|
| 1018 |
+
"macro_f1": 0.5571658615136876
|
| 1019 |
+
},
|
| 1020 |
+
"laya": {
|
| 1021 |
+
"accuracy": 0.65,
|
| 1022 |
+
"macro_f1": 0.6419437340153453
|
| 1023 |
+
},
|
| 1024 |
+
"gliner": {
|
| 1025 |
+
"accuracy": 0.52,
|
| 1026 |
+
"macro_f1": 0.39117199391172
|
| 1027 |
+
},
|
| 1028 |
+
"gemma": {
|
| 1029 |
+
"accuracy": 0.88,
|
| 1030 |
+
"macro_f1": 0.8798076923076923
|
| 1031 |
+
}
|
| 1032 |
+
},
|
| 1033 |
+
"B": {
|
| 1034 |
+
"ines": {
|
| 1035 |
+
"accuracy": 0.73,
|
| 1036 |
+
"macro_f1": 0.72997299729973
|
| 1037 |
+
},
|
| 1038 |
+
"julia": {
|
| 1039 |
+
"accuracy": 0.63,
|
| 1040 |
+
"macro_f1": 0.6189887756152817
|
| 1041 |
+
},
|
| 1042 |
+
"laya": {
|
| 1043 |
+
"accuracy": 0.83,
|
| 1044 |
+
"macro_f1": 0.8279178054458953
|
| 1045 |
+
},
|
| 1046 |
+
"gliner": {
|
| 1047 |
+
"accuracy": 0.62,
|
| 1048 |
+
"macro_f1": 0.5703301673450927
|
| 1049 |
+
},
|
| 1050 |
+
"gemma": null
|
| 1051 |
+
},
|
| 1052 |
+
"trackA_strategy": {
|
| 1053 |
+
"ines": "single request",
|
| 1054 |
+
"julia": "single request",
|
| 1055 |
+
"laya": "single request",
|
| 1056 |
+
"gliner": "single request",
|
| 1057 |
+
"gemma": "single request"
|
| 1058 |
+
}
|
| 1059 |
+
},
|
| 1060 |
+
"wikitoxic_threat": {
|
| 1061 |
+
"labels": 2,
|
| 1062 |
+
"in_ines_ft_mixture": true,
|
| 1063 |
+
"A": {
|
| 1064 |
+
"ines": {
|
| 1065 |
+
"accuracy": 0.47,
|
| 1066 |
+
"macro_f1": 0.3497730339835603
|
| 1067 |
+
},
|
| 1068 |
+
"julia": {
|
| 1069 |
+
"accuracy": 0.47,
|
| 1070 |
+
"macro_f1": 0.467390212038991
|
| 1071 |
+
},
|
| 1072 |
+
"laya": {
|
| 1073 |
+
"accuracy": 0.91,
|
| 1074 |
+
"macro_f1": 0.90999099909991
|
| 1075 |
+
},
|
| 1076 |
+
"gliner": {
|
| 1077 |
+
"accuracy": 0.53,
|
| 1078 |
+
"macro_f1": 0.39673982800667434
|
| 1079 |
+
},
|
| 1080 |
+
"gemma": {
|
| 1081 |
+
"accuracy": 0.97,
|
| 1082 |
+
"macro_f1": 0.96999699969997
|
| 1083 |
+
}
|
| 1084 |
+
},
|
| 1085 |
+
"B": {
|
| 1086 |
+
"ines": {
|
| 1087 |
+
"accuracy": 0.73,
|
| 1088 |
+
"macro_f1": 0.7253585596582239
|
| 1089 |
+
},
|
| 1090 |
+
"julia": {
|
| 1091 |
+
"accuracy": 0.55,
|
| 1092 |
+
"macro_f1": 0.54226426609704
|
| 1093 |
+
},
|
| 1094 |
+
"laya": {
|
| 1095 |
+
"accuracy": 0.77,
|
| 1096 |
+
"macro_f1": 0.7613860358958398
|
| 1097 |
+
},
|
| 1098 |
+
"gliner": {
|
| 1099 |
+
"accuracy": 0.52,
|
| 1100 |
+
"macro_f1": 0.37629937629937626
|
| 1101 |
+
},
|
| 1102 |
+
"gemma": null
|
| 1103 |
+
},
|
| 1104 |
+
"trackA_strategy": {
|
| 1105 |
+
"ines": "single request",
|
| 1106 |
+
"julia": "single request",
|
| 1107 |
+
"laya": "single request",
|
| 1108 |
+
"gliner": "single request",
|
| 1109 |
+
"gemma": "single request"
|
| 1110 |
+
}
|
| 1111 |
+
},
|
| 1112 |
+
"wikitoxic_toxicaggregated": {
|
| 1113 |
+
"labels": 2,
|
| 1114 |
+
"in_ines_ft_mixture": true,
|
| 1115 |
+
"A": {
|
| 1116 |
+
"ines": {
|
| 1117 |
+
"accuracy": 0.33,
|
| 1118 |
+
"macro_f1": 0.30490714804440294
|
| 1119 |
+
},
|
| 1120 |
+
"julia": {
|
| 1121 |
+
"accuracy": 0.57,
|
| 1122 |
+
"macro_f1": 0.5501621508525996
|
| 1123 |
+
},
|
| 1124 |
+
"laya": {
|
| 1125 |
+
"accuracy": 0.7,
|
| 1126 |
+
"macro_f1": 0.6921182266009853
|
| 1127 |
+
},
|
| 1128 |
+
"gliner": {
|
| 1129 |
+
"accuracy": 0.62,
|
| 1130 |
+
"macro_f1": 0.5766488413547237
|
| 1131 |
+
},
|
| 1132 |
+
"gemma": {
|
| 1133 |
+
"accuracy": 0.85,
|
| 1134 |
+
"macro_f1": 0.84998499849985
|
| 1135 |
+
}
|
| 1136 |
+
},
|
| 1137 |
+
"B": {
|
| 1138 |
+
"ines": {
|
| 1139 |
+
"accuracy": 0.77,
|
| 1140 |
+
"macro_f1": 0.7681217864704104
|
| 1141 |
+
},
|
| 1142 |
+
"julia": {
|
| 1143 |
+
"accuracy": 0.57,
|
| 1144 |
+
"macro_f1": 0.557203171661003
|
| 1145 |
+
},
|
| 1146 |
+
"laya": {
|
| 1147 |
+
"accuracy": 0.87,
|
| 1148 |
+
"macro_f1": 0.8677652324280338
|
| 1149 |
+
},
|
| 1150 |
+
"gliner": {
|
| 1151 |
+
"accuracy": 0.62,
|
| 1152 |
+
"macro_f1": 0.5766488413547237
|
| 1153 |
+
},
|
| 1154 |
+
"gemma": null
|
| 1155 |
+
},
|
| 1156 |
+
"trackA_strategy": {
|
| 1157 |
+
"ines": "single request",
|
| 1158 |
+
"julia": "single request",
|
| 1159 |
+
"laya": "single request",
|
| 1160 |
+
"gliner": "single request",
|
| 1161 |
+
"gemma": "single request"
|
| 1162 |
+
}
|
| 1163 |
+
}
|
| 1164 |
+
},
|
| 1165 |
+
"global": {
|
| 1166 |
+
"all 22": {
|
| 1167 |
+
"A": {
|
| 1168 |
+
"ines": 0.4739229278426516,
|
| 1169 |
+
"julia": 0.5245951186849342,
|
| 1170 |
+
"laya": 0.5301139532382007,
|
| 1171 |
+
"gliner": 0.536710369477031,
|
| 1172 |
+
"gemma": 0.7994690041839767
|
| 1173 |
+
},
|
| 1174 |
+
"B": {
|
| 1175 |
+
"ines": 0.45679397195871585,
|
| 1176 |
+
"julia": 0.3572179797391928,
|
| 1177 |
+
"laya": 0.6294614568454058,
|
| 1178 |
+
"gliner": 0.44952051094233986,
|
| 1179 |
+
"gemma": null
|
| 1180 |
+
}
|
| 1181 |
+
},
|
| 1182 |
+
"2 labels (13)": {
|
| 1183 |
+
"A": {
|
| 1184 |
+
"ines": 0.561538769269558,
|
| 1185 |
+
"julia": 0.5825291987072047,
|
| 1186 |
+
"laya": 0.7219571619118916,
|
| 1187 |
+
"gliner": 0.627447823733444,
|
| 1188 |
+
"gemma": 0.8973820729652007
|
| 1189 |
+
},
|
| 1190 |
+
"B": {
|
| 1191 |
+
"ines": 0.6295911198690077,
|
| 1192 |
+
"julia": 0.5172849909267092,
|
| 1193 |
+
"laya": 0.7758169828614861,
|
| 1194 |
+
"gliner": 0.5718255850749258,
|
| 1195 |
+
"gemma": null
|
| 1196 |
+
}
|
| 1197 |
+
},
|
| 1198 |
+
"3-20 labels (4)": {
|
| 1199 |
+
"A": {
|
| 1200 |
+
"ines": 0.48884952899904605,
|
| 1201 |
+
"julia": 0.6335545038498116,
|
| 1202 |
+
"laya": 0.5593357296978444,
|
| 1203 |
+
"gliner": 0.5412491344877212,
|
| 1204 |
+
"gemma": 0.6952123475409966
|
| 1205 |
+
},
|
| 1206 |
+
"B": {
|
| 1207 |
+
"ines": 0.28323609763701546,
|
| 1208 |
+
"julia": 0.23521866463768246,
|
| 1209 |
+
"laya": 0.6478880971986014,
|
| 1210 |
+
"gliner": 0.3784716661760502,
|
| 1211 |
+
"gemma": null
|
| 1212 |
+
}
|
| 1213 |
+
},
|
| 1214 |
+
"21-72 labels (5)": {
|
| 1215 |
+
"A": {
|
| 1216 |
+
"ines": 0.23418045920757943,
|
| 1217 |
+
"julia": 0.28679900249512835,
|
| 1218 |
+
"laya": 0.00794418951888898,
|
| 1219 |
+
"gliner": 0.29716197640180464,
|
| 1220 |
+
"gemma": 0.6283003506671788
|
| 1221 |
+
},
|
| 1222 |
+
"B": {
|
| 1223 |
+
"ines": 0.14636768684931714,
|
| 1224 |
+
"julia": 0.03864320273285837,
|
| 1225 |
+
"laya": 0.23419577692104018,
|
| 1226 |
+
"gliner": 0.18836639401064836,
|
| 1227 |
+
"gemma": null
|
| 1228 |
+
}
|
| 1229 |
+
},
|
| 1230 |
+
"absent from Ines-1 decision FT mixture (15)": {
|
| 1231 |
+
"A": {
|
| 1232 |
+
"ines": 0.4858805497724014,
|
| 1233 |
+
"julia": 0.5360250230979545,
|
| 1234 |
+
"laya": 0.5154640164128853,
|
| 1235 |
+
"gliner": 0.5280778437879452,
|
| 1236 |
+
"gemma": 0.7649777943405899
|
| 1237 |
+
},
|
| 1238 |
+
"B": {
|
| 1239 |
+
"ines": 0.4094576296280325,
|
| 1240 |
+
"julia": 0.3319018448171922,
|
| 1241 |
+
"laya": 0.6114586574566909,
|
| 1242 |
+
"gliner": 0.4284732740710553,
|
| 1243 |
+
"gemma": null
|
| 1244 |
+
}
|
| 1245 |
+
},
|
| 1246 |
+
"families in Ines-1 decision FT mixture (7)": {
|
| 1247 |
+
"A": {
|
| 1248 |
+
"ines": 0.44829945227890217,
|
| 1249 |
+
"julia": 0.500102466371319,
|
| 1250 |
+
"laya": 0.5615066750067336,
|
| 1251 |
+
"gliner": 0.5552086388107859,
|
| 1252 |
+
"gemma": 0.8733787395626631
|
| 1253 |
+
},
|
| 1254 |
+
"B": {
|
| 1255 |
+
"ines": 0.5582289912387514,
|
| 1256 |
+
"julia": 0.411466840286337,
|
| 1257 |
+
"laya": 0.6680388841069373,
|
| 1258 |
+
"gliner": 0.4946217328093781,
|
| 1259 |
+
"gemma": null
|
| 1260 |
+
}
|
| 1261 |
+
}
|
| 1262 |
+
},
|
| 1263 |
+
"bootstrap": {
|
| 1264 |
+
"all 22": {
|
| 1265 |
+
"A:julia - A:ines": {
|
| 1266 |
+
"difference": 0.050672190842282465,
|
| 1267 |
+
"ci95": [
|
| 1268 |
+
-0.026197789339992814,
|
| 1269 |
+
0.12574349835362167
|
| 1270 |
+
]
|
| 1271 |
+
},
|
| 1272 |
+
"A:laya - A:ines": {
|
| 1273 |
+
"difference": 0.05619102539554903,
|
| 1274 |
+
"ci95": [
|
| 1275 |
+
-0.03957741468351192,
|
| 1276 |
+
0.16332715413710902
|
| 1277 |
+
]
|
| 1278 |
+
},
|
| 1279 |
+
"A:gliner - A:ines": {
|
| 1280 |
+
"difference": 0.06278744163437923,
|
| 1281 |
+
"ci95": [
|
| 1282 |
+
0.014208626140706566,
|
| 1283 |
+
0.10706083324209556
|
| 1284 |
+
]
|
| 1285 |
+
},
|
| 1286 |
+
"A:gemma - A:ines": {
|
| 1287 |
+
"difference": 0.32554607634132515,
|
| 1288 |
+
"ci95": [
|
| 1289 |
+
0.2531753321872848,
|
| 1290 |
+
0.39198685847503467
|
| 1291 |
+
]
|
| 1292 |
+
},
|
| 1293 |
+
"B:julia - B:ines": {
|
| 1294 |
+
"difference": -0.099575992219523,
|
| 1295 |
+
"ci95": [
|
| 1296 |
+
-0.14637550498902815,
|
| 1297 |
+
-0.05017137416536893
|
| 1298 |
+
]
|
| 1299 |
+
},
|
| 1300 |
+
"B:laya - B:ines": {
|
| 1301 |
+
"difference": 0.17266748488669,
|
| 1302 |
+
"ci95": [
|
| 1303 |
+
0.11284488099527132,
|
| 1304 |
+
0.23744000375671462
|
| 1305 |
+
]
|
| 1306 |
+
},
|
| 1307 |
+
"B:gliner - B:ines": {
|
| 1308 |
+
"difference": -0.007273461016375937,
|
| 1309 |
+
"ci95": [
|
| 1310 |
+
-0.06982555475121606,
|
| 1311 |
+
0.05100715361862577
|
| 1312 |
+
]
|
| 1313 |
+
},
|
| 1314 |
+
"B:ines - A:ines": {
|
| 1315 |
+
"difference": -0.017128955883935842,
|
| 1316 |
+
"ci95": [
|
| 1317 |
+
-0.09913170297625284,
|
| 1318 |
+
0.07045906519324419
|
| 1319 |
+
]
|
| 1320 |
+
},
|
| 1321 |
+
"B:julia - A:julia": {
|
| 1322 |
+
"difference": -0.16737713894574127,
|
| 1323 |
+
"ci95": [
|
| 1324 |
+
-0.2550239493204071,
|
| 1325 |
+
-0.08802480974576297
|
| 1326 |
+
]
|
| 1327 |
+
},
|
| 1328 |
+
"B:laya - A:laya": {
|
| 1329 |
+
"difference": 0.09934750360720516,
|
| 1330 |
+
"ci95": [
|
| 1331 |
+
0.0454206858231911,
|
| 1332 |
+
0.14433272375703962
|
| 1333 |
+
]
|
| 1334 |
+
},
|
| 1335 |
+
"B:gliner - A:gliner": {
|
| 1336 |
+
"difference": -0.087189858534691,
|
| 1337 |
+
"ci95": [
|
| 1338 |
+
-0.13204415683257073,
|
| 1339 |
+
-0.0406803117967699
|
| 1340 |
+
]
|
| 1341 |
+
}
|
| 1342 |
+
},
|
| 1343 |
+
"2 labels (13)": {
|
| 1344 |
+
"A:julia - A:ines": {
|
| 1345 |
+
"difference": 0.020990429437646712,
|
| 1346 |
+
"ci95": [
|
| 1347 |
+
-0.06061145608979183,
|
| 1348 |
+
0.10740818043124331
|
| 1349 |
+
]
|
| 1350 |
+
},
|
| 1351 |
+
"A:laya - A:ines": {
|
| 1352 |
+
"difference": 0.16041839264233365,
|
| 1353 |
+
"ci95": [
|
| 1354 |
+
0.059458802921318936,
|
| 1355 |
+
0.27206775011759543
|
| 1356 |
+
]
|
| 1357 |
+
},
|
| 1358 |
+
"A:gliner - A:ines": {
|
| 1359 |
+
"difference": 0.0659090544638859,
|
| 1360 |
+
"ci95": [
|
| 1361 |
+
0.004675393764386895,
|
| 1362 |
+
0.12582906809147268
|
| 1363 |
+
]
|
| 1364 |
+
},
|
| 1365 |
+
"A:gemma - A:ines": {
|
| 1366 |
+
"difference": 0.33584330369564275,
|
| 1367 |
+
"ci95": [
|
| 1368 |
+
0.25214648887513785,
|
| 1369 |
+
0.4349238101458538
|
| 1370 |
+
]
|
| 1371 |
+
},
|
| 1372 |
+
"B:julia - B:ines": {
|
| 1373 |
+
"difference": -0.11230612894229851,
|
| 1374 |
+
"ci95": [
|
| 1375 |
+
-0.18080307502401038,
|
| 1376 |
+
-0.042581067377537536
|
| 1377 |
+
]
|
| 1378 |
+
},
|
| 1379 |
+
"B:laya - B:ines": {
|
| 1380 |
+
"difference": 0.14622586299247847,
|
| 1381 |
+
"ci95": [
|
| 1382 |
+
0.09520841847771228,
|
| 1383 |
+
0.20567775456268422
|
| 1384 |
+
]
|
| 1385 |
+
},
|
| 1386 |
+
"B:gliner - B:ines": {
|
| 1387 |
+
"difference": -0.057765534794081974,
|
| 1388 |
+
"ci95": [
|
| 1389 |
+
-0.1428996478267053,
|
| 1390 |
+
0.029876283921421498
|
| 1391 |
+
]
|
| 1392 |
+
},
|
| 1393 |
+
"B:ines - A:ines": {
|
| 1394 |
+
"difference": 0.06805235059944967,
|
| 1395 |
+
"ci95": [
|
| 1396 |
+
-0.04906521051519053,
|
| 1397 |
+
0.19755199873107607
|
| 1398 |
+
]
|
| 1399 |
+
},
|
| 1400 |
+
"B:julia - A:julia": {
|
| 1401 |
+
"difference": -0.06524420778049556,
|
| 1402 |
+
"ci95": [
|
| 1403 |
+
-0.12548066475124145,
|
| 1404 |
+
-0.012510733064128867
|
| 1405 |
+
]
|
| 1406 |
+
},
|
| 1407 |
+
"B:laya - A:laya": {
|
| 1408 |
+
"difference": 0.05385982094959447,
|
| 1409 |
+
"ci95": [
|
| 1410 |
+
-0.0009695761275055773,
|
| 1411 |
+
0.1065780484194153
|
| 1412 |
+
]
|
| 1413 |
+
},
|
| 1414 |
+
"B:gliner - A:gliner": {
|
| 1415 |
+
"difference": -0.0556222386585182,
|
| 1416 |
+
"ci95": [
|
| 1417 |
+
-0.11618959302529945,
|
| 1418 |
+
0.007571405517821818
|
| 1419 |
+
]
|
| 1420 |
+
}
|
| 1421 |
+
},
|
| 1422 |
+
"3-20 labels (4)": {
|
| 1423 |
+
"A:julia - A:ines": {
|
| 1424 |
+
"difference": 0.14470497485076556,
|
| 1425 |
+
"ci95": [
|
| 1426 |
+
-0.12895039644581718,
|
| 1427 |
+
0.42568245069405136
|
| 1428 |
+
]
|
| 1429 |
+
},
|
| 1430 |
+
"A:laya - A:ines": {
|
| 1431 |
+
"difference": 0.07048620069879838,
|
| 1432 |
+
"ci95": [
|
| 1433 |
+
-0.229964039114142,
|
| 1434 |
+
0.3531563180395061
|
| 1435 |
+
]
|
| 1436 |
+
},
|
| 1437 |
+
"A:gliner - A:ines": {
|
| 1438 |
+
"difference": 0.052399605488675075,
|
| 1439 |
+
"ci95": [
|
| 1440 |
+
-0.05581876515491288,
|
| 1441 |
+
0.1600753585487263
|
| 1442 |
+
]
|
| 1443 |
+
},
|
| 1444 |
+
"A:gemma - A:ines": {
|
| 1445 |
+
"difference": 0.2063628185419505,
|
| 1446 |
+
"ci95": [
|
| 1447 |
+
0.06159339629875036,
|
| 1448 |
+
0.3329951554044995
|
| 1449 |
+
]
|
| 1450 |
+
},
|
| 1451 |
+
"B:julia - B:ines": {
|
| 1452 |
+
"difference": -0.048017432999332976,
|
| 1453 |
+
"ci95": [
|
| 1454 |
+
-0.16164512046595278,
|
| 1455 |
+
0.05754769356829411
|
| 1456 |
+
]
|
| 1457 |
+
},
|
| 1458 |
+
"B:laya - B:ines": {
|
| 1459 |
+
"difference": 0.364651999561586,
|
| 1460 |
+
"ci95": [
|
| 1461 |
+
0.18329205393405798,
|
| 1462 |
+
0.5540391760322673
|
| 1463 |
+
]
|
| 1464 |
+
},
|
| 1465 |
+
"B:gliner - B:ines": {
|
| 1466 |
+
"difference": 0.09523556853903474,
|
| 1467 |
+
"ci95": [
|
| 1468 |
+
0.01900107622902177,
|
| 1469 |
+
0.17430955225269407
|
| 1470 |
+
]
|
| 1471 |
+
},
|
| 1472 |
+
"B:ines - A:ines": {
|
| 1473 |
+
"difference": -0.2056134313620306,
|
| 1474 |
+
"ci95": [
|
| 1475 |
+
-0.3164880910714138,
|
| 1476 |
+
-0.09244502441135767
|
| 1477 |
+
]
|
| 1478 |
+
},
|
| 1479 |
+
"B:julia - A:julia": {
|
| 1480 |
+
"difference": -0.3983358392121291,
|
| 1481 |
+
"ci95": [
|
| 1482 |
+
-0.5756158997267588,
|
| 1483 |
+
-0.20827382157801752
|
| 1484 |
+
]
|
| 1485 |
+
},
|
| 1486 |
+
"B:laya - A:laya": {
|
| 1487 |
+
"difference": 0.08855236750075707,
|
| 1488 |
+
"ci95": [
|
| 1489 |
+
-0.008752159824462086,
|
| 1490 |
+
0.1939655463264431
|
| 1491 |
+
]
|
| 1492 |
+
},
|
| 1493 |
+
"B:gliner - A:gliner": {
|
| 1494 |
+
"difference": -0.16277746831167095,
|
| 1495 |
+
"ci95": [
|
| 1496 |
+
-0.25637377988986315,
|
| 1497 |
+
-0.07162803351218687
|
| 1498 |
+
]
|
| 1499 |
+
}
|
| 1500 |
+
},
|
| 1501 |
+
"21-72 labels (5)": {
|
| 1502 |
+
"A:julia - A:ines": {
|
| 1503 |
+
"difference": 0.05261854328754896,
|
| 1504 |
+
"ci95": [
|
| 1505 |
+
-0.04301551177445001,
|
| 1506 |
+
0.15765204279382006
|
| 1507 |
+
]
|
| 1508 |
+
},
|
| 1509 |
+
"A:laya - A:ines": {
|
| 1510 |
+
"difference": -0.22623626968869046,
|
| 1511 |
+
"ci95": [
|
| 1512 |
+
-0.30032025667320744,
|
| 1513 |
+
-0.12187817221877785
|
| 1514 |
+
]
|
| 1515 |
+
},
|
| 1516 |
+
"A:gliner - A:ines": {
|
| 1517 |
+
"difference": 0.06298151719422522,
|
| 1518 |
+
"ci95": [
|
| 1519 |
+
-0.039063579916225874,
|
| 1520 |
+
0.1689149673736492
|
| 1521 |
+
]
|
| 1522 |
+
},
|
| 1523 |
+
"A:gemma - A:ines": {
|
| 1524 |
+
"difference": 0.39411989145959925,
|
| 1525 |
+
"ci95": [
|
| 1526 |
+
0.3019916275555586,
|
| 1527 |
+
0.473337812412976
|
| 1528 |
+
]
|
| 1529 |
+
},
|
| 1530 |
+
"B:julia - B:ines": {
|
| 1531 |
+
"difference": -0.10772448411645877,
|
| 1532 |
+
"ci95": [
|
| 1533 |
+
-0.15259487999276594,
|
| 1534 |
+
-0.04920310152464395
|
| 1535 |
+
]
|
| 1536 |
+
},
|
| 1537 |
+
"B:laya - B:ines": {
|
| 1538 |
+
"difference": 0.0878280900717231,
|
| 1539 |
+
"ci95": [
|
| 1540 |
+
0.011100563313016499,
|
| 1541 |
+
0.14839732102914302
|
| 1542 |
+
]
|
| 1543 |
+
},
|
| 1544 |
+
"B:gliner - B:ines": {
|
| 1545 |
+
"difference": 0.041998707161331236,
|
| 1546 |
+
"ci95": [
|
| 1547 |
+
-0.019529909671919345,
|
| 1548 |
+
0.11057761360469363
|
| 1549 |
+
]
|
| 1550 |
+
},
|
| 1551 |
+
"B:ines - A:ines": {
|
| 1552 |
+
"difference": -0.08781277235826232,
|
| 1553 |
+
"ci95": [
|
| 1554 |
+
-0.12624266008952723,
|
| 1555 |
+
-0.04849955912024958
|
| 1556 |
+
]
|
| 1557 |
+
},
|
| 1558 |
+
"B:julia - A:julia": {
|
| 1559 |
+
"difference": -0.24815579976227004,
|
| 1560 |
+
"ci95": [
|
| 1561 |
+
-0.3988206525705068,
|
| 1562 |
+
-0.09861247338677506
|
| 1563 |
+
]
|
| 1564 |
+
},
|
| 1565 |
+
"B:laya - A:laya": {
|
| 1566 |
+
"difference": 0.22625158740215126,
|
| 1567 |
+
"ci95": [
|
| 1568 |
+
0.12816497504348637,
|
| 1569 |
+
0.286499258070288
|
| 1570 |
+
]
|
| 1571 |
+
},
|
| 1572 |
+
"B:gliner - A:gliner": {
|
| 1573 |
+
"difference": -0.1087955823911563,
|
| 1574 |
+
"ci95": [
|
| 1575 |
+
-0.1700071752364474,
|
| 1576 |
+
-0.032048604061386245
|
| 1577 |
+
]
|
| 1578 |
+
}
|
| 1579 |
+
},
|
| 1580 |
+
"absent from Ines-1 decision FT mixture (15)": {
|
| 1581 |
+
"A:julia - A:ines": {
|
| 1582 |
+
"difference": 0.050144473325553045,
|
| 1583 |
+
"ci95": [
|
| 1584 |
+
-0.028885054771255515,
|
| 1585 |
+
0.1500566532935792
|
| 1586 |
+
]
|
| 1587 |
+
},
|
| 1588 |
+
"A:laya - A:ines": {
|
| 1589 |
+
"difference": 0.029583466640483884,
|
| 1590 |
+
"ci95": [
|
| 1591 |
+
-0.08075161791194013,
|
| 1592 |
+
0.14421513826606108
|
| 1593 |
+
]
|
| 1594 |
+
},
|
| 1595 |
+
"A:gliner - A:ines": {
|
| 1596 |
+
"difference": 0.04219729401554383,
|
| 1597 |
+
"ci95": [
|
| 1598 |
+
-0.006600902674686908,
|
| 1599 |
+
0.09437438464034366
|
| 1600 |
+
]
|
| 1601 |
+
},
|
| 1602 |
+
"A:gemma - A:ines": {
|
| 1603 |
+
"difference": 0.2790972445681884,
|
| 1604 |
+
"ci95": [
|
| 1605 |
+
0.211381616754686,
|
| 1606 |
+
0.3521485454260135
|
| 1607 |
+
]
|
| 1608 |
+
},
|
| 1609 |
+
"B:julia - B:ines": {
|
| 1610 |
+
"difference": -0.0775557848108404,
|
| 1611 |
+
"ci95": [
|
| 1612 |
+
-0.14064220126958463,
|
| 1613 |
+
-0.013517268561773333
|
| 1614 |
+
]
|
| 1615 |
+
},
|
| 1616 |
+
"B:laya - B:ines": {
|
| 1617 |
+
"difference": 0.2020010278286585,
|
| 1618 |
+
"ci95": [
|
| 1619 |
+
0.1269579464717378,
|
| 1620 |
+
0.2900629696566373
|
| 1621 |
+
]
|
| 1622 |
+
},
|
| 1623 |
+
"B:gliner - B:ines": {
|
| 1624 |
+
"difference": 0.019015644443022794,
|
| 1625 |
+
"ci95": [
|
| 1626 |
+
-0.03174972794352044,
|
| 1627 |
+
0.07223435697834166
|
| 1628 |
+
]
|
| 1629 |
+
},
|
| 1630 |
+
"B:ines - A:ines": {
|
| 1631 |
+
"difference": -0.0764229201443689,
|
| 1632 |
+
"ci95": [
|
| 1633 |
+
-0.15017289217507543,
|
| 1634 |
+
-0.009184764537925973
|
| 1635 |
+
]
|
| 1636 |
+
},
|
| 1637 |
+
"B:julia - A:julia": {
|
| 1638 |
+
"difference": -0.20412317828076224,
|
| 1639 |
+
"ci95": [
|
| 1640 |
+
-0.2998785466263466,
|
| 1641 |
+
-0.11592881332069362
|
| 1642 |
+
]
|
| 1643 |
+
},
|
| 1644 |
+
"B:laya - A:laya": {
|
| 1645 |
+
"difference": 0.09599464104380576,
|
| 1646 |
+
"ci95": [
|
| 1647 |
+
0.0459047217540702,
|
| 1648 |
+
0.1442013728576654
|
| 1649 |
+
]
|
| 1650 |
+
},
|
| 1651 |
+
"B:gliner - A:gliner": {
|
| 1652 |
+
"difference": -0.09960456971688988,
|
| 1653 |
+
"ci95": [
|
| 1654 |
+
-0.14649500828750583,
|
| 1655 |
+
-0.04960181354566028
|
| 1656 |
+
]
|
| 1657 |
+
}
|
| 1658 |
+
},
|
| 1659 |
+
"families in Ines-1 decision FT mixture (7)": {
|
| 1660 |
+
"A:julia - A:ines": {
|
| 1661 |
+
"difference": 0.051803014092416944,
|
| 1662 |
+
"ci95": [
|
| 1663 |
+
-0.10019425716510703,
|
| 1664 |
+
0.1778425425557031
|
| 1665 |
+
]
|
| 1666 |
+
},
|
| 1667 |
+
"A:laya - A:ines": {
|
| 1668 |
+
"difference": 0.11320722272783149,
|
| 1669 |
+
"ci95": [
|
| 1670 |
+
-0.12479087223953271,
|
| 1671 |
+
0.3295312555301822
|
| 1672 |
+
]
|
| 1673 |
+
},
|
| 1674 |
+
"A:gliner - A:ines": {
|
| 1675 |
+
"difference": 0.10690918653188367,
|
| 1676 |
+
"ci95": [
|
| 1677 |
+
0.019280206335977032,
|
| 1678 |
+
0.19106643449909544
|
| 1679 |
+
]
|
| 1680 |
+
},
|
| 1681 |
+
"A:gemma - A:ines": {
|
| 1682 |
+
"difference": 0.42507928728376093,
|
| 1683 |
+
"ci95": [
|
| 1684 |
+
0.28040693225491636,
|
| 1685 |
+
0.5352594526522105
|
| 1686 |
+
]
|
| 1687 |
+
},
|
| 1688 |
+
"B:julia - B:ines": {
|
| 1689 |
+
"difference": -0.14676215095241435,
|
| 1690 |
+
"ci95": [
|
| 1691 |
+
-0.1986636251903599,
|
| 1692 |
+
-0.08740265140015356
|
| 1693 |
+
]
|
| 1694 |
+
},
|
| 1695 |
+
"B:laya - B:ines": {
|
| 1696 |
+
"difference": 0.10980989286818602,
|
| 1697 |
+
"ci95": [
|
| 1698 |
+
0.03243625252632493,
|
| 1699 |
+
0.19414880936010162
|
| 1700 |
+
]
|
| 1701 |
+
},
|
| 1702 |
+
"B:gliner - B:ines": {
|
| 1703 |
+
"difference": -0.06360725842937322,
|
| 1704 |
+
"ci95": [
|
| 1705 |
+
-0.21886027228587646,
|
| 1706 |
+
0.08094216222892091
|
| 1707 |
+
]
|
| 1708 |
+
},
|
| 1709 |
+
"B:ines - A:ines": {
|
| 1710 |
+
"difference": 0.10992953895984924,
|
| 1711 |
+
"ci95": [
|
| 1712 |
+
-0.07562649281899102,
|
| 1713 |
+
0.2927199100383737
|
| 1714 |
+
]
|
| 1715 |
+
},
|
| 1716 |
+
"B:julia - A:julia": {
|
| 1717 |
+
"difference": -0.08863562608498207,
|
| 1718 |
+
"ci95": [
|
| 1719 |
+
-0.24495270619294315,
|
| 1720 |
+
0.033542744774361914
|
| 1721 |
+
]
|
| 1722 |
+
},
|
| 1723 |
+
"B:laya - A:laya": {
|
| 1724 |
+
"difference": 0.10653220910020378,
|
| 1725 |
+
"ci95": [
|
| 1726 |
+
-0.015590406007179902,
|
| 1727 |
+
0.20797413933948058
|
| 1728 |
+
]
|
| 1729 |
+
},
|
| 1730 |
+
"B:gliner - A:gliner": {
|
| 1731 |
+
"difference": -0.06058690600140765,
|
| 1732 |
+
"ci95": [
|
| 1733 |
+
-0.14995189706107254,
|
| 1734 |
+
0.032545729267887305
|
| 1735 |
+
]
|
| 1736 |
+
}
|
| 1737 |
+
}
|
| 1738 |
+
},
|
| 1739 |
+
"calibration": {
|
| 1740 |
+
"native (Track A, full native distributions, <=20 labels, 17 datasets)": {
|
| 1741 |
+
"ines": {
|
| 1742 |
+
"brier": 0.5014338202978408,
|
| 1743 |
+
"nll": 0.8000206904458703,
|
| 1744 |
+
"ece": 0.1270706365126021,
|
| 1745 |
+
"coverage_at_error_budget": 0.08941176470588236,
|
| 1746 |
+
"mean_confidence": 0.6118652016888647,
|
| 1747 |
+
"datasets": 17
|
| 1748 |
+
},
|
| 1749 |
+
"julia": {
|
| 1750 |
+
"brier": 0.645280763707045,
|
| 1751 |
+
"nll": 1.6618435294062641,
|
| 1752 |
+
"ece": 0.2983721069531266,
|
| 1753 |
+
"coverage_at_error_budget": 0.1288235294117647,
|
| 1754 |
+
"mean_confidence": 0.8941619627630315,
|
| 1755 |
+
"datasets": 17
|
| 1756 |
+
},
|
| 1757 |
+
"laya": {
|
| 1758 |
+
"brier": 0.4006447652764705,
|
| 1759 |
+
"nll": 0.6925248427789752,
|
| 1760 |
+
"ece": 0.09332229411764704,
|
| 1761 |
+
"coverage_at_error_budget": 0.25588235294117645,
|
| 1762 |
+
"mean_confidence": 0.6874152352941176,
|
| 1763 |
+
"datasets": 17
|
| 1764 |
+
},
|
| 1765 |
+
"gliner": {
|
| 1766 |
+
"brier": 0.48542838104085484,
|
| 1767 |
+
"nll": 0.8599173300405782,
|
| 1768 |
+
"ece": 0.1585389382487745,
|
| 1769 |
+
"coverage_at_error_budget": 0.19705882352941173,
|
| 1770 |
+
"mean_confidence": 0.7310719386879542,
|
| 1771 |
+
"datasets": 17
|
| 1772 |
+
},
|
| 1773 |
+
"gemma": {
|
| 1774 |
+
"brier": 0.24999771502464618,
|
| 1775 |
+
"nll": 0.6128898920857043,
|
| 1776 |
+
"ece": 0.1156891273014901,
|
| 1777 |
+
"coverage_at_error_budget": 0.693529411764706,
|
| 1778 |
+
"mean_confidence": 0.9580137741914275,
|
| 1779 |
+
"datasets": 17
|
| 1780 |
+
}
|
| 1781 |
+
},
|
| 1782 |
+
"derived (Track B softmax of log-odds, same 17 datasets)": {
|
| 1783 |
+
"ines": {
|
| 1784 |
+
"brier": 0.5266089613872135,
|
| 1785 |
+
"nll": 0.8626628971730973,
|
| 1786 |
+
"ece": 0.11876053221277025,
|
| 1787 |
+
"coverage_at_error_budget": 0.1335294117647059,
|
| 1788 |
+
"mean_confidence": 0.4924301710578631,
|
| 1789 |
+
"datasets": 17
|
| 1790 |
+
},
|
| 1791 |
+
"julia": {
|
| 1792 |
+
"brier": 0.7685502933200778,
|
| 1793 |
+
"nll": 1.5719702226511314,
|
| 1794 |
+
"ece": 0.31343832319933446,
|
| 1795 |
+
"coverage_at_error_budget": 0.014117647058823528,
|
| 1796 |
+
"mean_confidence": 0.7770917395785171,
|
| 1797 |
+
"datasets": 17
|
| 1798 |
+
},
|
| 1799 |
+
"laya": {
|
| 1800 |
+
"brier": 0.35941673886508035,
|
| 1801 |
+
"nll": 0.6807608663005745,
|
| 1802 |
+
"ece": 0.12656927310814167,
|
| 1803 |
+
"coverage_at_error_budget": 0.38176470588235295,
|
| 1804 |
+
"mean_confidence": 0.7694262437653563,
|
| 1805 |
+
"datasets": 17
|
| 1806 |
+
},
|
| 1807 |
+
"gliner": {
|
| 1808 |
+
"brier": 0.5872442915941616,
|
| 1809 |
+
"nll": 1.0308428769681912,
|
| 1810 |
+
"ece": 0.2029575594285007,
|
| 1811 |
+
"coverage_at_error_budget": 0.03705882352941177,
|
| 1812 |
+
"mean_confidence": 0.7019368747032854,
|
| 1813 |
+
"datasets": 17
|
| 1814 |
+
}
|
| 1815 |
+
},
|
| 1816 |
+
"derived (Track B softmax of log-odds, all 22 datasets)": {
|
| 1817 |
+
"ines": {
|
| 1818 |
+
"brier": 0.6231953131399431,
|
| 1819 |
+
"nll": 1.4522999819886815,
|
| 1820 |
+
"ece": 0.12354853986128157,
|
| 1821 |
+
"coverage_at_error_budget": 0.10590909090909091,
|
| 1822 |
+
"mean_confidence": 0.39202771650296064,
|
| 1823 |
+
"datasets": 22
|
| 1824 |
+
},
|
| 1825 |
+
"julia": {
|
| 1826 |
+
"brier": 0.8554517402951244,
|
| 1827 |
+
"nll": 2.468274766821923,
|
| 1828 |
+
"ece": 0.31679964418123413,
|
| 1829 |
+
"coverage_at_error_budget": 0.011363636363636364,
|
| 1830 |
+
"mean_confidence": 0.687741939307485,
|
| 1831 |
+
"datasets": 22
|
| 1832 |
+
},
|
| 1833 |
+
"laya": {
|
| 1834 |
+
"brier": 0.47884708479764443,
|
| 1835 |
+
"nll": 1.203588850522752,
|
| 1836 |
+
"ece": 0.12330875423821963,
|
| 1837 |
+
"coverage_at_error_budget": 0.29681818181818187,
|
| 1838 |
+
"mean_confidence": 0.6690071844208508,
|
| 1839 |
+
"datasets": 22
|
| 1840 |
+
},
|
| 1841 |
+
"gliner": {
|
| 1842 |
+
"brier": 0.6643794367941797,
|
| 1843 |
+
"nll": 1.5673671850720794,
|
| 1844 |
+
"ece": 0.18732672413963056,
|
| 1845 |
+
"coverage_at_error_budget": 0.031363636363636364,
|
| 1846 |
+
"mean_confidence": 0.5655688289154754,
|
| 1847 |
+
"datasets": 22
|
| 1848 |
+
}
|
| 1849 |
+
}
|
| 1850 |
+
},
|
| 1851 |
+
"ines_ablation": {
|
| 1852 |
+
"agnews": {
|
| 1853 |
+
"labels": 4,
|
| 1854 |
+
"accuracy": {
|
| 1855 |
+
"native(<=26)": 0.59,
|
| 1856 |
+
"knockout-20": null,
|
| 1857 |
+
"knockout-2": 0.64,
|
| 1858 |
+
"one-vs-rest": 0.42
|
| 1859 |
+
},
|
| 1860 |
+
"macro_f1": {
|
| 1861 |
+
"native(<=26)": 0.5843451916892106,
|
| 1862 |
+
"knockout-20": null,
|
| 1863 |
+
"knockout-2": 0.639691081949832,
|
| 1864 |
+
"one-vs-rest": 0.32652938961350175
|
| 1865 |
+
}
|
| 1866 |
+
},
|
| 1867 |
+
"emotiondair": {
|
| 1868 |
+
"labels": 6,
|
| 1869 |
+
"accuracy": {
|
| 1870 |
+
"native(<=26)": 0.43,
|
| 1871 |
+
"knockout-20": null,
|
| 1872 |
+
"knockout-2": 0.43,
|
| 1873 |
+
"one-vs-rest": 0.34
|
| 1874 |
+
},
|
| 1875 |
+
"macro_f1": {
|
| 1876 |
+
"native(<=26)": 0.39042230395613853,
|
| 1877 |
+
"knockout-20": null,
|
| 1878 |
+
"knockout-2": 0.39217171717171717,
|
| 1879 |
+
"one-vs-rest": 0.29478472669962036
|
| 1880 |
+
}
|
| 1881 |
+
},
|
| 1882 |
+
"banking77": {
|
| 1883 |
+
"labels": 72,
|
| 1884 |
+
"accuracy": {
|
| 1885 |
+
"native(<=26)": null,
|
| 1886 |
+
"knockout-20": 0.39,
|
| 1887 |
+
"knockout-2": 0.45,
|
| 1888 |
+
"one-vs-rest": 0.29
|
| 1889 |
+
},
|
| 1890 |
+
"macro_f1": {
|
| 1891 |
+
"native(<=26)": null,
|
| 1892 |
+
"knockout-20": 0.34265873015873016,
|
| 1893 |
+
"knockout-2": 0.4010912698412698,
|
| 1894 |
+
"one-vs-rest": 0.2234898589065256
|
| 1895 |
+
}
|
| 1896 |
+
},
|
| 1897 |
+
"financialphrasebank": {
|
| 1898 |
+
"labels": 3,
|
| 1899 |
+
"accuracy": {
|
| 1900 |
+
"native(<=26)": 0.51,
|
| 1901 |
+
"knockout-20": null,
|
| 1902 |
+
"knockout-2": 0.44,
|
| 1903 |
+
"one-vs-rest": 0.45
|
| 1904 |
+
},
|
| 1905 |
+
"macro_f1": {
|
| 1906 |
+
"native(<=26)": 0.48911853494169083,
|
| 1907 |
+
"knockout-20": null,
|
| 1908 |
+
"knockout-2": 0.3744570258331727,
|
| 1909 |
+
"one-vs-rest": 0.388558378810103
|
| 1910 |
+
}
|
| 1911 |
+
},
|
| 1912 |
+
"empathetic": {
|
| 1913 |
+
"labels": 32,
|
| 1914 |
+
"accuracy": {
|
| 1915 |
+
"native(<=26)": null,
|
| 1916 |
+
"knockout-20": 0.25,
|
| 1917 |
+
"knockout-2": 0.16,
|
| 1918 |
+
"one-vs-rest": 0.18
|
| 1919 |
+
},
|
| 1920 |
+
"macro_f1": {
|
| 1921 |
+
"native(<=26)": null,
|
| 1922 |
+
"knockout-20": 0.1909003551273288,
|
| 1923 |
+
"knockout-2": 0.10230728593731689,
|
| 1924 |
+
"one-vs-rest": 0.14416745535166586
|
| 1925 |
+
}
|
| 1926 |
+
},
|
| 1927 |
+
"massive": {
|
| 1928 |
+
"labels": 59,
|
| 1929 |
+
"accuracy": {
|
| 1930 |
+
"native(<=26)": null,
|
| 1931 |
+
"knockout-20": 0.35,
|
| 1932 |
+
"knockout-2": 0.41,
|
| 1933 |
+
"one-vs-rest": 0.2
|
| 1934 |
+
},
|
| 1935 |
+
"macro_f1": {
|
| 1936 |
+
"native(<=26)": null,
|
| 1937 |
+
"knockout-20": 0.2502433536331841,
|
| 1938 |
+
"knockout-2": 0.3015873015873015,
|
| 1939 |
+
"one-vs-rest": 0.15282852740479855
|
| 1940 |
+
}
|
| 1941 |
+
},
|
| 1942 |
+
"yahootopics": {
|
| 1943 |
+
"labels": 10,
|
| 1944 |
+
"accuracy": {
|
| 1945 |
+
"native(<=26)": 0.5,
|
| 1946 |
+
"knockout-20": null,
|
| 1947 |
+
"knockout-2": 0.57,
|
| 1948 |
+
"one-vs-rest": 0.18
|
| 1949 |
+
},
|
| 1950 |
+
"macro_f1": {
|
| 1951 |
+
"native(<=26)": 0.4915120854091442,
|
| 1952 |
+
"knockout-20": null,
|
| 1953 |
+
"knockout-2": 0.5584408865987813,
|
| 1954 |
+
"one-vs-rest": 0.12307189542483658
|
| 1955 |
+
}
|
| 1956 |
+
},
|
| 1957 |
+
"manifesto": {
|
| 1958 |
+
"labels": 56,
|
| 1959 |
+
"accuracy": {
|
| 1960 |
+
"native(<=26)": null,
|
| 1961 |
+
"knockout-20": 0.11,
|
| 1962 |
+
"knockout-2": 0.18,
|
| 1963 |
+
"one-vs-rest": 0.05
|
| 1964 |
+
},
|
| 1965 |
+
"macro_f1": {
|
| 1966 |
+
"native(<=26)": null,
|
| 1967 |
+
"knockout-20": 0.07504578754578754,
|
| 1968 |
+
"knockout-2": 0.14917929292929294,
|
| 1969 |
+
"one-vs-rest": 0.032902735562310034
|
| 1970 |
+
}
|
| 1971 |
+
},
|
| 1972 |
+
"capsotu": {
|
| 1973 |
+
"labels": 21,
|
| 1974 |
+
"accuracy": {
|
| 1975 |
+
"native(<=26)": 0.36,
|
| 1976 |
+
"knockout-20": 0.37,
|
| 1977 |
+
"knockout-2": 0.39,
|
| 1978 |
+
"one-vs-rest": 0.21
|
| 1979 |
+
},
|
| 1980 |
+
"macro_f1": {
|
| 1981 |
+
"native(<=26)": 0.3120540695728665,
|
| 1982 |
+
"knockout-20": 0.3317682178584434,
|
| 1983 |
+
"knockout-2": 0.3739835503810659,
|
| 1984 |
+
"one-vs-rest": 0.17844985702128557
|
| 1985 |
+
}
|
| 1986 |
+
}
|
| 1987 |
+
},
|
| 1988 |
+
"anchors": {
|
| 1989 |
+
"GLiNER2.5 Track A vs published pilot (accuracy)": {
|
| 1990 |
+
"agnews": {
|
| 1991 |
+
"ours": 0.7,
|
| 1992 |
+
"published": 0.7
|
| 1993 |
+
},
|
| 1994 |
+
"emotiondair": {
|
| 1995 |
+
"ours": 0.44,
|
| 1996 |
+
"published": 0.44
|
| 1997 |
+
},
|
| 1998 |
+
"banking77": {
|
| 1999 |
+
"ours": 0.61,
|
| 2000 |
+
"published": 0.61
|
| 2001 |
+
}
|
| 2002 |
+
},
|
| 2003 |
+
"Julia-1 Track A vs its model card (accuracy)": {
|
| 2004 |
+
"agnews": {
|
| 2005 |
+
"ours": 0.94,
|
| 2006 |
+
"card": 0.94,
|
| 2007 |
+
"ours_strategy": "single request"
|
| 2008 |
+
},
|
| 2009 |
+
"emotiondair": {
|
| 2010 |
+
"ours": 0.86,
|
| 2011 |
+
"card": 0.86,
|
| 2012 |
+
"ours_strategy": "single request"
|
| 2013 |
+
},
|
| 2014 |
+
"banking77": {
|
| 2015 |
+
"ours": 0.63,
|
| 2016 |
+
"card": 0.64,
|
| 2017 |
+
"ours_strategy": "Router(width=20, survivors=2)"
|
| 2018 |
+
}
|
| 2019 |
+
}
|
| 2020 |
+
},
|
| 2021 |
+
"ovr_ties": {
|
| 2022 |
+
"ines": {
|
| 2023 |
+
"examples_with_tie_at_max": 268,
|
| 2024 |
+
"n": 2200
|
| 2025 |
+
},
|
| 2026 |
+
"julia": {
|
| 2027 |
+
"examples_with_tie_at_max": 4,
|
| 2028 |
+
"n": 2200
|
| 2029 |
+
},
|
| 2030 |
+
"laya": {
|
| 2031 |
+
"examples_with_tie_at_max": 11,
|
| 2032 |
+
"n": 2200
|
| 2033 |
+
},
|
| 2034 |
+
"gliner": {
|
| 2035 |
+
"examples_with_tie_at_max": 0,
|
| 2036 |
+
"n": 2200
|
| 2037 |
+
}
|
| 2038 |
+
},
|
| 2039 |
+
"addon": {
|
| 2040 |
+
"date": "2026-10-06",
|
| 2041 |
+
"model": "RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic (DiffusionGemma-26B-A4B-it, FP8) via razorback16/openjev:0.4.0, model openjev-0.1, think 0",
|
| 2042 |
+
"scope": "Track A (native) only; added after the btzsc22-ext-v1 freeze; the four original models are recomputed by the same code and match results.json exactly"
|
| 2043 |
+
}
|
| 2044 |
+
}
|
eval/btzsc22/addon_gemma/trackA_native.gemma.jsonl
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
eval/speed.md
CHANGED
|
@@ -34,3 +34,36 @@
|
|
| 34 |
~180k prompt tokens/s here, close to its isolated prefill ceiling (~230k).
|
| 35 |
- Not a general speed ranking: different serving stacks (1 vs 4 processes, graphs vs none), different prompt
|
| 36 |
lengths for the same request, one GPU.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
~180k prompt tokens/s here, close to its isolated prefill ceiling (~230k).
|
| 35 |
- Not a general speed ranking: different serving stacks (1 vs 4 processes, graphs vs none), different prompt
|
| 36 |
lengths for the same request, one GPU.
|
| 37 |
+
|
| 38 |
+
## Run 3 (2026-10-06): four models, one GPU, with memory
|
| 39 |
+
|
| 40 |
+
Same load generator (`scripts/bench_jev.py`, closed loop, 16 client processes, HTTP/1.1 without keep-alive), same two
|
| 41 |
+
request sets and levels as above, **one free H200 per model in turn** (GPU 2 of the node, nothing else on it). The
|
| 42 |
+
only change to the load generator: the request's `model` field is taken from `BENCH_MODEL` (OpenJev rejects requests
|
| 43 |
+
without a known model; the other servers ignore it). GPU memory was sampled every 0.5 s with `nvidia-smi` (whole GPU).
|
| 44 |
+
Raw output: `speed_runs/2026-10-06_part1_gemma_laya.txt`, `speed_runs/2026-10-06_part2_gliner_ines.txt` (the first
|
| 45 |
+
GLiNER and Ines-1 start attempts failed on environment issues before any request and were rerun in part 2).
|
| 46 |
+
|
| 47 |
+
| model (params total / active) | server | short c=1 p50 / p95 | short max q/s (c) | typed c=1 p50 / p95 | typed max q/s (c) | GPU memory idle / peak |
|
| 48 |
+
|---|---|---:|---:|---:|---:|---:|
|
| 49 |
+
| **Ines-1** (1.59B / 405M, bf16) | `scripts/serve_jev.py`, 1 process, CUDA graphs, dynamic batching | **9.4 / 9.7 ms** | **1,896** (128) | 30.6 / **40.5 ms** | 544 (32) | 9.7 / 21.9 GB |
|
| 50 |
+
| laya-multilingual arch. (322M, bf16) | production Jev server, 4 processes, micro-batching | 13.0 / 25.9 ms | 1,848 (256) | **20.6** / 89.2 ms | **840** (128) | 8.0 / 24.1 GB |
|
| 51 |
+
| GLiNER2.5-multi arch. (287M) | GLiNER Jev server, 4 workers, **no batching** | 48.8 / 49.3 ms | 68 (64) | 123.1 / 129.3 ms | 49 (32) | 6.7 / 26.5 GB |
|
| 52 |
+
| DiffusionGemma-26B-A4B-it FP8 (~26B / ~4B) | OpenJev 0.4.0 (vLLM), max 64 sequences | 91.7 / 94.0 ms | 79 (32) | 158.8 / 170.2 ms | 143 (32) | weights 25.8 GiB; **120.6 / 135.1 GB reserved** |
|
| 53 |
+
|
| 54 |
+
Latency at the best-throughput level (p50 / p95): Ines-1 118 / 197 ms (short), 282 / 386 ms (typed); laya 212 / 483 ms,
|
| 55 |
+
604 / 913 ms; GLiNER 1,729 / 3,805 ms, 2,070 / 6,871 ms; Gemma 718 / 1,439 ms, 1,008 / 1,545 ms.
|
| 56 |
+
|
| 57 |
+
Reading:
|
| 58 |
+
- **Ines-1 has the lowest single-request latency on short requests and matches the 322M encoder's peak short
|
| 59 |
+
throughput** with one process. On ~300-token questions the encoder is still faster (840 vs 544 questions/s,
|
| 60 |
+
20.6 vs 30.6 ms median), but Ines-1's tail is tighter (p95 40.5 vs 89.2 ms), as in runs 1–2.
|
| 61 |
+
- **Gemma is 5–10× slower at c=1 and ~4–24× lower in peak throughput** than Ines-1 or laya, and returned HTTP 503 on
|
| 62 |
+
12 of the ~12,000 short requests sent at c ≥ 32 (none on typed cases). Its memory figure is vLLM's preallocation (`gpu_memory_utilization`
|
| 63 |
+
0.80: 25.8 GiB of weights + 83 GiB of KV cache), not a minimum; it needs at least the ~26 GiB of weights.
|
| 64 |
+
- **GLiNER's numbers measure its server, not the architecture's ceiling**: that server answers one request per
|
| 65 |
+
worker with no batching, so throughput stays flat as concurrency rises. Input tokens per question are counted
|
| 66 |
+
differently by each server (`usage`): Ines-1 61.5 / 316.2, laya 35.0 / 311.6, GLiNER 8.5 / 37.0, Gemma 67.0 / 136.7.
|
| 67 |
+
- Peak memory includes activations at the highest concurrency; idle is after load and warm-up. The encoder rows use
|
| 68 |
+
domain fine-tunes of those architectures (speed is the architecture's), as in runs 1–2.
|
| 69 |
+
- Not a general speed ranking: different serving stacks, prompt templates and process counts, one GPU.
|
eval/speed_runs/2026-10-06_part1_gemma_laya.txt
ADDED
|
@@ -0,0 +1,40 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
===== gemma 11:23:26
|
| 2 |
+
VRAM_IDLE_MiB 120557
|
| 3 |
+
== short (2 questions)
|
| 4 |
+
c=1 n=3000 11 req/s 22 questions/s p50 91.7 ms p90 93.1 p95 94.0 p99 95.9 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 3000}
|
| 5 |
+
c=8 n=3000 32 req/s 64 questions/s p50 266.5 ms p90 292.5 p95 300.6 p99 357.7 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 3000}
|
| 6 |
+
c=32 n=2992 40 req/s 79 questions/s p50 718.2 ms p90 1260.2 p95 1438.8 p99 1856.4 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2985, 503: 7}
|
| 7 |
+
c=64 n=2992 38 req/s 76 questions/s p50 1506.4 ms p90 2498.3 p95 2818.2 p99 3528.7 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2991, 503: 1}
|
| 8 |
+
c=128 n=2992 36 req/s 71 questions/s p50 3376.2 ms p90 4356.4 p95 4636.1 p99 5255.2 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2988, 503: 4}
|
| 9 |
+
c=256 n=2992 36 req/s 72 questions/s p50 6448.7 ms p90 7529.1 p95 7880.5 p99 8527.4 input tok/question 67.0 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 10 |
+
== typed-decisions ES (5 questions per case)
|
| 11 |
+
c=1 n=800 6 req/s 32 questions/s p50 158.8 ms p90 166.4 p95 170.2 p99 177.3 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
|
| 12 |
+
c=8 n=800 21 req/s 107 questions/s p50 382.0 ms p90 410.2 p95 429.2 p99 522.9 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
|
| 13 |
+
c=32 n=800 29 req/s 143 questions/s p50 1007.7 ms p90 1382.2 p95 1544.9 p99 1773.0 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
|
| 14 |
+
c=128 n=800 25 req/s 127 questions/s p50 4236.0 ms p90 4894.6 p95 5139.9 p99 5471.4 input tok/question 136.7 rows/forward 0.0 in forward 0 % {200: 800}
|
| 15 |
+
VRAM_PEAK_MiB 135149
|
| 16 |
+
(APIServer pid=560) INFO 10-06 11:07:52 [api_utils.py:286] non-default args: {'model_tag': 'RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic', 'enable_auto_tool_choice': True, 'tool_call_parser': 'gemma4', 'host': '127.0.0.1', 'model': 'RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic', 'max_model_len': 65536, 'max_logprobs': 32, 'served_model_name': ['dgemma'], 'override_generation_config': {'max_new_tokens': None}, 'attention_backend': 'TRITON_ATTN', 'reasoning_parser': 'gemma4', 'gpu_memory_utilization': 0.8, 'enable_prefix_caching': True, 'limit_mm_per_prompt': {'image': 8, 'video': 0}, 'max_num_seqs': 64, 'async_scheduling': True, 'diffusion_config': {'canvas_length': 64}}
|
| 17 |
+
(EngineCore pid=1209) INFO 10-06 11:08:27 [default_loader.py:430] Loading weights took 16.33 seconds
|
| 18 |
+
(EngineCore pid=1209) INFO 10-06 11:08:28 [model_runner.py:408] Model loading took 25.83 GiB memory and 20.123655 seconds
|
| 19 |
+
(EngineCore pid=1209) INFO 10-06 11:08:28 [utils.py:320] Using LBNHC KV cache layout.
|
| 20 |
+
(EngineCore pid=1209) INFO 10-06 11:08:50 [gpu_worker.py:674] Available KV cache memory: 83.17 GiB
|
| 21 |
+
(EngineCore pid=1209) INFO 10-06 11:08:50 [gpu_worker.py:689] CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.8000 is equivalent to --gpu-memory-utilization=0.7968 without CUDA graph memory profiling. To maintain the same effective KV cache size as before, increase --gpu-memory-utilization to 0.8032. To disable, set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0.
|
| 22 |
+
(EngineCore pid=1209) INFO 10-06 11:08:50 [kv_cache_utils.py:2432] GPU KV cache size: 1,191,787 tokens, Maximum concurrency for 65,536 tokens per request: 18.19x
|
| 23 |
+
(EngineCore pid=1209) INFO 10-06 11:09:07 [gpu_worker.py:860] CUDA graph pool memory: 0.13 GiB (actual), 0.44 GiB (estimated), difference: 0.31 GiB (231.6%).
|
| 24 |
+
===== laya 11:39:11
|
| 25 |
+
VRAM_IDLE_MiB 7986
|
| 26 |
+
== short (2 questions)
|
| 27 |
+
c=1 n=3000 61 req/s 122 questions/s p50 13.0 ms p90 25.0 p95 25.9 p99 29.0 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 3000}
|
| 28 |
+
c=8 n=3000 284 req/s 567 questions/s p50 28.0 ms p90 33.1 p95 34.9 p99 107.4 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 3000}
|
| 29 |
+
c=32 n=2992 598 req/s 1197 questions/s p50 37.8 ms p90 107.6 p95 142.1 p99 213.6 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 30 |
+
c=64 n=2992 765 req/s 1529 questions/s p50 50.2 ms p90 182.6 p95 220.8 p99 261.2 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 31 |
+
c=128 n=2992 873 req/s 1747 questions/s p50 115.0 ms p90 219.9 p95 239.6 p99 272.7 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 32 |
+
c=256 n=2992 924 req/s 1848 questions/s p50 211.6 ms p90 344.3 p95 483.3 p99 566.3 input tok/question 35.0 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 33 |
+
== typed-decisions ES (5 questions per case)
|
| 34 |
+
c=1 n=800 30 req/s 152 questions/s p50 20.6 ms p90 85.2 p95 89.2 p99 101.3 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
|
| 35 |
+
c=8 n=800 107 req/s 533 questions/s p50 33.1 ms p90 174.0 p95 208.0 p99 270.4 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
|
| 36 |
+
c=32 n=800 150 req/s 750 questions/s p50 178.9 ms p90 351.6 p95 454.6 p99 652.9 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
|
| 37 |
+
c=128 n=800 168 req/s 840 questions/s p50 603.9 ms p90 873.1 p95 912.8 p99 1285.5 input tok/question 311.6 rows/forward 0.0 in forward 0 % {200: 800}
|
| 38 |
+
VRAM_PEAK_MiB 24076
|
| 39 |
+
NO ARRANCA http://localhost:30043
|
| 40 |
+
Error response from daemon: No such container: banco-gliner
|
eval/speed_runs/2026-10-06_part2_gliner_ines.txt
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
===== gliner 11:57:45
|
| 2 |
+
VRAM_IDLE_MiB 6650
|
| 3 |
+
== short (2 questions)
|
| 4 |
+
c=1 n=3000 20 req/s 40 questions/s p50 48.8 ms p90 49.1 p95 49.3 p99 51.2 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 3000}
|
| 5 |
+
c=8 n=3000 21 req/s 43 questions/s p50 419.9 ms p90 487.7 p95 510.9 p99 555.9 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 3000}
|
| 6 |
+
c=32 n=2992 21 req/s 42 questions/s p50 1672.7 ms p90 2033.1 p95 2096.2 p99 2262.7 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 7 |
+
c=64 n=2992 34 req/s 68 questions/s p50 1729.4 ms p90 3542.9 p95 3804.8 p99 4448.3 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 8 |
+
c=128 n=2992 33 req/s 65 questions/s p50 2407.2 ms p90 9051.0 p95 9471.4 p99 9858.6 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 9 |
+
c=256 n=2992 26 req/s 53 questions/s p50 5912.3 ms p90 17059.3 p95 17745.1 p99 17961.1 input tok/question 8.5 rows/forward 0.0 in forward 0 % {200: 2992}
|
| 10 |
+
== typed-decisions ES (5 questions per case)
|
| 11 |
+
c=1 n=800 8 req/s 40 questions/s p50 123.1 ms p90 125.5 p95 129.3 p99 198.9 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
|
| 12 |
+
c=8 n=800 8 req/s 41 questions/s p50 997.9 ms p90 1161.4 p95 1297.9 p99 1472.9 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
|
| 13 |
+
c=32 n=800 10 req/s 49 questions/s p50 2069.6 ms p90 6033.7 p95 6870.5 p99 7986.3 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
|
| 14 |
+
c=128 n=800 9 req/s 45 questions/s p50 11465.9 ms p90 17167.3 p95 19262.3 p99 20541.8 input tok/question 37.0 rows/forward 0.0 in forward 0 % {200: 800}
|
| 15 |
+
VRAM_PEAK_MiB 26544
|
| 16 |
+
===== ines 12:16:37
|
| 17 |
+
VRAM_IDLE_MiB 9747
|
| 18 |
+
== short (2 questions)
|
| 19 |
+
c=1 n=3000 105 req/s 210 questions/s p50 9.4 ms p90 9.7 p95 9.7 p99 9.8 input tok/question 61.5 rows/forward 2.0 in forward 70 % {200: 3000}
|
| 20 |
+
c=8 n=3000 443 req/s 887 questions/s p50 17.8 ms p90 18.0 p95 18.1 p99 18.4 input tok/question 61.5 rows/forward 8.0 in forward 98 % {200: 3000}
|
| 21 |
+
c=32 n=2992 664 req/s 1328 questions/s p50 45.2 ms p90 47.2 p95 52.9 p99 143.1 input tok/question 61.5 rows/forward 30.2 in forward 95 % {200: 2992}
|
| 22 |
+
c=64 n=2992 891 req/s 1782 questions/s p50 66.2 ms p90 69.4 p95 83.9 p99 142.0 input tok/question 61.5 rows/forward 41.6 in forward 96 % {200: 2992}
|
| 23 |
+
c=128 n=2992 948 req/s 1896 questions/s p50 117.8 ms p90 153.1 p95 197.1 p99 198.6 input tok/question 61.5 rows/forward 49.9 in forward 95 % {200: 2992}
|
| 24 |
+
c=256 n=2992 908 req/s 1816 questions/s p50 224.6 ms p90 307.5 p95 343.7 p99 353.6 input tok/question 61.5 rows/forward 53.2 in forward 86 % {200: 2992}
|
| 25 |
+
== typed-decisions ES (5 questions per case)
|
| 26 |
+
c=1 n=800 32 req/s 162 questions/s p50 30.6 ms p90 38.4 p95 40.5 p99 48.4 input tok/question 316.2 rows/forward 5.0 in forward 55 % {200: 800}
|
| 27 |
+
c=8 n=800 93 req/s 465 questions/s p50 85.4 ms p90 110.2 p95 116.2 p99 124.0 input tok/question 316.2 rows/forward 17.6 in forward 99 % {200: 800}
|
| 28 |
+
c=32 n=800 109 req/s 544 questions/s p50 282.3 ms p90 370.3 p95 385.5 p99 457.8 input tok/question 316.2 rows/forward 48.4 in forward 99 % {200: 800}
|
| 29 |
+
c=128 n=800 106 req/s 528 questions/s p50 1022.8 ms p90 1304.1 p95 1306.1 p99 1308.2 input tok/question 316.2 rows/forward 54.0 in forward 98 % {200: 800}
|
| 30 |
+
VRAM_PEAK_MiB 21871
|
| 31 |
+
FIN 12:18:17
|