Download evaluation/QUESTION-SCALING.md from llm-semantic-router/Decision-1.0-Nox-4B: direct link, hf CLI and curl.
- Browser
- Download file 1.46 kB
-
https://huggingface.co/llm-semantic-router/Decision-1.0-Nox-4B/resolve/main/evaluation/QUESTION-SCALING.md
- Command line
-
hf download hf://llm-semantic-router/Decision-1.0-Nox-4B/evaluation/QUESTION-SCALING.md
-
curl -L -o QUESTION-SCALING.md https://huggingface.co/llm-semantic-router/Decision-1.0-Nox-4B/resolve/main/evaluation/QUESTION-SCALING.md
Request latency
Nox · distinct Choice questions with 499 input tokens per question. Only the number of questions changes.
| Questions | p50 ms ↓ | p95 ms ↓ | Peak allocated GiB |
|---|---|---|---|
| 1 | 32.678 | 33.719 | 8.079 |
| 2 | 36.370 | 36.647 | 8.146 |
| 4 | 60.415 | 60.952 | 8.280 |
| 8 | 103.619 | 104.181 | 8.549 |
| 16 | 206.991 | 207.624 | 8.549 |
| 32 | 414.199 | 415.371 | 8.549 |
Six independently loaded process blocks supply 30 measured requests per point after warmup. End-to-end Python request latency includes rendering, tokenization, inference, output assembly and final synchronization. Model loading and network are excluded. Requests are sequential, with no cross-request prefix cache.
These measurements precede the Choice null-description update. Every option in this workload supplies a description, so the update does not change these rendered inputs. Current Nox fills a null description with its option ID; this curve does not measure that case. The measured runtime passed an offline Hub-download proof and full 3,160-answer regression; this historical result does not validate a different serving runtime. Measured revision and runtime scope.
These fixed short-input Choice measurements do not establish concurrent HTTP throughput, long-context scaling, other question-type performance or a cross-hardware speed ranking.
