Decision-1.0-Nox-4B / evaluation /QUESTION-SCALING.md
Xunzhuo's picture
Publish clean Decision model repository
23fdf61
|
Raw History Blame Contribute Delete
1.46 kB

Request latency

Nox · distinct Choice questions with 499 input tokens per question. Only the number of questions changes.

Question scaling

Questions p50 ms ↓ p95 ms ↓ Peak allocated GiB
1 32.678 33.719 8.079
2 36.370 36.647 8.146
4 60.415 60.952 8.280
8 103.619 104.181 8.549
16 206.991 207.624 8.549
32 414.199 415.371 8.549

Six independently loaded process blocks supply 30 measured requests per point after warmup. End-to-end Python request latency includes rendering, tokenization, inference, output assembly and final synchronization. Model loading and network are excluded. Requests are sequential, with no cross-request prefix cache.

These measurements precede the Choice null-description update. Every option in this workload supplies a description, so the update does not change these rendered inputs. Current Nox fills a null description with its option ID; this curve does not measure that case. The measured runtime passed an offline Hub-download proof and full 3,160-answer regression; this historical result does not validate a different serving runtime. Measured revision and runtime scope.

These fixed short-input Choice measurements do not establish concurrent HTTP throughput, long-context scaling, other question-type performance or a cross-hardware speed ranking.