--- library_name: onnxruntime tags: - onnx - onnxruntime - webgpu - browser base_model: convaiinnovations/laya-typed-decisions license: apache-2.0 --- # typed-decisions Laya ONNX Standard ONNX Runtime files are under `onnx/`. Browser-specific graphs with large gather tensors split for WebGPU are under `onnxruntime/webgpu/`. The graph uses fixed batch size 1 with dynamic sequence length and option count. Run multiple questions as separate inference calls. See `manifest.json` for pinned source revisions, checksums, and validation status. Standard and transformed FP32 graphs passed CPU parity checks; browser WebGPU execution has not been validated. Pure FP16-compute exports are unavailable because conversion produced an invalid graph. ## Weight-only INT8 (CPU / WASM target) `onnxruntime/wasm/int8-block-64/model.onnx` uses symmetric 8-bit encoder MatMul weights, block size 64, FP32 activations and decision-head computation. Exactly representable embedding/head weights use FP16 storage with FP32 compute; this is not an FP16-compute model. Recipe informed by [laya-web](https://github.com/nvkudva/laya-web). Quantized encoder MatMuls: 96. First/last encoder layers retained at higher precision: 2 each. Graph plus external weights: 561.28 MB. CPU conversion fidelity against 34 current PyTorch cases: argmax agreement 100.0%, max calibrated probability error 0.016499, mean KL 5.62498e-05. These are parity fixtures, not a task accuracy benchmark. Browser WASM execution has not been validated. This variant is not for WebGPU. Use ONNX Runtime 1.30.0 for the tested CPU path; load the external data file beside the graph. See the accompanying reference, quantization, and validation JSON files. ### Conversion fidelity by weight variant These results measure reproduction of PyTorch outputs, **not task accuracy**. | Weights | Tested backend | Reference cases | Result | | --- | --- | ---: | --- | | Standard FP32 | ONNX Runtime CPU | 34 | Logit parity passed (atol/rtol 0.001); 100.0% argmax agreement; max probability drift 7.79176e-06 | | WebGPU-transformed FP32 | ONNX Runtime CPU only | 3 | Logit parity passed (atol/rtol 0.001); browser untested | | Mixed INT8 weight-only / FP16 storage, FP32 compute | ONNX Runtime CPU | 34 | 100.0% argmax agreement; max probability drift 0.016499; mean KL 5.62498e-05 | FP16-compute weights are not available. The mixed variant's exact retained layers and storage conversions are listed in quantization.json. ### Reproduce the conversion-fidelity test Download this repository, then run from `onnxruntime/wasm/int8-block-64/`: ```sh python -m pip install numpy==2.5.3 onnxruntime==1.30.0 python evaluate.py --standard ../../../onnx/model.onnx ``` [Evaluator](onnxruntime/wasm/int8-block-64/evaluate.py) · [PyTorch references and inputs](onnxruntime/wasm/int8-block-64/reference.json) · [Per-case results](onnxruntime/wasm/int8-block-64/validation.json) · [Quantization recipe](onnxruntime/wasm/int8-block-64/quantization.json) · [All attempts](onnxruntime/wasm/int8-block-64/attempts.json). Requires Python 3.12. The evaluator exits nonzero if a gate fails. Argmax agreement measures agreement with PyTorch, not correctness against human labels. Probability error is measured after the checkpoint's temperature calibration; mean KL measures distribution drift. These 34 synthetic cases cover question types, cardinalities, multilingual text, empty state and truncation. They have no ground-truth labels. The precision fallback was selected on this same set, so these results do not establish held-out accuracy or calibration quality. To regenerate PyTorch references from the pinned checkpoint, use `laya-models int8 --model typed-decisions` from [laya-webgpu](https://github.com/ti3x/laya-webgpu).