LocalLLaMA/typed-decisions
Benchmark • Updated • 3.2k • 34.5k • 157
How to use nuhmanpk/pawn-17m with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
A 18M-parameter calibrated decision model, distilled from Rook-1 and Laya's typed-decisions checkpoint.
Give it evidence and typed questions (choice, score, noul), and it returns a calibrated probability for every option in one forward pass.
| Rook SecOps test accuracy (6,005 questions) | 0.752 (ECE 0.041) |
| typed-decisions test accuracy (2,000 questions) | 0.686 (ECE 0.139) |
| Latency, T4 GPU | 15.0 ms median / 17.5 ms p95 per case |
| Latency, CPU (Kaggle 4 vCPU) | 72 ms median / 112 ms p95 per case |
| Weights | 37 MB (fp16 safetensors) |
| Base encoder | jhu-clsp/ettin-encoder-17m |
| Metric | v1 | v1.1 | Change |
|---|---|---|---|
| Rook test accuracy | 0.693 | 0.752 | +0.059 |
| typed-decisions accuracy | 0.656 | 0.686 | +0.030 |
| typed-decisions ECE (lower is better) | 0.113 | 0.139 | +0.026 |
The changes since v1:
pip install laya
import laya
pawn = laya.load("nuhmanpk/pawn-17m")
q = {"phishing": {"type": "noul", "instructions": "Is this message a phishing, smishing or scam attempt?",
"criteria": {"false": "no: a legitimate message with no attempt to deceive",
"true": "yes: it tries to trick the reader into giving up credentials, money or data"}}}
print(pawn.predict({"message": "Your account is locked. Verify here: http://secure-login.example"}, q)["answers"]["phishing"]["noul"])
Pawn runs fully inside a web page with ONNX Runtime Web, with no server and no API key, and the data never leaves the device. Live demo | Browser docs
import * as ort from "https://cdn.jsdelivr.net/npm/[email protected]/dist/ort.all.min.mjs";
import { AutoTokenizer } from "https://cdn.jsdelivr.net/npm/@huggingface/[email protected]";
import { Pawn } from "./pawn.js"; // web/pawn.js in this repo
const pawn = await Pawn.load("nuhmanpk/pawn-17m", { ort, AutoTokenizer });
const out = await pawn.predict({ message: "Verify your password: http://x.example" },
{ phishing: { type: "noul", instructions: "Is this message a phishing, smishing or scam attempt?" } });
| Browser weights | Size | Agreement |
|---|---|---|
fp32 (onnx/model.onnx) |
74 MB | 100.0% of decisions match PyTorch |
int8 (onnx/model_int8.onnx) |
59 MB | 31.8% match fp32, not published (below 98%) |
| workflow | accuracy |
|---|---|
alert_triage |
1.000 |
phishing_message |
0.980 |
cwe_root_cause |
0.943 |
prompt_attack |
0.937 |
phishing_url |
0.915 |
attack_technique |
0.784 |
kev_ransomware |
0.779 |
security_incidents |
0.716 |
vuln_exploitation |
0.713 |
vuln_severity |
0.663 |
guidance_domain |
0.621 |
mitigation_select |
0.590 |
alert_triage reproduces rule-based labels (WitFoo). It measures rule learning, not analyst judgment.Built by Nuhman PK on Laya (Convai Innovations), Ettin encoders (Johns Hopkins University) and CIRCL vulnerability data. The training data is credited on the dataset card.
Base model
jhu-clsp/ettin-encoder-17m