Text Classification
Laya
ONNX
Safetensors
English
typed-decisions
system-one
calibration
distillation
edge
cybersecurity
pawn

Pawn-17M: calibrated typed decisions, small enough to run in your browser

Pawn-17M (v1.1)

A 18M-parameter calibrated decision model, distilled from Rook-1 and Laya's typed-decisions checkpoint. Give it evidence and typed questions (choice, score, noul), and it returns a calibrated probability for every option in one forward pass.

Rook SecOps test accuracy (6,005 questions) 0.752 (ECE 0.041)
typed-decisions test accuracy (2,000 questions) 0.686 (ECE 0.139)
Latency, T4 GPU 15.0 ms median / 17.5 ms p95 per case
Latency, CPU (Kaggle 4 vCPU) 72 ms median / 112 ms p95 per case
Weights 37 MB (fp16 safetensors)
Base encoder jhu-clsp/ettin-encoder-17m

v1.1: what changed

Metric v1 v1.1 Change
Rook test accuracy 0.693 0.752 +0.059
typed-decisions accuracy 0.656 0.686 +0.030
typed-decisions ECE (lower is better) 0.113 0.139 +0.026

The changes since v1:

  • Second teacher: Laya's typed-decisions checkpoint now soft-labels the business workflows, and those cases are repeated 2x.
  • More severity data: 30,000 extra CVE descriptions from CIRCL with real CVSS labels in the natural class mix. They exclude every CVE in the Rook validation and test splits and in CTI-Bench.
  • Softer teacher targets (distillation temperature 2.0).
  • Training schedule: 4 epochs at a higher learning rate (encoder 0.00015, head 0.0004).
  • Calibration on held-out security and business data together.

Usage

pip install laya
import laya
pawn = laya.load("nuhmanpk/pawn-17m")
q = {"phishing": {"type": "noul", "instructions": "Is this message a phishing, smishing or scam attempt?",
      "criteria": {"false": "no: a legitimate message with no attempt to deceive",
                    "true": "yes: it tries to trick the reader into giving up credentials, money or data"}}}
print(pawn.predict({"message": "Your account is locked. Verify here: http://secure-login.example"}, q)["answers"]["phishing"]["noul"])

Use it in the browser

Pawn runs fully inside a web page with ONNX Runtime Web, with no server and no API key, and the data never leaves the device. Live demo | Browser docs

import * as ort from "https://cdn.jsdelivr.net/npm/[email protected]/dist/ort.all.min.mjs";
import { AutoTokenizer } from "https://cdn.jsdelivr.net/npm/@huggingface/[email protected]";
import { Pawn } from "./pawn.js";   // web/pawn.js in this repo
const pawn = await Pawn.load("nuhmanpk/pawn-17m", { ort, AutoTokenizer });
const out = await pawn.predict({ message: "Verify your password: http://x.example" },
  { phishing: { type: "noul", instructions: "Is this message a phishing, smishing or scam attempt?" } });
Browser weights Size Agreement
fp32 (onnx/model.onnx) 74 MB 100.0% of decisions match PyTorch
int8 (onnx/model_int8.onnx) 59 MB 31.8% match fp32, not published (below 98%)

Training

  • Teachers: nuhmanpk/rook-1 for the security workflows and extra CVEs. Laya typed-decisions for the business workflows.
  • Targets: gold-to-teacher mix of 0.5 (security), 0.7 (extra CVEs) and 0.5 (business); teacher softened with tau = 2.0.
  • Recipe: Laya RLCD (proper-scoring-rule policy gradient plus soft cross-entropy) on Kaggle 2x T4 with DDP and fp16. 69,240 training questions; 14 minutes.
  • Calibration: temperatures choice 0.867, score 0.797, noul 0.727.

Rook test accuracy by workflow

workflow accuracy
alert_triage 1.000
phishing_message 0.980
cwe_root_cause 0.943
prompt_attack 0.937
phishing_url 0.915
attack_technique 0.784
kev_ransomware 0.779
security_incidents 0.716
vuln_exploitation 0.713
vuln_severity 0.663
guidance_domain 0.621
mitigation_select 0.590

Limitations

  • Smaller than its teachers. Knowledge-heavy tasks such as CVSS severity and mitigation selection remain the hardest.
  • alert_triage reproduces rule-based labels (WitFoo). It measures rule learning, not analyst judgment.
  • English only.
  • Defensive use only. Keep a human in the loop for consequential actions.
  • Can be steered by adversarial text placed in the state.

Credits

Built by Nuhman PK on Laya (Convai Innovations), Ettin encoders (Johns Hopkins University) and CIRCL vulnerability data. The training data is credited on the dataset card.

Downloads last month
28
Safetensors
Model size
18.5M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nuhmanpk/pawn-17m

Quantized
(9)
this model

Datasets used to train nuhmanpk/pawn-17m

Space using nuhmanpk/pawn-17m 1

Collection including nuhmanpk/pawn-17m