Text Classification
Laya
ONNX
Safetensors
English
typed-decisions
system-one
calibration
distillation
edge
cybersecurity
pawn

Pawn-68M: calibrated typed decisions, small enough to run in your browser

Pawn-68M (v1.1)

A 75M-parameter calibrated decision model, distilled from Rook-1 and Laya's typed-decisions checkpoint. Give it evidence and typed questions (choice, score, noul), and it returns a calibrated probability for every option in one forward pass.

Rook SecOps test accuracy (6,005 questions) 0.783 (ECE 0.060)
typed-decisions test accuracy (2,000 questions) 0.706 (ECE 0.141)
Latency, T4 GPU 28.1 ms median / 32.1 ms p95 per case
Latency, CPU (Kaggle 4 vCPU) 376 ms median / 589 ms p95 per case
Weights 150 MB (fp16 safetensors)
Base encoder jhu-clsp/ettin-encoder-68m

v1.1: what changed

Metric v1 v1.1 Change
Rook test accuracy 0.737 0.783 +0.046
typed-decisions accuracy 0.668 0.706 +0.037
typed-decisions ECE (lower is better) 0.121 0.141 +0.020

The changes since v1:

  • Second teacher: Laya's typed-decisions checkpoint now soft-labels the business workflows, and those cases are repeated 2x.
  • More severity data: 30,000 extra CVE descriptions from CIRCL with real CVSS labels in the natural class mix. They exclude every CVE in the Rook validation and test splits and in CTI-Bench.
  • Softer teacher targets (distillation temperature 2.0).
  • Training schedule: 4 epochs at a higher learning rate (encoder 8e-05, head 0.0003).
  • Calibration on held-out security and business data together.

Usage

pip install laya
import laya
pawn = laya.load("nuhmanpk/pawn-68m")
q = {"phishing": {"type": "noul", "instructions": "Is this message a phishing, smishing or scam attempt?",
      "criteria": {"false": "no: a legitimate message with no attempt to deceive",
                    "true": "yes: it tries to trick the reader into giving up credentials, money or data"}}}
print(pawn.predict({"message": "Your account is locked. Verify here: http://secure-login.example"}, q)["answers"]["phishing"]["noul"])

Use it in the browser

Pawn runs fully inside a web page with ONNX Runtime Web, with no server and no API key, and the data never leaves the device. Live demo | Browser docs

import * as ort from "https://cdn.jsdelivr.net/npm/[email protected]/dist/ort.all.min.mjs";
import { AutoTokenizer } from "https://cdn.jsdelivr.net/npm/@huggingface/[email protected]";
import { Pawn } from "./pawn.js";   // web/pawn.js in this repo
const pawn = await Pawn.load("nuhmanpk/pawn-68m", { ort, AutoTokenizer });
const out = await pawn.predict({ message: "Verify your password: http://x.example" },
  { phishing: { type: "noul", instructions: "Is this message a phishing, smishing or scam attempt?" } });
Browser weights Size Agreement
fp32 (onnx/model.onnx) 300 MB 100.0% of decisions match PyTorch
int8 (onnx/model_int8.onnx) 158 MB 37.7% match fp32, not published (below 98%)

Training

  • Teachers: nuhmanpk/rook-1 for the security workflows and extra CVEs. Laya typed-decisions for the business workflows.
  • Targets: gold-to-teacher mix of 0.5 (security), 0.7 (extra CVEs) and 0.5 (business); teacher softened with tau = 2.0.
  • Recipe: Laya RLCD (proper-scoring-rule policy gradient plus soft cross-entropy) on Kaggle 2x T4 with DDP and fp16. 69,240 training questions; 56 minutes.
  • Calibration: temperatures choice 0.793, score 0.767, noul 0.742.

Rook test accuracy by workflow

workflow accuracy
alert_triage 1.000
phishing_message 0.982
prompt_attack 0.951
cwe_root_cause 0.943
phishing_url 0.928
attack_technique 0.920
kev_ransomware 0.779
guidance_domain 0.772
vuln_exploitation 0.731
security_incidents 0.700
vuln_severity 0.677
mitigation_select 0.639

Limitations

  • Smaller than its teachers. Knowledge-heavy tasks such as CVSS severity and mitigation selection remain the hardest.
  • alert_triage reproduces rule-based labels (WitFoo). It measures rule learning, not analyst judgment.
  • English only.
  • Defensive use only. Keep a human in the loop for consequential actions.
  • Can be steered by adversarial text placed in the state.

Credits

Built by Nuhman PK on Laya (Convai Innovations), Ettin encoders (Johns Hopkins University) and CIRCL vulnerability data. The training data is credited on the dataset card.

Downloads last month
28
Safetensors
Model size
74.8M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nuhmanpk/pawn-68m

Quantized
(10)
this model

Datasets used to train nuhmanpk/pawn-68m

Space using nuhmanpk/pawn-68m 1

Collection including nuhmanpk/pawn-68m