---
license: apache-2.0
language:
- en
base_model: jhu-clsp/ettin-encoder-17m
datasets:
- nuhmanpk/rook-secops-decisions
- LocalLLaMA/typed-decisions
- CIRCL/vulnerability-scores
pipeline_tag: text-classification
library_name: laya
tags:
- typed-decisions
- system-one
- calibration
- distillation
- edge
- cybersecurity
- pawn
---
# Pawn-17M (v1.1)
**A 18M-parameter calibrated decision model, distilled from [Rook-1](https://huggingface.co/nuhmanpk/rook-1) and Laya's typed-decisions checkpoint.**
Give it evidence and typed questions (`choice`, `score`, `noul`), and it returns a calibrated probability for every option in one forward pass.
| | |
|---|---|
| Rook SecOps test accuracy (6,005 questions) | **0.752** (ECE 0.041) |
| typed-decisions test accuracy (2,000 questions) | **0.686** (ECE 0.139) |
| Latency, T4 GPU | 15.0 ms median / 17.5 ms p95 per case |
| Latency, CPU (Kaggle 4 vCPU) | 72 ms median / 112 ms p95 per case |
| Weights | 37 MB (fp16 safetensors) |
| Base encoder | [jhu-clsp/ettin-encoder-17m](https://huggingface.co/jhu-clsp/ettin-encoder-17m) |
## v1.1: what changed
| Metric | v1 | v1.1 | Change |
|---|---:|---:|---:|
| Rook test accuracy | 0.693 | 0.752 | +0.059 |
| typed-decisions accuracy | 0.656 | 0.686 | +0.030 |
| typed-decisions ECE (lower is better) | 0.113 | 0.139 | +0.026 |
The changes since v1:
- **Second teacher:** Laya's typed-decisions checkpoint now soft-labels the business workflows, and those cases are repeated 2x.
- **More severity data:** 30,000 extra CVE descriptions from CIRCL with real CVSS labels in the natural class mix. They exclude every CVE in the Rook validation and test splits and in CTI-Bench.
- **Softer teacher targets** (distillation temperature 2.0).
- **Training schedule:** 4 epochs at a higher learning rate (encoder 0.00015, head 0.0004).
- **Calibration** on held-out security and business data together.
## Usage
```bash
pip install laya
```
```python
import laya
pawn = laya.load("nuhmanpk/pawn-17m")
q = {"phishing": {"type": "noul", "instructions": "Is this message a phishing, smishing or scam attempt?",
"criteria": {"false": "no: a legitimate message with no attempt to deceive",
"true": "yes: it tries to trick the reader into giving up credentials, money or data"}}}
print(pawn.predict({"message": "Your account is locked. Verify here: http://secure-login.example"}, q)["answers"]["phishing"]["noul"])
```
## Use it in the browser
Pawn runs fully inside a web page with ONNX Runtime Web, with no server and no API key, and the data never leaves the device.
**[Live demo](https://huggingface.co/spaces/nuhmanpk/pawn-browser)** | **[Browser docs](BROWSER.md)**
```js
import * as ort from "https://cdn.jsdelivr.net/npm/onnxruntime-web@1.20.1/dist/ort.all.min.mjs";
import { AutoTokenizer } from "https://cdn.jsdelivr.net/npm/@huggingface/transformers@3.8.1";
import { Pawn } from "./pawn.js"; // web/pawn.js in this repo
const pawn = await Pawn.load("nuhmanpk/pawn-17m", { ort, AutoTokenizer });
const out = await pawn.predict({ message: "Verify your password: http://x.example" },
{ phishing: { type: "noul", instructions: "Is this message a phishing, smishing or scam attempt?" } });
```
| Browser weights | Size | Agreement |
|---|---:|---|
| fp32 (`onnx/model.onnx`) | 74 MB | 100.0% of decisions match PyTorch |
| int8 (`onnx/model_int8.onnx`) | 59 MB | 31.8% match fp32, not published (below 98%) |
## Training
- **Teachers:** [nuhmanpk/rook-1](https://huggingface.co/nuhmanpk/rook-1) for the security workflows and extra CVEs. Laya typed-decisions for the business workflows.
- **Targets:** gold-to-teacher mix of 0.5 (security), 0.7 (extra CVEs) and 0.5 (business); teacher softened with tau = 2.0.
- **Recipe:** Laya RLCD (proper-scoring-rule policy gradient plus soft cross-entropy) on Kaggle 2x T4 with DDP and fp16. 69,240 training questions; 14 minutes.
- **Calibration:** temperatures choice 0.867, score 0.797, noul 0.727.
## Rook test accuracy by workflow
| workflow | accuracy |
|---|---:|
| `alert_triage` | 1.000 |
| `phishing_message` | 0.980 |
| `cwe_root_cause` | 0.943 |
| `prompt_attack` | 0.937 |
| `phishing_url` | 0.915 |
| `attack_technique` | 0.784 |
| `kev_ransomware` | 0.779 |
| `security_incidents` | 0.716 |
| `vuln_exploitation` | 0.713 |
| `vuln_severity` | 0.663 |
| `guidance_domain` | 0.621 |
| `mitigation_select` | 0.590 |
## Limitations
- **Smaller than its teachers.** Knowledge-heavy tasks such as CVSS severity and mitigation selection remain the hardest.
- **`alert_triage` reproduces rule-based labels** (WitFoo). It measures rule learning, not analyst judgment.
- **English only.**
- **Defensive use only.** Keep a human in the loop for consequential actions.
- **Can be steered by adversarial text** placed in the state.
## Credits
Built by **Nuhman PK** on [Laya](https://github.com/NandhaKishorM/laya) (Convai Innovations), [Ettin](https://huggingface.co/jhu-clsp) encoders (Johns Hopkins University) and [CIRCL](https://huggingface.co/CIRCL) vulnerability data.
The training data is credited on the [dataset card](https://huggingface.co/datasets/nuhmanpk/rook-secops-decisions).