Rainier-VL-2B-MTP
Paper: Rainier-VL: A Compact Vision-Language Model for Industrial Inspection Across Edge GPUs
Myan Sudharsanan, Hemmanth Puppala, Gavin Daher, and Jack Rosenbloom
Denali Advanced Integration Physical AI Team
Abstract
Efficiently running defect detection via visual inspection on a production line or on returns benches is a problem that is the intersection of three concerns, speed, accuracy, and cost. In the domain of defect detection, no other Vision Language Model (VLM) can satisfy the constraints for the areas of concern out of the box. Hence, we present Rainier-VL, a 2.1 B-parameter VLM that meets that operating point across the edge GPU range, from a 50 W, 8 GB workstation card to a Blackwell-class part. There are two novel architectural advancements with Rainier-VL. The OrthoBernstein projector treats the vision→language connector as Householder-orthogonal layers composed with a learnable Bernstein activation, and an SVD warm-started to a dense teacher. This results in giving 24.2× fewer parameters and 12.2× fewer forward FLOPs than the dense equivalent. A hybrid state-space backbone replaces the growing KV cache with a fixed recurrent state, making per-frame cost constant on a continuous camera feed. Rainier-VL reaches 0.949 image-AUROC on VisA defect detection, ahead of every open-source 2–10 B model we evaluate in the same harness, including a 10.3 B model at 0.852, at 4.6× less inference memory than that model and no quantization anywhere.
This repository
This is the Rainier-VL-2B checkpoint with multi-token prediction (MTP), packaged with its own inference engine and an OpenAI-compatible server. One image plus a short instruction returns a defect verdict with a calibrated probability, per-defect segmentation masks, bounding boxes, a dense defect-probability map, or a one-sentence description. It runs on an 8 GB GPU such as the RTX A1000.
Model details
| Parameters | 2.094 B unique (0.878 B vision + 1.215 B language + 1.08 M connector); 1.644 B active at inference |
| Vision encoder | SigLIP-SO400M, 384 px, 729 visual tokens per frame |
| Projector | OrthoBernstein (Householder-orthogonal + Bernstein activation), 1.08 M parameters |
| Language model | Zamba2-1.2B (hybrid Mamba2 / attention) |
| Extra heads | 3 dense mask heads, 1 box head (10.07 M total); MTP, 4 heads (16.78 M, 33.6 MB) |
| Weights on disk | 3.99 GB, bf16, no quantization |
| Input | one RGB image, longest side 384 px, JPEG q90 |
| Output | JSON (verdicts, RLE masks, boxes, scores) or free text |
One vision forward per image produces a shared projected grid that the language model, the box head and the mask heads all read, so asking for a verdict, a box and a mask on the same frame pays the vision cost once.
Tasks
The task is selected by a tag at the start of the prompt.
| task | prompt | returns |
|---|---|---|
detect |
[detect] |
answer (yes/no), p_yes (calibrated probability), raw generated JSON |
seg |
[seg] Segment the defect. |
per-defect instances (mask, box, label, severity), union mask_rle, object silhouette |
od |
[od] Localize the defect. |
one box per defect plus a primary bbox |
anomaly |
[anomaly] Localize the defect. |
defect-probability map and anomaly_score |
caption |
[caption] Describe this <noun> in one short sentence, noting any visible damage or defects. |
one sentence of text |
For detect the engine fills in the canonical question itself; send only the tag. od and
anomaly reuse the seg pass on the same frame, so asking for all three costs about one seg
call.
Benchmarks
Defect detection vs open-source VLMs
Image-AUROC across four domains. Every row is scored in the same harness on the same images (per-category cap 40). Parameters are unique parameters.
| model | params (B) | VisA | BSData | BTAD | MVTec (leather) |
|---|---|---|---|---|---|
| Rainier-VL-2B | 2.1 | 0.949 | 0.974 | 0.975 | 0.995 |
| GLM-4.6V-Flash | 10.3 | 0.852 | 0.943 | 0.898 | 0.987 |
| Qwen3-VL-4B | 4.4 | 0.840 | 0.915 | 0.571 | 0.997 |
| Qwen3-VL-2B | 2.1 | 0.802 | 0.916 | 0.531 | 0.996 |
| Cosmos-Reason2-2B | 2.4 | 0.801 | 0.923 | 0.666 | 0.989 |
| Holo-3.1-4B | 4.5 | 0.728 | 0.855 | 0.637 | 0.999 |
| NuExtract3-4B | 4.5 | 0.709 | 0.911 | 0.270 | 0.991 |
| Qwen2.5-VL-3B | 3.8 | 0.702 | 0.619 | 0.785 | 0.988 |
| Qwen3.5-4B | 4.5 | 0.683 | 0.794 | 0.766 | 0.998 |
| Gemma-4-E2B | 5.1 | 0.625 | 0.933 | 0.321 | 0.996 |
Logical anomalies (MVTec-LOCO) are excluded: every model evaluated, Rainier-VL included, scores between 0.47 and 0.56 there. The strongest peer on VisA (GLM-4.6V-Flash) needs 20.9 GB of GPU memory; Rainier-VL peaks at 6.4 GB.
Full held-out reserves
On the complete held-out reserves (no per-category cap) this checkpoint reaches:
| benchmark | image-AUROC |
|---|---|
| VisA | 0.968 |
| BSData | 0.973 |
| BTAD | 0.968 |
Segmentation heads
Validation IoU of the mask heads shipped in weights/heads/:
| head | role | IoU |
|---|---|---|
defect_cisba |
defect masks and anomaly map | 0.858 |
typed_text |
per-instance defect typing | 0.735 |
generic_coco |
object silhouette | 0.626 |
Decode cost vs open-source VLMs
Per-token decode cost on an RTX PRO 6000, best of 3 after three warm-ups. Rainier-VL is measured
with the bundled rainier_engine (static hybrid cache, fused multi-token verify, MTP speculative
decoding); the other models are served through vLLM with CUDA graphs.
| model | params (B) | ms / token | relative |
|---|---|---|---|
| Rainier-VL-2B | 2.1 | 3.40 | 1.0× |
| Qwen3-VL-2B | 2.1 | 8.99 | 2.6× |
| Qwen2.5-VL-3B | 3.8 | 15.74 | 4.6× |
| Qwen3-VL-4B | 4.4 | 20.57 | 6.0× |
Speed and latency
RTX A1000 8 GB (50 W, deployment target)
Measured by the bundled engine per request, single stream, one 384 px frame, fresh compute (caches excluded), card under sustained load.
| task | p50 | p95 | notes |
|---|---|---|---|
detect |
671 ms | 800 ms | time to first token 558 ms; speculative decode, 82 tok/s |
seg |
746 ms | 1,486 ms | mask head 323 ms of the total; includes silhouette and defect typing |
od |
+2 ms | +3 ms | reuses the seg pass on the same frame |
anomaly |
+2 ms | +3 ms | reuses the seg pass on the same frame |
caption |
882 ms | 2,334 ms | greedy decode, 30 tok/s, 64-token cap |
Prefill of a 739-token sequence takes about 390 ms at the default chunk_size=256. Peak GPU memory
is about 6.4 GB with MTP and all heads loaded.
In a deployed inspection station (capture, inference, overlay and write, 2,747 frames over 65 minutes, two camera streams) the full application runs at 908 ms per frame end to end, 1.101 frames/s sustained, with per-camera revisit at 1,591 ms (p10 1,504 ms, p90 1,669 ms). Adding mask and box output to a verdict-only frame costs 39 ms.
RTX PRO 6000 (Blackwell)
The engine's cost per request is a fixed part (preprocess, vision encode, prefill) plus a per-token part, so answer length is the biggest lever you control.
| request | generated tokens | latency | frames / s |
|---|---|---|---|
| verdict only (P(yes), no generation) | 0 | ~59 ms | ~17 |
{"defect": "no"} |
7 | ~79 ms | ~12.6 |
| one sentence | 48 | ~219 ms | ~4.6 |
Fixed cost ≈ 55.6 ms plus ≈ 3.40 ms per generated token. Batching two frames in one call amortises about 57% of the per-request overhead.
Vision encoder alone, batch 1 / 4 / 8: 19.9 / 36.5 / 70.3 ms.
Throughput
| setting | figure |
|---|---|
detect, A1000, fresh frames |
~89 requests / min |
seg, A1000, fresh frames |
~80 requests / min |
| deployed station, A1000, end to end | 1.101 frames / s over 2 streams |
| full five-task sweep on a static scene, A1000, warm cache | p50 27 ms, ~67 frames / min |
| verdict-only stream, RTX PRO 6000 | ~17 frames / s |
| vision encode, RTX PRO 6000, batch 1 / 4 / 8 | 50.2 / 109.6 / 113.8 img / s |
| sustained decode, RTX PRO 6000 | 198 tok / s |
On a static scene the server answers repeats from its scene cache, so a monitoring loop runs far above the fresh-compute rate and pays the fresh cost once when the scene changes.
Multi-token prediction (MTP)
MTP adds four small heads to the language model, each predicting one of the next four tokens from the model's last hidden state. They are trained jointly with the base model and share its embedding table, so they add only 16.8 M parameters. At inference MTP proposes several tokens at once and the base model checks them all in a single forward pass, keeping only the tokens that match its own greedy choice. Output is identical to plain greedy decoding, and each forward pass commits more than three tokens on average.
| MTP acceptance by position (+1 / +2 / +3 / +4) | 0.780 / 0.802 / 0.812 / 0.777 |
| Accepted tokens per verify round | 3.31 |
| Verify step, RTX A1000 | 24.8 ms |
| Realized decode speedup, RTX A1000 | 1.695× |
| Decode speedup at k = 8, RTX PRO 6000 | 2.05× (6.30 ms per token with MTP, 12.96 ms greedy) |
Determinism
Outputs are bit-identical within a process and across server restarts when
CUBLAS_WORKSPACE_CONFIG=:4096:8 and a persistent TRITON_CACHE_DIR are set before launch.
Requirements
- NVIDIA GPU with 8 GB VRAM or more, Ampere or newer. Peak use is about 6.4 GB. The 4 GB RTX A1000 variant cannot run this model, and there is no quantized version.
- CUDA 12.8-capable driver, Python 3.12
torch==2.10.0,transformers==5.5.4(pinned; other versions change the SigLIP module layout),mamba-ssmandcausal-conv1dCUDA wheels. Seerequirements.txt.- For the HTTP server additionally:
fastapi,uvicorn,scipy. - Access to
google/siglip-so400m-patch14-384andZyphra/Zamba2-1.2Bon the Hugging Face Hub. The loader builds both towers from these configs and then loads Rainier's weights over them. After the first download they are read from the local HF cache.
Quick start
1. Download and install
hf download Denali-AI/Rainier-VL-2B-MTP --local-dir rainier-vl-2b-mtp
cd rainier-vl-2b-mtp
python3.12 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install fastapi uvicorn scipy
2. Check the GPU and run the smoke test
PYTHONPATH=code python -m ortho_vlm.kernel_probe # reports which decode kernel tier this GPU gets
python smoke_test.py # loads the weights and checks for a valid JSON answer
smoke_test.py must end with PASS. It verifies that the weights and MTP heads load and that
the model returns a parseable {"defect": ...} answer.
3. Python API
import sys; sys.path.insert(0, "code")
from PIL import Image
from rainier_engine import Rainier
m = Rainier.from_pretrained("weights/vlm.pt", drafter="weights/drafter/head.pt")
img = Image.open("part.jpg").convert("RGB")
out = m.generate(img, "Describe any defect in this image.", max_new_tokens=64)
print(out["texts"][0])
The Python API covers generation (detect and caption). The mask heads, object profiles, confidence gate and scene cache live in the server.
4. Start the server
export CUBLAS_WORKSPACE_CONFIG=:4096:8 # deterministic outputs across restarts
export TRITON_CACHE_DIR=$PWD/.triton-cache # persistent compiled kernels
python serve/engine_shim.py --model-dir . --port 8150
The server warms up before it starts listening; wait for ready on :8150. Then:
curl localhost:8150/health # lists loaded modes and spec-decode status
For an offline host, copy a populated HF cache alongside the model and use
serve/serve_standalone.sh (it sets HF_HUB_OFFLINE=1); set MODEL_PATH to this directory
and PYTHON to your venv's interpreter.
5. Send a request
import base64, io, json, requests
from PIL import Image
img = Image.open("part.jpg").convert("RGB")
img.thumbnail((384, 384), Image.Resampling.BOX) # longest side 384, never upscale
buf = io.BytesIO(); img.save(buf, "JPEG", quality=90)
uri = "data:image/jpeg;base64," + base64.b64encode(buf.getvalue()).decode()
def ask(text, max_tokens=24):
r = requests.post("http://localhost:8150/v1/chat/completions", json={
"model": "rainier-vl-2b-jointmtp-v2-a1000",
"temperature": 0, "max_tokens": max_tokens,
"messages": [{"role": "user", "content": [
{"type": "text", "text": text},
{"type": "image_url", "image_url": {"url": uri}}]}]})
return r.json()["choices"][0]["message"]["content"]
verdict = json.loads(ask("[detect] [pf:metal]"))
print(verdict["answer"], verdict["p_yes"])
seg = json.loads(ask("[seg] [pf:metal] Segment the defect."))
for inst in seg.get("instances", []):
print(inst.get("label"), inst.get("severity"), inst["bbox"])
print(ask("[caption] Describe this metal part in one short sentence, "
"noting any visible damage or defects.", max_tokens=64))
detect, seg, od and anomaly return a JSON string in message.content; caption
returns plain text. Boxes are normalized [x0, y0, x1, y1]. Masks are run-length encoded as
{"size": [h, w], "counts": [...]}, row-major, alternating 0-runs and 1-runs starting with a
0-run. "stream": true is supported and returns the whole answer as one SSE chunk.
Prompting rules
- Send the bare prompt. Do not wrap it in
USER:/ASSISTANT:or any chat-role template. The model was trained on bare prompts and a role wrapper breaks the answer format. - Put tags first, then the task text. Tags are stripped before the text reaches the model.
- Use
temperature: 0.
Control tags
| tag | effect |
|---|---|
[pf:<profile>] |
object profile: sets defect vocabulary and object noun (see below). Default generic |
[gate:on|off|<p>] |
confidence gate: per-camera calibration, off, or a fixed P(yes) threshold such as 0.90 |
[cam:<id>] |
camera id, used to look up per-camera gate thresholds |
[tg:<type>] |
limit seg/od/anomaly output to one defect type (defect = all) |
[dw:x] |
mask tightness, 0 to 1 |
[ds:x] |
defect sensitivity, 0 to 1 |
[es:x] |
edge sensitivity, 0 to 1 |
[as:x] |
anomaly sensitivity, 0 to 1 |
[bs:x] |
box size, 0 to 1 |
[tp:type=dw,ds,es,as,bs;...] |
per-defect-type values for the five sliders above |
[fresh] |
bypass all caches for this request |
Slider values of 0.5 are neutral.
Object profiles
A profile tells the model what kind of object is in frame and which defect types to look for.
| profile | object | defect types |
|---|---|---|
generic |
product | scratch, crack, dent, stain, dirt, tear, contamination, missing_part |
shoes |
shoe | stain, dirt, scratch, tear, glue_residue, discoloration, crease |
packaging |
box | tear, dent, open_seam, stain, hole, contamination |
pcb |
circuit board | missing_part, contamination, scratch, stain |
electronics |
electronic device | scratch, crack, dent, missing_part, burn_mark, contamination |
clothing |
garment | stain, tear, missing_part, contamination |
metal |
metal part | scratch, dent, corrosion, crack, contamination |
weld |
weld joint | crack, porosity, undercut, spatter, incomplete_fusion |
painted_frame |
painted metal tube | paint_run, orange_peel, blister, chip, contamination |
The full registry, including the probe questions used per defect type, is in
hyperparameters/object_profiles.engine.json. New profiles are added to PROFILES in
serve/engine_shim.py.
Calibrating for a new line
The generated yes/no verdict should be calibrated per camera before it is trusted on a new
production line. Collect about 20 clean and 20 defective frames per camera, measure how well
p_yes separates them, and write the per-camera thresholds to serve/detect_calib.json
(Denali's calibrate_detect.py tool produces this file). With [gate:on] the server then
applies the threshold with a 3-frame rolling median per camera, and masks follow the gated
decision. Enable the gate only on cameras where calibration AUROC is 0.9 or higher.
No calibration file ships with this repository. Without one, [gate:on] has no effect and
the server returns the model's generated verdict.
Configuration
Server behaviour can be tuned through environment variables on the serve process. Defaults
and descriptions are in hyperparameters/engine_env_defaults.json. The most commonly changed:
| variable | default | effect |
|---|---|---|
RAINIER_MAX_NEW_CAP |
64 |
maximum tokens for captions and free-text answers |
RAINIER_SPEC |
adaptive |
speculative decoding policy |
RAINIER_DETECT_CALIB |
serve/detect_calib.json |
path to per-camera gate thresholds (reloaded on change) |
RAINIER_GATE_MEDIAN_N |
3 |
rolling-median window for the gate |
Limitations
- Defects smaller than about 10 px at 384 px are usually detected but not localized. Use a closer crop or a zoomed camera for them.
- The scene cache matches frames with a perceptual hash. On a static view, a small new defect
can be served the previous clean result for up to five requests; send
[fresh]when every frame must be recomputed. - The yes/no verdict is not calibrated for new lines out of the box (see above).
- Speculative decoding speeds up the short JSON answers of
detect; free-form captions fall back to greedy decoding. - The server processes one request at a time on one GPU.
Files
| path | contents |
|---|---|
weights/vlm.pt |
base model weights |
weights/drafter/ |
MTP heads |
weights/heads/ |
mask heads (defect_cisba, generic_coco, typed_text), each with its paired projector |
weights/box_head.pt |
box head, used only when no mask head is loaded |
weights/tokenizer/ |
tokenizer |
code/ |
rainier_engine (inference API) and ortho_vlm (model code and kernels) |
serve/ |
engine_shim.py server, serve_standalone.sh offline launcher, helper modules |
hyperparameters/ |
object-profile registry and server environment defaults |
smoke_test.py |
install check |
config.json |
model configuration summary |
MANIFEST.json |
sha256 of every weight file |
paper/ |
the Rainier-VL paper (PDF) |
Citation
@techreport{sudharsanan2026rainiervl,
title = {Rainier-VL: A Compact Vision-Language Model for Industrial Inspection Across Edge GPUs},
author = {Sudharsanan, Myan and Puppala, Hemmanth and Daher, Gavin and Rosenbloom, Jack},
institution = {Denali Advanced Integration},
year = {2026},
note = {Model: https://huggingface.co/Denali-AI/Rainier-VL-2B-MTP}
}
License
Released under the Apache License 2.0. Copyright 2026 Denali AI.
- Downloads last month
- 157



