inblick YOLOv8n: quantization-aware trained int8 detectors

Six YOLOv8n networks (plus two larger YOLOv8s, below) trained with quantization-aware training (QAT) for a custom int8 FPGA CNN accelerator (the inblick project, AMD AC701 / Artix-7). Two run at 320x320; four run at the non-square 16:9 sizes 640x384 and 480x288 (no letterboxing into a square, so a 16:9 camera wastes about 6 % of the input instead of 44 %). The QAT simulates the accelerator's exact int8 arithmetic:

  • per-tensor asymmetric int8 activations with half-to-even rounding;
  • per-channel symmetric int8 weights;
  • int32 bias;
  • SiLU as Q(x), Q(sigmoid), Q(x*sigmoid) through a lookup table.

The exported int8 graphs run bit-exact on the hardware.

Folder Network Classes Use it for
coco_qat_320/ standard YOLOv8n, QAT fine-tuned on COCO 80 COCO general objects
crowdhuman_qat_320/ YOLOv8n with a 1-class head, trained on CrowdHuman full-body boxes, then QAT person people, crowds, dance, pedestrians
coco_qat_640x384/ the COCO network, QAT at 640x384 80 COCO general objects, best accuracy (19.7 fps on the board)
coco_qat_480x288/ the COCO network, QAT at 480x288 80 COCO general objects, 34.5 fps
crowdhuman_qat_640x384/ the CrowdHuman network, fp32 fine-tune then QAT at 640x384 person people, best accuracy (21.6 fps)
crowdhuman_qat_480x288/ the CrowdHuman network, QAT at 480x288 person people, 37.9 fps
crowdhuman_v8s_qat_480x288/ YOLOv8s, 1-class CrowdHuman, QAT at 480x288 person people, more accurate, 37.6 fps on the KV260
crowdhuman_v8s_qat_640x384/ YOLOv8s, 1-class CrowdHuman, QAT at 640x384 person people, best accuracy, 21.7 fps on the KV260

Licence: AGPL-3.0 (derived from Ultralytics YOLOv8). The training and export source code is in source/ (AGPL-3.0). The CrowdHuman model is additionally bound by the CrowdHuman dataset terms: non-commercial research use only. See LICENSE-NOTICE.md.

Results (320x320, int8)

COCO val2017 (5000 images), mAP50-95 / mAP50

Model mAP50-95 mAP50
official YOLOv8n fp32 28.49 41.47
post-training quantization (PTQ) int8 27.07 39.47
coco_qat_320 (QAT int8) 28.68 42.39

QAT recovers the whole int8 loss and ends slightly above fp32.

CrowdHuman val (4370 images), AP50-95 / AP50, int8: crowdhuman_qat_320 scores 38.15 / 68.66.

Multi-object tracking, held-out MOT17 (05/10/11/13) and DanceTrack (12 sequences), at 46 fps, with the ByteTrack/OC-SORT-style tracker of the project. Values are HOTA / IDF1:

Detector MOT17 DanceTrack
YOLOv8n COCO, PTQ int8 (previous baseline) 30.5 / 34.9 24.4 / 27.0
coco_qat_320 31.3 / 37.3 26.3 / 28.5
crowdhuman_qat_320 36.1 / 43.8 28.5 / 31.2

Rows 2 and 3 use the presets in each folder's tracker_presets.json, which keep the board's on-board post-processing at full frame rate. With 4 and 12 held-out sequences, differences under about 1 HOTA are noise.

Non-square 16:9 models (640x384 and 480x288)

Accuracy (int8, onnxruntime; COCO val2017 mAP50-95 / mAP50 and CrowdHuman val AP50-95 / AP50), accelerator speed on the AC701 board (measured cycles per frame at 150 MHz) and held-out tracking at that frame rate with each folder's tuned preset (HOTA / IDF1; MOT17 and DanceTrack):

Model AP Board fps MOT17 DanceTrack
coco_qat_320 (reference) 28.68 / 42.39 46.3 31.3 / 37.3 (46 fps) 26.3 / 28.5
coco_qat_480x288 31.13 / 45.38 34.5 34.1 / 42.2 29.7 / 31.9
coco_qat_640x384 34.45 / 49.52 19.7 36.6 / 45.5 32.7 / 36.1
crowdhuman_qat_320 (reference) 38.15 / 68.66 50.7 36.1 / 43.8 (46 fps) 28.5 / 31.2
crowdhuman_qat_480x288 43.05 / 74.24 37.9 39.4 / 48.9 31.5 / 34.2
crowdhuman_qat_640x384 47.77 / 79.10 21.6 40.8 / 50.2 34.7 / 38.1

The COCO val images are 4:3, so their AP at 640x384 understates what a 16:9 camera gets. Each folder's README compares three ways of getting from the 320 network to the new size: (a) the 320 QAT weights and scales on the new graph, (b) a fresh PTQ export of them, (c) a short QAT re-train (for CrowdHuman, also an fp32 fine-tune at the new size first). All of them compile for the same accelerator bitstream; 640 is the widest input it can take (in_width * ceil(Cin / 16) of the stride-16 layer reaches the 960-word row buffer).

YOLOv8s models (KV260)

Two larger YOLOv8s detectors, 1-class CrowdHuman, same QAT arithmetic and export as above. They need the KV260 accelerator target (cnn_accel_v1_kv260, partial-sum ISA 3.2); the board speed is for the 4-core 250 MHz build. CrowdHuman val (4370 images), int8, onnxruntime:

Model AP50-95 / AP50 int8 fp32 Board fps YOLOv8n of the same size (AP50-95 / AP50)
crowdhuman_v8s_qat_480x288 47.67 / 78.60 48.05 / 78.83 37.6 43.05 / 74.24
crowdhuman_v8s_qat_640x384 52.93 / 82.95 53.01 / 83.03 21.7 47.77 / 79.10

Recipe: yolov8s.pt, 1-class head re-initialised, 60 fp32 epochs, 15 QAT epochs (lr 1e-4, batch 32 at 480x288 and 16 at 640x384), no 320 stage. No tracking benchmark was run for them; their tracker_presets.json are the YOLOv8n presets of the same size.

Tracker presets and gmc

tracker_presets.json of the 640x384 and 480x288 CrowdHuman folders (YOLOv8n and YOLOv8s) now sets "gmc": 2 in the _board_dt and _board_cmc presets: image camera-motion compensation, a similarity estimated by block matching on the detector's input image (bit-exact on the board; _board_cmc keeps cmc 3 as the fallback). Held-out MOT17 (YOLOv8n, board model): 480x288 41.8 / 51.3 HOTA / IDF1 with gmc 2 against 39.4 / 48.9 without; 640x384 43.9 / 55.1 against 40.8 / 50.2. The plain _board presets are unchanged.

Files (per folder)

File What
*.pt trained weights, Ultralytics checkpoint (YOLOv8n: BN-fused fp16; the recipe is in train_args)
qdq.onnx int8 QDQ ONNX graph with the learned scales; the raw detection heads (no NMS in the graph)
manifest.json export manifest: input quantisation, head layout, class names, weights sha256, QAT recipe
sample_input.npy a sample int8 input for golden-model checks
metrics.json all evaluation numbers
tracker_presets.json tracker presets tuned for this detector
README.md the detailed recipe, results and verification

Using the weights outside inblick

The .pt files are ordinary Ultralytics YOLOv8n checkpoints:

from ultralytics import YOLO
model = YOLO("crowdhuman_qat_320/crowdhuman_qat_320.pt")
results = model("image.jpg", imgsz=320)

They were fine-tuned to suit int8 rounding, so in fp32 they are slightly below the official weights. Their benefit shows when they are quantized the same way: per-tensor asymmetric activations and per-channel symmetric weights at 320x320. qdq.onnx runs in onnxruntime and outputs the six raw YOLOv8 heads (3 box-distribution and 3 class-score maps); decode them as YOLOv8 does (DFL with reg_max 16, then NMS).

Training

  • Hardware: one RTX 5090. About 35 min for the COCO QAT, and about 27 min for the CrowdHuman model (fp32 fine-tune plus QAT).
  • QAT optimiser: AdamW, lr 1e-4 for the weights and 2e-4 for the log activation scales, cosine schedule, EMA 0.999, batch 64, seed 0. Ultralytics 8.4.164, torch 2.14.0.
  • COCO: 10 QAT epochs on train2017 from the official yolov8n.pt; epoch 6 kept.
  • CrowdHuman:
    • start: the official yolov8n.pt with a re-initialised 1-class head;
    • fp32 fine-tune: 60 epochs on full-body (fbox) boxes, with mask and ignore regions dropped;
    • then 15 QAT epochs.
    • The CrowdHuman data came from the Hugging Face mirror sshao0516/CrowdHuman of the official release.

Full details are in each folder's README.md.

Source

source/ holds the QAT training and export code that produced these weights (AGPL-3.0): the fake-quantised QDQ network, the trainer, the YOLOv8n int8 export and the COCO evaluation, with its environment in requirements-train.txt. See source/README.md.

Downloads last month
508
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train ru551n/inblick-yolov8n