BrightGuo's picture
ProcObject-10K object-centric SFT checkpoint (Qwen3-VL-4B) + model card
d6fb04d verified
|
Raw History Blame Contribute Delete
4.14 kB
---
license: cc-by-nc-4.0
base_model: Qwen/Qwen3-VL-4B-Instruct
library_name: transformers
pipeline_tag: video-text-to-text
datasets:
- BrightGuo/ProcObject-10K
language:
- en
tags:
- qwen3-vl
- video-qa
- temporal-grounding
- procedural-video
- object-centric
---
# ProcObject-Qwen3-VL-4B
**Qwen3-VL-4B-Instruct fine-tuned with object-centric SFT on ProcObject-10K** — the fine-tuned model
of *ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos*
(NeurIPS 2026, Evaluations & Datasets Track).
[Paper](https://arxiv.org/abs/2512.03479) · [Code](https://github.com/WenliangGuo/ProcObject-10K) ·
[Dataset](https://huggingface.co/datasets/BrightGuo/ProcObject-10K)
Given frames of an instructional video clip and a question about an object, the model answers and
localizes the supporting moments:
```json
{"answer": "The tortilla starts whole on the table, is torn into two pieces, and finally placed into a bowl.", "evidence": [[0, 8], [14, 16]]}
```
`evidence` intervals are in seconds from the start of the clip.
## Training
Two stages on the 9,472 ProcObject-10K training QA (8 GPUs, bf16):
1. **Evidence-prediction SFT** — LoRA (r=16, α=32) on the language model's linear layers, trained on
the JSON answer + evidence target; 3 epochs, lr 5e-5; 2 FPS, ≤48 frames.
2. **Object-centric SFT** — continues the same LoRA, adds LoRA on the attention layers of the top
third of the vision blocks, tunes the vision merger, and trains two auxiliary heads: a spatial
head supervised by soft patch masks from Grounding DINO boxes of LLM-extracted object phrases,
and a temporal head supervised by frame-in-evidence labels.
`L = L_gen + 0.05·L_spl + 0.10·L_tmp`, 3 epochs.
The auxiliary heads are training-only and are **not** part of this checkpoint: it is a standard
merged Qwen3-VL model with no extra inference cost. Full recipe and code:
[github.com/WenliangGuo/ProcObject-10K](https://github.com/WenliangGuo/ProcObject-10K).
## Results on the ProcObject-10K test set
With the released evaluation harness (48 frames; J. = 0–5 LLM judge, mean of Qwen3-4B and
Llama-3.2-3B):
| | S. | B. | J. | mIoU | mIoP | mIoG | [email protected] |
|---|:-:|:-:|:-:|:-:|:-:|:-:|:-:|
| Qwen3-VL-4B-Instruct (zero-shot, paper) | 73.3 | 89.2 | – | 38.6 | 63.1 | 54.1 | – |
| **this model** | **80.8** | **92.0** | 3.25 | **45.3** | **66.8** | **56.4** | 63.5 |
mIoU is 51.9 on Multi-hop Reasoning and 35.9 on Needle-in-a-Haystack questions.
## Usage
The reported numbers use the benchmark protocol: 48 frames sampled uniformly from the clip, resized
to 512×512, each sent as an image preceded by `[Timestamp: <t>s]`, after the system prompt in
[`benchmark/prompts/system_prompt.txt`](https://github.com/WenliangGuo/ProcObject-10K/blob/master/benchmark/prompts/system_prompt.txt),
greedy decoding. The easiest way to reproduce them is the repository's harness:
```bash
git clone https://github.com/WenliangGuo/ProcObject-10K && cd ProcObject-10K
pip install -r requirements/eval.txt
# build data/clips/ first (preprocess/README.md), then:
MODEL=BrightGuo/ProcObject-Qwen3-VL-4B NAME=procobject_qwen3vl_4b NUM_FRAMES=48 \
bash benchmark/scripts/run_sharded.sh 0
bash benchmark/scripts/evaluate.sh benchmark/pred_results/procobject_qwen3vl_4b_predictions.json
```
Or serve it with vLLM (`vllm serve BrightGuo/ProcObject-Qwen3-VL-4B --limit-mm-per-prompt '{"image": 64}'`)
and send the same message layout through the OpenAI-compatible API (`benchmark/models/runners.py`,
`VLLMRunner`).
Tested with vLLM 0.17.1 / transformers 4.57.6.
## License
CC BY-NC 4.0, following the ProcObject-10K annotations it was trained on; the base model
Qwen3-VL-4B-Instruct is Apache-2.0.
## Citation
```bibtex
@inproceedings{guo2026procobject,
title = {{ProcObject-10K}: Benchmarking Object-Centric Procedural Understanding in Instructional Videos},
author = {Guo, Wenliang and Kong, Yu},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
year = {2026},
url = {https://arxiv.org/abs/2512.03479}
}
```