Video-Text-to-Text
Transformers
Safetensors
English
qwen3_vl
image-text-to-text
qwen3-vl
video-qa
temporal-grounding
procedural-video
object-centric
Instructions to use BrightGuo/ProcObject-Qwen3-VL-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BrightGuo/ProcObject-Qwen3-VL-4B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BrightGuo/ProcObject-Qwen3-VL-4B") model = AutoModelForMultimodalLM.from_pretrained("BrightGuo/ProcObject-Qwen3-VL-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from BrightGuo/ProcObject-Qwen3-VL-4B: direct link, hf CLI and curl.
- Browser
- Download file 4.14 kB
-
https://huggingface.co/BrightGuo/ProcObject-Qwen3-VL-4B/resolve/main/README.md
- Command line
-
hf download hf://BrightGuo/ProcObject-Qwen3-VL-4B/README.md
-
curl -L -o README.md https://huggingface.co/BrightGuo/ProcObject-Qwen3-VL-4B/resolve/main/README.md
4.14 kB
| license: cc-by-nc-4.0 | |
| base_model: Qwen/Qwen3-VL-4B-Instruct | |
| library_name: transformers | |
| pipeline_tag: video-text-to-text | |
| datasets: | |
| - BrightGuo/ProcObject-10K | |
| language: | |
| - en | |
| tags: | |
| - qwen3-vl | |
| - video-qa | |
| - temporal-grounding | |
| - procedural-video | |
| - object-centric | |
| # ProcObject-Qwen3-VL-4B | |
| **Qwen3-VL-4B-Instruct fine-tuned with object-centric SFT on ProcObject-10K** — the fine-tuned model | |
| of *ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos* | |
| (NeurIPS 2026, Evaluations & Datasets Track). | |
| [Paper](https://arxiv.org/abs/2512.03479) · [Code](https://github.com/WenliangGuo/ProcObject-10K) · | |
| [Dataset](https://huggingface.co/datasets/BrightGuo/ProcObject-10K) | |
| Given frames of an instructional video clip and a question about an object, the model answers and | |
| localizes the supporting moments: | |
| ```json | |
| {"answer": "The tortilla starts whole on the table, is torn into two pieces, and finally placed into a bowl.", "evidence": [[0, 8], [14, 16]]} | |
| ``` | |
| `evidence` intervals are in seconds from the start of the clip. | |
| ## Training | |
| Two stages on the 9,472 ProcObject-10K training QA (8 GPUs, bf16): | |
| 1. **Evidence-prediction SFT** — LoRA (r=16, α=32) on the language model's linear layers, trained on | |
| the JSON answer + evidence target; 3 epochs, lr 5e-5; 2 FPS, ≤48 frames. | |
| 2. **Object-centric SFT** — continues the same LoRA, adds LoRA on the attention layers of the top | |
| third of the vision blocks, tunes the vision merger, and trains two auxiliary heads: a spatial | |
| head supervised by soft patch masks from Grounding DINO boxes of LLM-extracted object phrases, | |
| and a temporal head supervised by frame-in-evidence labels. | |
| `L = L_gen + 0.05·L_spl + 0.10·L_tmp`, 3 epochs. | |
| The auxiliary heads are training-only and are **not** part of this checkpoint: it is a standard | |
| merged Qwen3-VL model with no extra inference cost. Full recipe and code: | |
| [github.com/WenliangGuo/ProcObject-10K](https://github.com/WenliangGuo/ProcObject-10K). | |
| ## Results on the ProcObject-10K test set | |
| With the released evaluation harness (48 frames; J. = 0–5 LLM judge, mean of Qwen3-4B and | |
| Llama-3.2-3B): | |
| | | S. | B. | J. | mIoU | mIoP | mIoG | [email protected] | | |
| |---|:-:|:-:|:-:|:-:|:-:|:-:|:-:| | |
| | Qwen3-VL-4B-Instruct (zero-shot, paper) | 73.3 | 89.2 | – | 38.6 | 63.1 | 54.1 | – | | |
| | **this model** | **80.8** | **92.0** | 3.25 | **45.3** | **66.8** | **56.4** | 63.5 | | |
| mIoU is 51.9 on Multi-hop Reasoning and 35.9 on Needle-in-a-Haystack questions. | |
| ## Usage | |
| The reported numbers use the benchmark protocol: 48 frames sampled uniformly from the clip, resized | |
| to 512×512, each sent as an image preceded by `[Timestamp: <t>s]`, after the system prompt in | |
| [`benchmark/prompts/system_prompt.txt`](https://github.com/WenliangGuo/ProcObject-10K/blob/master/benchmark/prompts/system_prompt.txt), | |
| greedy decoding. The easiest way to reproduce them is the repository's harness: | |
| ```bash | |
| git clone https://github.com/WenliangGuo/ProcObject-10K && cd ProcObject-10K | |
| pip install -r requirements/eval.txt | |
| # build data/clips/ first (preprocess/README.md), then: | |
| MODEL=BrightGuo/ProcObject-Qwen3-VL-4B NAME=procobject_qwen3vl_4b NUM_FRAMES=48 \ | |
| bash benchmark/scripts/run_sharded.sh 0 | |
| bash benchmark/scripts/evaluate.sh benchmark/pred_results/procobject_qwen3vl_4b_predictions.json | |
| ``` | |
| Or serve it with vLLM (`vllm serve BrightGuo/ProcObject-Qwen3-VL-4B --limit-mm-per-prompt '{"image": 64}'`) | |
| and send the same message layout through the OpenAI-compatible API (`benchmark/models/runners.py`, | |
| `VLLMRunner`). | |
| Tested with vLLM 0.17.1 / transformers 4.57.6. | |
| ## License | |
| CC BY-NC 4.0, following the ProcObject-10K annotations it was trained on; the base model | |
| Qwen3-VL-4B-Instruct is Apache-2.0. | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{guo2026procobject, | |
| title = {{ProcObject-10K}: Benchmarking Object-Centric Procedural Understanding in Instructional Videos}, | |
| author = {Guo, Wenliang and Kong, Yu}, | |
| booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track}, | |
| year = {2026}, | |
| url = {https://arxiv.org/abs/2512.03479} | |
| } | |
| ``` | |