PhotoFramer: Multi-modal Image Composition Instruction
Paper • 2512.00993 • Published
Composition Assessment Model ( project page / codes / paper ) weights fully fine-tuned on AVA, CADB, GAIC, and KUPCP datasets.
The metrics for composition assessment is SRCC / PLCC, and the metric for composition classification is accuracy.
| Dataset | CADB Assessment | GAIC Assessment | AVA Assessment | CADB Classification |
|---|---|---|---|---|
| VFN | 0.052 / 0.049 | 0.152 / 0.162 | 0.139 / 0.142 | - |
| VEN | 0.084 / 0.082 | 0.410 / 0.428 | 0.232 / 0.241 | - |
| AutoPhoto | 0.065 / 0.079 | 0.407 / 0.427 | 0.604 / 0.613 | - |
| Q-Align | 0.561 / 0.557 | 0.169 / 0.178 | 0.809 / 0.804 | - |
| Qwen2.5-VL-32B | 0.420 / 0.426 | 0.195 / 0.205 | 0.527 / 0.492 | 0.101 |
| Our Model (7B) | 0.763 / 0.777 | 0.795 / 0.805 | 0.825 / 0.828 | 0.583 |
If you find our work useful for your research and applications, please cite using the BibTeX:
@inproceedings{photoframer,
title={PhotoFramer: Multi-modal Image Composition Instruction},
author={You, Zhiyuan and Wang, Ke and Zhang, He and Cai, Xin and Gu, Jinjin and Xue, Tianfan and Dong, Chao and Zhang, Zhoutong},
booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2026}
}
Base model
Qwen/Qwen2.5-VL-7B-Instruct