InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic Images
Xiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen, Tong Zhu, Leida Li
摘要
Aesthetic Image Cropping (AIC) aims to improve the visual appeal of images by removing redundant content while preserving attractive elements. Despite the encouraging progresses achieved in data-driven approaches, most existing models struggle to understand user intentions, particularly for diversified scenes with multiple subjects. Moreover, they can only provide cropping results without explanations, which further restricts their usability in real-world applications. Motivated by the above facts, we introduce InstructCrop : a multimodal large language model (MLLM)-based AIC framework, which can understand user instructions and provide explanatory reasons for cropping results. Specifically, we first build a multimodal Image Cropping Instruction Tuning (ICIT) dataset through a cost-effective paradigm by generating high-quality instruction tuning data based on the existing cropping datasets. Then, we embed dynamic domain knowledge into the cropping model by integrating cropping-aware experts of aesthetic assessment and composition classification. Finally, we adapt MLLMs to generate the cropping results and corresponding explanations. Quantitative and qualitative experiments on three benchmark datasets demonstrate that InstructCrop enables effective and interpretable image cropping, which aligns better with user intentions. Data and code are available at https://github.com/sxfly99/InstructCrop.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level UnderstandingShuo Cao, Nan Ma, Jiayang Li, Xiaohui Li 等CVPR 2026 · 被引用 38 次
- AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics PerceptionYipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan 等ACM MM 2024 · 被引用 34 次
- ProCrop: Learning Aesthetic Image Cropping from Professional CompositionsKe Zhang, Tianyu Ding, Jiachen Jiang, Tianyi Chen 等AAAI 2026 · 被引用 3 次
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
- Cropper: Vision-Language Model for Image Cropping through In-Context LearningSeung Hyun Lee, Jijun Jiang, Yiran Xu, Zhuofang Li 等CVPR 2025
