ProCrop: Learning Aesthetic Image Cropping from Professional Compositions
Ke Zhang, Tianyu Ding, Jiachen Jiang, Tianyi Chen, Ilya Zharkov, Vishal M. Patel, Luming Liang
Abstract
Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rulebased and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a retrieval-based method that leverages professional photography to guide cropping decisions. By fusing features from professional photographs with those of the query image, ProCrop learns from professional compositions, significantly boosting performance. Additionally, we present a large-scale dataset of 242K weakly-annotated images, generated by out-painting professional images and iteratively refining diverse crop proposals. This composition-aware dataset generation offers diverse high-quality crop proposals guided by aesthetic principles and becomes the largest publicly available dataset for image cropping. Extensive experiments show that ProCrop significantly outperforms existing methods in both supervised and weakly-supervised settings. Notably, when trained on the new dataset, our Pro-Crop surpasses previous weakly-supervised methods and even matches fully supervised approaches. Both the code and dataset will be made publicly available to advance research in image aesthetics and composition analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PhotoFramer: Multi-modal Image Composition InstructionZhiyuan You, Ke Wang, He Zhang, Xin Cai et al.CVPR 2026 · 8 citations
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and CroppingTianxiang Du, Hulingxiao He, Yuxin PengCVPR 2026 · 3 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
Related papers
- InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic ImagesXiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen et al.ACM MM 2025
- Photography Perspective Composition: Towards Aesthetic Perspective RecommendationLujian Yao, Siming Zheng, Xinbin Yuan, Zhuoxuan Cai et al.NeurIPS 2025 · 2 citations
- Composing Photos Like a PhotographerChaoyi Hong, Shuaiyuan Du, Ke Xian, Hao Lu et al.CVPR 2021
- Image Cropping with Composition and Saliency Aware Aesthetic Score MapYi Tu, Li Niu, Weijie Zhao, Dawei Cheng et al.AAAI 2020 · 55 citations
- Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and UnderstandingZhaoran Zhao, Peng Lu, Anran Zhang, Peipei Li et al.CVPR 2025
