Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping
Tianxiang Du, Hulingxiao He, Yuxin Peng
Abstract
The cropped image focuses on the architectural details of the building, emphasizing grandeur and symmetry. By removing the sky and lower portion, the composition becomes more balanced and centered, drawing attention to the columns and dome for a stronger visual impact. Venus (ours):The composition appears somewhat lacking in a clear focal point and a sense of emotional depth, as the linear perspective of the benches holds potential but is not fully utilized. To enhance the overall visual impact, it is recommended to adjust the shooting angle at night, utilizing the circular ring as a framing element to guide the viewer's gaze toward the distant skyscrapers, thereby strengthening the visual focus and enhancing the sense of depth.
AesExpert:The clarity of this image is high, and the details are depicted vividly and accurately, with sharp edges. The color variety is abundant, and the combination is natural and harmonious, resulting in a pleasing visual impression. The composition adopts a central composition, highlighting the subject. Therefore, this image looks quite beautiful.
Re-captured Photo Reshoot Following Aesthetic Guidance Current MLLMs Task 2: Aesthetic Cropping (after capture) Re-captured Photo GPT-4o AesExpert Venus (ours) Fail to Identify Issues
GPT-4o: This photo effectively uses perspective and composition to create depth and interest. The leading lines of the wooden planks draw the viewer's eye toward the background, where tall buildings and trees form a dynamic contrast between urban and natural elements… The composition harmonizes these details into a balanced and engaging image.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e20ca4d-c93f-47c9-8d66-3a5874daed55Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen et al.ICML 2024 · 499 citations
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen et al.ICLR 2024 · 258 citations
- Personalized Image Aesthetics Assessment with Rich AttributesYuzhe Yang, Liwu Xu, Leida Li, Nan Qie et al.CVPR 2022 · 84 citations
Related papers
- InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic ImagesXiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen et al.ACM MM 2025
- ProCrop: Learning Aesthetic Image Cropping from Professional CompositionsKe Zhang, Tianyu Ding, Jiachen Jiang, Tianyi Chen et al.AAAI 2026 · 3 citations
- Doppelgangers++: Improved Visual Disambiguation with Geometric 3D FeaturesYuanbo Xiangli, Ruojin Cai, Hanyu Chen, Jeffrey Byrne et al.CVPR 2025
- Photography Perspective Composition: Towards Aesthetic Perspective RecommendationLujian Yao, Siming Zheng, Xinbin Yuan, Zhuoxuan Cai et al.NeurIPS 2025 · 2 citations
- Cropper: Vision-Language Model for Image Cropping through In-Context LearningSeung Hyun Lee, Jijun Jiang, Yiran Xu, Zhuofang Li et al.CVPR 2025
