Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
Ruixiang Jiang, Chang Wen Chen
摘要
The rapid technical progress of generative art (GenArt) has democratized the creation of visually appealing imagery. However, achieving genuine artistic impact - the kind that resonates with viewers on a deeper, more meaningful level - remains formidable as it requires a sophisticated aesthetic sensibility. This sensibility involves a multifaceted cognitive process extending beyond mere visual appeal, which is often overlooked by current computational methods. This paper pioneers an approach to capture this complex process by investigating how the reasoning capabilities of Multimodal LLMs (MLLMs) can be effectively elicited to perform aesthetic judgment. Our analysis reveals a critical challenge: MLLMs exhibit a tendency towards hallucinations during aesthetic reasoning, characterized by subjective opinions and unsubstantiated artistic interpretations. We further demonstrate that these hallucinations can be suppressed by employing an evidence-based and objective reasoning process, as substantiated by our proposed baseline, ArtCoT. MLLMs prompted by this principle produce multifaceted, in-depth aesthetic reasoning that aligns significantly better with human judgment. These findings have direct applications in areas such as AI art tutoring and as reward models for image generation. Ultimately, we hope this work paves the way for AI systems that can truly understand, appreciate, and contribute to art that aligns with human aesthetic values. Project homepage: https://github.com/songrise/MLLM4Art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Can Vision-Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset PerspectiveRuichuan An, Shizhao Sun, Danqing Huang, Mingxi Cheng 等ICLR 2026 · 被引用 6 次
- Fine-grained Zero-Shot Object DetectionHongxu Ma, Chenbo Zhang, Lu Zhang, Jiaogen Zhou 等ACM MM 2025 · 被引用 4 次
- DiffArtist: Towards Structure and Appearance Controllable Image StylizationRuixiang Jiang, Chang Wen ChenACM MM 2025 · 被引用 4 次
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and CroppingTianxiang Du, Hulingxiao He, Yuxin PengCVPR 2026 · 被引用 3 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy OptimizationBoyang Liu, Yifan Hu, Senjie Jin, Shihan Dou 等ICLR 2026 · 被引用 6 次
- InstructCrop: Teaching Multimodal Large Language Models to Crop Aesthetic ImagesXiangfei Sheng, Pangu Xie, Weidong Zou, Pengfei Chen 等ACM MM 2025
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference AlignmentRishab Parthasarathy, Jasmine Collins, Cory StephensonAAAI 2026
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan 等ICLR 2026 · 被引用 9 次
- Cantor: Inspiring Multimodal Chain-of-Thought of MLLMTimin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu 等ACM MM 2024 · 被引用 20 次
