AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
Sohan Patnaik, Rishabh Jain, Balaji Krishnamurthy, Mausoom Sarkar
Abstract
Visual layouts are essential in graphic design fields such as advertising, posters, and web interfaces. The application of generative models for content-aware layout generation has recently gained traction. However, these models fail to understand the contextual aesthetic requirements of layout design and do not align with human-like preferences, primarily treating it as a prediction task without considering the final rendered output. To overcome these problems, we offer Aesthetic-Aware Preference Alignment (AAPA), a novel technique to train a Multi-modal Large Language Model (MLLM) for layout prediction that uses MLLM's aesthetic preferences for Direct Preference Optimization over graphic layouts. We propose a data filtering protocol utilizing our layout-quality heuristics for AAPA to ensure training happens on high-quality layouts. Additionally, we introduce a novel evaluation metric that uses another MLLM to compute the win rate of the generated layout against the ground-truth layout based on aesthetics criteria. We also demonstrate the applicability of AAPA for MLLMs of varying scales (1B to 8B parameters) and LLM families (Qwen, Phi, InternLM). By conducting thorough qualitative and quantitative analyses, we verify the efficacy of our approach on two challenging benchmarks -Crello and Webui, showcasing 17%, and 16% improvement over current State-of-The-Art methods, thereby highlighting the potential of MLLMs in aesthetic-aware layout generation. Results can be downloaded from Project Webpage
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design GenerationJianyu LAI, Sixiang Chen, Jialin Gao, Hengyu Shi et al.CVPR 2026 · 5 citations
- LayerD: Decomposing Raster Graphic Designs into LayersTomoyuki Suzuki, Kang-Jun Liu, Naoto Inoue, Kota YamaguchiICCV 2025 · 3 citations
Builds on23
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
Related papers
- Seeing is Improving: Visual Feedback for Iterative Text Layout RefinementJunrong Guo, Shancheng Fang, Yadong Qu, Hongtao XieCVPR 2026 · 2 citations
- Two-stage Content-Aware Layout Generation for Poster DesignsShang Chai, Liansheng Zhuang, Fengying Yan, Zihan ZhouACM MM 2023 · 3 citations
- PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout GenerationHsiaoYuan Hsu, Yuxin PengCVPR 2025
- Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and AlgorithmsMiaosen Zhang, Yixuan Wei, Zhen Xing, Yifei Ma et al.NeurIPS 2024 · 2 citations
- Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMsZitian Wang, Yue Liao, Kang Rong, Fengyun Rao et al.ICCV 2025
