EAT: An Enhancer for Aesthetics-Oriented Transformers
Shuai He, Anlong Ming, Shuntian Zheng, Haobin Zhong, Huadong Ma
Abstract
Transformers have shown great potential in various vision tasks, but none of them have surpassed the best CNN model on image aesthetics assessment (IAA) tasks. IAA is a challenging task in multimedia systems that requires attention to both foreground and background, as well as robustness to noisy and redundant labels. The global and dense attention mechanism of Transformers, designed for saliency-oriented tasks, may miss important aesthetic information in the background, increase the computational cost and slow down the convergence on IAA tasks. To address these issues, we propose an Enhancer for Aesthetics-Oriented Transformers (EAT). EAT uses a deformable, sparse and data-dependent attention mechanism that learns where to focus and how to refine attention by offsets. EAT also guides the offsets to balance the attention between foreground and background according to dedicated rules. Our EAT-enhanced Transformers outperform the previous methods on four representative datasets with fewer training epochs. Code is available in https://github.com/woshidandan/Image-Aesthetics-Assessment
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers11
- ELTA: An Enhancer against Long-Tail for Aesthetics-oriented ModelsLimin Liu, Shuai He, Anlong Ming, Rui Xie et al.ICML 2024 · 13 citations
- Rethinking No-reference Image Exposure Assessment from Holism to Pixel: Models, Datasets and BenchmarksShuai He, Shuntian Zheng, Anlong Ming, Banyu Wu et al.NeurIPS 2024 · 4 citations
- TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center LearningJinglun Li, Xinyu Zhou, Kaixun Jiang, Lingyi Hong et al.ACM MM 2024 · 1 citation
- FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face RetouchingHui Li, Xiaoyu Ren, Hongjiu Yu, Ying Chen et al.ICCV 2025 · 1 citation
- Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language ModelsXingyuan Ma, Shuai He, Anlong Ming, Haobin Zhong et al.AAAI 2026
Related papers
- Thinking Image Color Aesthetics Assessment: Models, Datasets and BenchmarksShuai He, Anlong Ming, Yaqi Li, Jinyuan Sun et al.ICCV 2023 · 36 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- Image Harmonization with TransformerZonghui Guo, Dongsheng Guo, Haiyong Zheng, Zhaorui Gu et al.ICCV 2021 · 95 citations
- Accurate Image Restoration with Attention Retractable TransformerJiale Zhang, Yulun Zhang, Jinjin Gu, Yongbing Zhang et al.ICLR 2023 · 47 citations
- Object-level Attention for Aesthetic Rating Distribution PredictionJingwen Hou, Sheng Yang, Weisi LinACM MM 2020 · 31 citations
