EAT: An Enhancer for Aesthetics-Oriented Transformers
Shuai He, Anlong Ming, Shuntian Zheng, Haobin Zhong, Huadong Ma
摘要
Transformers have shown great potential in various vision tasks, but none of them have surpassed the best CNN model on image aesthetics assessment (IAA) tasks. IAA is a challenging task in multimedia systems that requires attention to both foreground and background, as well as robustness to noisy and redundant labels. The global and dense attention mechanism of Transformers, designed for saliency-oriented tasks, may miss important aesthetic information in the background, increase the computational cost and slow down the convergence on IAA tasks. To address these issues, we propose an Enhancer for Aesthetics-Oriented Transformers (EAT). EAT uses a deformable, sparse and data-dependent attention mechanism that learns where to focus and how to refine attention by offsets. EAT also guides the offsets to balance the attention between foreground and background according to dedicated rules. Our EAT-enhanced Transformers outperform the previous methods on four representative datasets with fewer training epochs. Code is available in https://github.com/woshidandan/Image-Aesthetics-Assessment
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper11
- ELTA: An Enhancer against Long-Tail for Aesthetics-oriented ModelsLimin Liu, Shuai He, Anlong Ming, Rui Xie 等ICML 2024 · 被引用 13 次
- Rethinking No-reference Image Exposure Assessment from Holism to Pixel: Models, Datasets and BenchmarksShuai He, Shuntian Zheng, Anlong Ming, Banyu Wu 等NeurIPS 2024 · 被引用 4 次
- TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center LearningJinglun Li, Xinyu Zhou, Kaixun Jiang, Lingyi Hong 等ACM MM 2024 · 被引用 1 次
- FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face RetouchingHui Li, Xiaoyu Ren, Hongjiu Yu, Ying Chen 等ICCV 2025 · 被引用 1 次
- Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language ModelsXingyuan Ma, Shuai He, Anlong Ming, Haobin Zhong 等AAAI 2026
相关 Paper
- Thinking Image Color Aesthetics Assessment: Models, Datasets and BenchmarksShuai He, Anlong Ming, Yaqi Li, Jinyuan Sun 等ICCV 2023 · 被引用 36 次
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li 等CVPR 2022 · 被引用 835 次
- Image Harmonization with TransformerZonghui Guo, Dongsheng Guo, Haiyong Zheng, Zhaorui Gu 等ICCV 2021 · 被引用 95 次
- Accurate Image Restoration with Attention Retractable TransformerJiale Zhang, Yulun Zhang, Jinjin Gu, Yongbing Zhang 等ICLR 2023 · 被引用 47 次
- Object-level Attention for Aesthetic Rating Distribution PredictionJingwen Hou, Sheng Yang, Weisi LinACM MM 2020 · 被引用 31 次
