Does text attract attention on e-commerce images: A novel saliency prediction dataset and method
Lai Jiang, Yifei Li, Shengxi Li, Mai Xu, Se Lei, Yichen Guo, Bo Huang
Abstract
E-commerce images are playing a central role in attracting people's attention when retailing and shopping online, and an accurate attention prediction is of significant importance for both customers and retailers, where its research is yet to start. In this paper, we establish the first dataset of saliency e-commerce images (SalECI), which allows for learning to predict saliency on the e-commerce images. We then provide specialized and thorough analysis by high-lighting the distinct features of e-commerce images, e.g., non-locality and correlation to text regions. Correspondingly, taking advantages of the non-local and self-attention mechanisms, we propose a salient SWin-Transformer back-bone, followed by a multi-task learning with saliency and text detection heads, where an information flow mechanism is proposed to further benefit both tasks. Experimental results have verified the state-of-the-art performances of our work in the e-commerce scenario.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction RecognitionYuchen Zhou, Linkai Liu, Chao GouCVPR 2024 · 13 citations
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural ImagesOzgur Kara, Harris Nisar, James M. RehgNeurIPS 2025 · 7 citations
- Unsupervised Salient Instance DetectionXin Tian, Ke Xu, Rynson W. H. LauCVPR 2024 · 2 citations
- Attend to Anything: Foundation Model for Unified Human Attention ModelingWenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun ZhaoICML 2026
- DINN360: Deformable Invertible Neural Network for Latitude-aware 360° Image RescalingYichen Guo, Mai Xu, Lai Jiang, Leonid Sigal et al.CVPR 2023
Builds on4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
- DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modelingAkis Linardos, Matthias Kümmerer, Ori Press, Matthias BethgeICCV 2021 · 98 citations
- ContourNet: Taking a Further Step Toward Accurate Arbitrary-Shaped Scene Text DetectionYuxin Wang, Hongtao Xie, Zheng-Jun Zha, Mengting Xing et al.CVPR 2020
Related papers
- Learning Selective Self-Mutual Attention for RGB-D Saliency DetectionNian Liu, Ni Zhang, Junwei HanCVPR 2020
- Multi-Type Self-Attention Guided Degraded Saliency DetectionZiqi Zhou, Zheng Wang, Huchuan Lu, Song Wang et al.AAAI 2020 · 22 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- Inferring Attention Shift Ranks of Objects for Image SaliencyAvishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie et al.CVPR 2020
- Video Object of Interest SegmentationSiyuan Zhou, Chunru Zhan, Biao Wang, Tiezheng Ge et al.AAAI 2023
