Saliency in Augmented Reality
Huiyu Duan, Wei Shen, Xiongkuo Min, Danyang Tu, Jing Li, Guangtao Zhai
Abstract
With the rapid development of multimedia technology, Augmented Reality (AR) has become a promising next-generation mobile platform. The primary theory underlying AR is human visual confusion, which allows users to perceive the real-world scenes and augmented contents (virtual-world scenes) simultaneously by superimposing them together. To achieve good Quality of Experience (QoE), it is important to understand the interaction between two scenarios, and harmoniously display AR contents. However, studies on how this superimposition will influence the human visual attention are lacking. Therefore, in this paper, we mainly analyze the interaction effect between background (BG) scenes and AR contents, and study the saliency prediction problem in AR. Specifically, we first construct a Saliency in AR Dataset (SARD), which contains 450 BG images, 450 AR images, as well as 1350 superimposed images generated by superimposing BG and AR images in pair with three mixing levels. A large-scale eye-tracking experiment among 60 subjects is conducted to collect eye movement data. To better predict the saliency in AR, we propose a vector quantized saliency prediction method and generalize it for AR saliency prediction. For comparison, three benchmark methods are proposed and evaluated together with our proposed method on our SARD. Experimental results demonstrate the superiority of our proposed method on both of the common saliency prediction problem and the AR saliency prediction problem over benchmark methods. Our dataset and code are available at: https://github.com/DuanHuiyu/ARSaliency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented RealityYanming Xiu, Tim Scargill, Maria GorlatovaIEEE VR 2025 · 14 citations
- ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial ImagesXilei Zhu, Liu Yang, Huiyu Duan, Xiongkuo Min et al.IEEE VR 2025 · 10 citations
- Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationZhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu et al.ACM MM 2025
Builds on4
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo et al.CVPR 2022 · 69 citations
- Perceptual Quality Assessment of Omnidirectional ImagesYuming Fang, Liping Huang, Jiebin Yan, Xuelin Liu et al.AAAI 2022 · 16 citations
- Taming Transformers for High-Resolution Image SynthesisPatrick Esser, Robin Rombach, Björn OmmerCVPR 2021
Related papers
- Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D GraphicsKaiwei Zhang, Dandan Zhu, Xiongkuo Min, Guangtao ZhaiAAAI 2025 · 1 citation
- Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependentXuemei Zhou, Irene Viola, Silvia Rossi, Pablo CésarIEEE VR 2025 · 5 citations
- FixationNet: Forecasting Eye Fixations in Task-Oriented Virtual EnvironmentsZhiming Hu, Andreas Bulling, Sheng Li, Guoping WangIEEE VR 2021 · 76 citations
- SalBiNet360: Saliency Prediction on 360° Images with Local-Global Bifurcated Deep NetworkDongwen Chen, Chunmei Qing, Xiangmin Xu, Huansheng ZhuIEEE VR 2020 · 3 citations
- SalientVR: saliency-driven mobile 360-degree video streaming with gaze informationShibo Wang, Shusen Yang, Hailiang Li, Xiaodan Zhang et al.MobiCom 2022 · 43 citations
