Process Only Where You Look: Hardware and Algorithm Co-optimization for Efficient Gaze-Tracked Foveated Rendering in Virtual Reality
Haiyu Wang, Wenxuan Liu, Kenneth Chen, Qi Sun, Sai Qian Zhang
Abstract
Virtual reality (VR) plays a crucial role in advancing immersive, interactive experiences that transform learning, work, and entertainment by enhancing user engagement and expanding possibilities across various fields. Image rendering is one of the most crucial application in VR, as it produces high-quality, realistic visuals that are vital for maintaining immersive user experiences and preventing visual discomfort or motion sickness. However, the cost of image rendering in VR environment is considerable, primarily due to the demands of high-quality visual experiences from users. This challenge is even greater in real-time applications, where maintaining low latency further increases the complexity of the rendering process. On the other hand, VR devices, such as head-mounted displays (HMDs), are intrinsically linked to human behavior, using insights from perception and cognition to enhance user experience.
In this work, we aim to reduce the high computational costs of the rendering process in VR by leveraging natural human eye dynamics and focusing on processing only where you look (POLO). This involves co-optimizing AI algorithms with underlying hardware for greater efficiency. We introduce POLONet, an efficient multitask deep learning framework designed to track human eye movements with minimal latency. Integrated with the POLO accelerator as a plug-in for VR HMD SoCs, this approach significantly lowers image rendering costs, achieving up to a 3.9× reduction in end-to-end latency compared to the latest gaze tracking methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 497b620a-ad07-45c2-9d7c-32966a1fc835Cited by top-tier papers2
- SPLATONIC: Architectural Support for 3D Gaussian Splatting SLAM via Sparse ProcessingXiaotong Huang, He Zhu, Tianrui Ma, Yuxiang Xiong et al.HPCA 2026 · 1 citation
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan et al.NeurIPS 2021 · 295 citations
- Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision TransformerYifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng et al.AAAI 2022 · 288 citations
- FovVideoVDP: a visible difference predictor for wide field-of-view videoRafal K. Mantiuk, Gyorgy Denes, Alexandre Chapiro, Anton Kaplanyan et al.SIGGRAPH 2021 · 158 citations
Related papers
- FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Efficient Foveated Rendering in Virtual RealityWenxuan Liu, Budmonde Duinkharjav, Qi Sun, Sai Qian ZhangIEEE VR 2025 · 14 citations
- Q-VR: system-level design for future mobile collaborative virtual realityChenhao Xie, Xie Li, Yang Hu, Huwan Peng et al.ASPLOS 2021 · 36 citations
- Power, Performance, and Image Quality Tradeoffs in Foveated RenderingRahul Singh, Muhammad Huzaifa, Jeffrey Liu, Anjul Patney et al.IEEE VR 2023 · 31 citations
- Low-Latency Ocular Parallax Rendering and Investigation of Its Effect on Depth Perception in Virtual RealityYuri Mikawa, Taiki FukiageIEEE VR 2024 · 3 citations
- Deep-Saliency Foveated Ray Tracing For Real-time VR RenderingYang Gao, Wencan Li, Shiyu Liang, Weizichuan Feng et al.IEEE VR 2026
