Process Only Where You Look: Hardware and Algorithm Co-optimization for Efficient Gaze-Tracked Foveated Rendering in Virtual Reality
Haiyu Wang, Wenxuan Liu, Kenneth Chen, Qi Sun, Sai Qian Zhang
摘要
Virtual reality (VR) plays a crucial role in advancing immersive, interactive experiences that transform learning, work, and entertainment by enhancing user engagement and expanding possibilities across various fields. Image rendering is one of the most crucial application in VR, as it produces high-quality, realistic visuals that are vital for maintaining immersive user experiences and preventing visual discomfort or motion sickness. However, the cost of image rendering in VR environment is considerable, primarily due to the demands of high-quality visual experiences from users. This challenge is even greater in real-time applications, where maintaining low latency further increases the complexity of the rendering process. On the other hand, VR devices, such as head-mounted displays (HMDs), are intrinsically linked to human behavior, using insights from perception and cognition to enhance user experience.
In this work, we aim to reduce the high computational costs of the rendering process in VR by leveraging natural human eye dynamics and focusing on processing only where you look (POLO). This involves co-optimizing AI algorithms with underlying hardware for greater efficiency. We introduce POLONet, an efficient multitask deep learning framework designed to track human eye movements with minimal latency. Integrated with the POLO accelerator as a plug-in for VR HMD SoCs, this approach significantly lowers image rendering costs, achieving up to a 3.9× reduction in end-to-end latency compared to the latest gaze tracking methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SPLATONIC: Architectural Support for 3D Gaussian Splatting SLAM via Sparse ProcessingXiaotong Huang, He Zhu, Tianrui Ma, Yuxiang Xiong 等HPCA 2026 · 被引用 1 次
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 被引用 662 次
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan 等NeurIPS 2021 · 被引用 295 次
- Evo-ViT: Slow-Fast Token Evolution for Dynamic Vision TransformerYifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng 等AAAI 2022 · 被引用 288 次
- FovVideoVDP: a visible difference predictor for wide field-of-view videoRafal K. Mantiuk, Gyorgy Denes, Alexandre Chapiro, Anton Kaplanyan 等SIGGRAPH 2021 · 被引用 158 次
相关 Paper
- FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Efficient Foveated Rendering in Virtual RealityWenxuan Liu, Budmonde Duinkharjav, Qi Sun, Sai Qian ZhangIEEE VR 2025 · 被引用 14 次
- Q-VR: system-level design for future mobile collaborative virtual realityChenhao Xie, Xie Li, Yang Hu, Huwan Peng 等ASPLOS 2021 · 被引用 36 次
- Power, Performance, and Image Quality Tradeoffs in Foveated RenderingRahul Singh, Muhammad Huzaifa, Jeffrey Liu, Anjul Patney 等IEEE VR 2023 · 被引用 31 次
- Low-Latency Ocular Parallax Rendering and Investigation of Its Effect on Depth Perception in Virtual RealityYuri Mikawa, Taiki FukiageIEEE VR 2024 · 被引用 3 次
- Deep-Saliency Foveated Ray Tracing For Real-time VR RenderingYang Gao, Wencan Li, Shiyu Liang, Weizichuan Feng 等IEEE VR 2026
