HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets
Yili Jin, Xize Duan, Fangxin Wang, Xue Liu
Abstract
Virtual Reality (VR) has become increasingly popular for remote collaboration, but video conferencing poses challenges when the user's face is covered by the headset. Existing solutions have limitations in terms of accessibility. In this paper, we propose HeadsetOff, a novel system that achieves photorealistic video conferencing on economical VR headsets by leveraging voice-driven face reconstruction. HeadsetOff consists of three main components: a multimodal predictor, a generator, and an adaptive controller. The predictor effectively predicts user future behavior based on different modalities. The generator employs voice, head motion, and eye blink to animate the human face. The adaptive controller dynamically selects the appropriate generator model based on the trade-off between video quality and delay. Experimental results demonstrate the effectiveness of HeadsetOff in achieving high-quality, low-latency video conferencing on economical VR headsets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8da7289f-eaa4-49b1-9bb9-7a222a8f4516Cited by top-tier papers1
Ask how each one uses itBuilds on12
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- CaV3: Cache-assisted Viewport Adaptive Volumetric Video StreamingJunhua Liu, Boxiang Zhu, Fangxin Wang, Yili Jin et al.IEEE VR 2023 · 40 citations
- Understanding User Behavior in Volumetric Video Watching: Dataset, Analysis and PredictionKaiyuan Hu, Haowen Yang, Yili Jin, Junhua Liu et al.ACM MM 2023 · 37 citations
- Where Are You Looking?: A Large-Scale Dataset of Head and Gaze Behavior for 360-Degree Videos and a Pilot StudyYili Jin, Junhua Liu, Fangxin Wang, Shuguang CuiACM MM 2022 · 37 citations
Related papers
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- Universal Facial Encoding of Codec Avatars from VR HeadsetsShaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh et al.SIGGRAPH 2024 · 4 citations
- Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile TelepresenceYonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih et al.CVPR 2023
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang et al.CVPR 2023
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
