Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile Telepresence
Yonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih, Peizhao Zhang, Xiaoliang Dai, Yingyan Celine Lin
Abstract
Real-time and robust photorealistic avatars for telepresence in AR/VR have been highly desired for enabling immersive photorealistic telepresence. However, there still exists one key bottleneck: the considerable computational expense needed to accurately infer facial expressions captured from headset-mounted cameras with a quality level that can match the realism of the avatar's human appearance. To this end, we propose a framework called Auto-CARD, which for the first time enables real-time and robust driving of Codec Avatars when exclusively using merely on-device computing resources. This is achieved by minimizing two sources of redundancy. First, we develop a dedicated neural architecture search technique called AVE-NAS for avatar encoding in AR/VR, which explicitly boosts both the searched architectures' robustness in the presence of extreme facial expressions and hardware friendliness on fast evolving AR/VR headsets. Second, we leverage the temporal redundancy in consecutively captured images during continuous rendering and develop a mechanism dubbed LATEX to skip the computation of redundant frames. Specifically, we first identify an opportunity from the linearity of the latent space derived by the avatar decoder and then propose to perform adaptive latent extrapolation for redundant frames. For evaluation, we demonstrate the efficacy of our Auto-CARD framework in real-time Codec Avatar driving settings, where we achieve a 5.05× speed-up on Meta Quest 2 while maintaining a comparable or even better animation quality than state-of-the-art avatar encoder designs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93efdd0d-d791-48cd-a784-85231f2cd4deCited by top-tier papers5
- VOODOO 3D: Volumetric Portrait Disentanglement for One-Shot 3D Head ReenactmentPhong Tran, Egor Zakharov, Long-Nhat Ho, Anh Tuan Tran et al.CVPR 2024 · 15 citations
- Arbitrary-Scale 3D Gaussian Super-ResolutionHuimin Zeng, Yue Bai, Yun FuAAAI 2026 · 2 citations
- Physically Inspired Gaussian Splatting for HDR Novel View SynthesisHuimin Zeng, Yue Bai, hailing wang, Yun FuCVPR 2026 · 1 citation
- Streamlined Facial Data Collection Based on Utterance and Emotional Data for Human-to-Avatar ReconstructionSeoyoung Kang, Seokhwan Yang, Hail Song, Boram Yoon et al.IEEE VR 2026
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
Builds on12
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- FasterSeg: Searching for Faster Real-time Semantic SegmentationWuyang Chen, Xinyu Gong, Xianming Liu, Qian Zhang et al.ICLR 2020 · 206 citations
- HW-NAS-Bench: Hardware-Aware Neural Architecture Search BenchmarkChaojian Li, Zhongzhi Yu, Yonggan Fu, Yongan Zhang et al.ICLR 2021 · 128 citations
Related papers
- Universal Facial Encoding of Codec Avatars from VR HeadsetsShaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh et al.SIGGRAPH 2024 · 4 citations
- F-CAD: A Framework to Explore Hardware Accelerators for Codec Avatar DecodingXiaofan Zhang, Dawei Wang, Pierce Chuang, Shugao Ma et al.DAC 2021 · 9 citations
- Pixel Codec AvatarsShugao Ma, Tomas Simon, Jason M. Saragih, Dawei Wang et al.CVPR 2021
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong et al.CVPR 2025
- Driving-signal aware full-body avatarsTimur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabián Prada et al.SIGGRAPH 2021 · 71 citations
