Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile Telepresence
Yonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih, Peizhao Zhang, Xiaoliang Dai, Yingyan Celine Lin
摘要
Real-time and robust photorealistic avatars for telepresence in AR/VR have been highly desired for enabling immersive photorealistic telepresence. However, there still exists one key bottleneck: the considerable computational expense needed to accurately infer facial expressions captured from headset-mounted cameras with a quality level that can match the realism of the avatar's human appearance. To this end, we propose a framework called Auto-CARD, which for the first time enables real-time and robust driving of Codec Avatars when exclusively using merely on-device computing resources. This is achieved by minimizing two sources of redundancy. First, we develop a dedicated neural architecture search technique called AVE-NAS for avatar encoding in AR/VR, which explicitly boosts both the searched architectures' robustness in the presence of extreme facial expressions and hardware friendliness on fast evolving AR/VR headsets. Second, we leverage the temporal redundancy in consecutively captured images during continuous rendering and develop a mechanism dubbed LATEX to skip the computation of redundant frames. Specifically, we first identify an opportunity from the linearity of the latent space derived by the avatar decoder and then propose to perform adaptive latent extrapolation for redundant frames. For evaluation, we demonstrate the efficacy of our Auto-CARD framework in real-time Codec Avatar driving settings, where we achieve a 5.05× speed-up on Meta Quest 2 while maintaining a comparable or even better animation quality than state-of-the-art avatar encoder designs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- VOODOO 3D: Volumetric Portrait Disentanglement for One-Shot 3D Head ReenactmentPhong Tran, Egor Zakharov, Long-Nhat Ho, Anh Tuan Tran 等CVPR 2024 · 被引用 15 次
- Arbitrary-Scale 3D Gaussian Super-ResolutionHuimin Zeng, Yue Bai, Yun FuAAAI 2026 · 被引用 2 次
- Physically Inspired Gaussian Splatting for HDR Novel View SynthesisHuimin Zeng, Yue Bai, hailing wang, Yun FuCVPR 2026 · 被引用 1 次
- Streamlined Facial Data Collection Based on Utterance and Emotional Data for Human-to-Avatar ReconstructionSeoyoung Kang, Seokhwan Yang, Hail Song, Boram Yoon 等IEEE VR 2026
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
它引用的顶会 Paper12
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- FasterSeg: Searching for Faster Real-time Semantic SegmentationWuyang Chen, Xinyu Gong, Xianming Liu, Qian Zhang 等ICLR 2020 · 被引用 206 次
- HW-NAS-Bench: Hardware-Aware Neural Architecture Search BenchmarkChaojian Li, Zhongzhi Yu, Yonggan Fu, Yongan Zhang 等ICLR 2021 · 被引用 128 次
相关 Paper
- Universal Facial Encoding of Codec Avatars from VR HeadsetsShaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh 等SIGGRAPH 2024 · 被引用 4 次
- F-CAD: A Framework to Explore Hardware Accelerators for Codec Avatar DecodingXiaofan Zhang, Dawei Wang, Pierce Chuang, Shugao Ma 等DAC 2021 · 被引用 9 次
- Pixel Codec AvatarsShugao Ma, Tomas Simon, Jason M. Saragih, Dawei Wang 等CVPR 2021
- LUCAS: Layered Universal Codec AvatarsDi Liu, Teng Deng, Giljoo Nam, Yu Rong 等CVPR 2025
- Driving-signal aware full-body avatarsTimur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabián Prada 等SIGGRAPH 2021 · 被引用 71 次
