F-CAD: A Framework to Explore Hardware Accelerators for Codec Avatar Decoding
Xiaofan Zhang, Dawei Wang, Pierce Chuang, Shugao Ma, Deming Chen, Yuecheng Li
摘要
Creating virtual avatars with realistic rendering is one of the most essential and challenging tasks to provide highly immersive virtual reality (VR) experiences. It requires not only sophisticated deep neural network (DNN) based codec avatar decoders to ensure high visual quality and precise motion expression, but also efficient hardware accelerators to guarantee smooth real-time rendering using lightweight edge devices, like untethered VR headsets. Existing hardware accelerators, however, fail to deliver sufficient performance and efficiency targeting such decoders which consist of multi-branch DNNs and require demanding compute and memory resources. To address these problems, we propose an automation framework, called F-CAD (Facebook Codec avatar Accelerator Design), to explore and deliver optimized hardware accelerators for codec avatar decoding. Novel technologies include 1) a new accelerator architecture to efficiently handle multi-branch DNNs; 2) a multi-branch dynamic design space to enable fine-grained architecture configurations; and 3) an efficient architecture search for picking the optimized hardware design based on both application-specific demands and hardware resource constraints. To the best of our knowledge, F-CAD is the first automation tool that supports the whole design flow of hardware acceleration of codec avatar decoders, allowing joint optimization on decoder designs in popular machine learning frameworks and corresponding customized accelerator design with cycle-accurate evaluation. Results show that the accelerators generated by F-CAD can deliver up to 122.1 frames per second (FPS) and 91.6% hardware efficiency when running the latest codec avatar decoder. Compared to the state-of-the-art designs, F-CAD achieves 4.0× and 2.8× higher throughput, 62.5% and 21.2% higher efficiency than DNNBuilder [1] and HybridDNN [2] by targeting the same hardware device.
• This is the first work that focuses on building electronic design automation tools and providing rapid hardware accelerator design and exploration to leverage VR avatar applications for resourceconstrained devices.
• We propose a novel elastic architecture to flexibly support multibranch DNNs with complicated layer dependencies and a wellconstructed architecture unit to support up to three-dimensional parallelism (3D parallelism) for high throughput and efficiency.
• We define a multi-branch dynamic design space to cover all possible hardware design combinations, which allows F-CAD to obtain the optimized solution with the best achievable performance.
• We integrate a DSE (design space exploration) engine to leverage efficient explorations within the predefined space and deliver the best accelerator by considering various customized constraints, such as available resources, maximum parallelism, maximum batch size, different branch priority, etc.
Codec avatar is formulated as a view-dependent Variational Auto-Encoder (VAE) framework [3], [4]. As described in Fig. 2, the
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
- Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile TelepresenceYonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih 等CVPR 2023
- Best of Both Worlds: AutoML Codesign of a CNN and its Hardware AcceleratorMohamed S. Abdelfattah, Lukasz Dudziak, Thomas Chau, Royson Lee 等DAC 2020 · 被引用 78 次
- Uni-Render: A Unified Accelerator for Real-Time Rendering Across Diverse Neural RenderersChaojian Li, Sixu Li, Linrui Jiang, Jingqun Zhang 等HPCA 2025 · 被引用 6 次
- DANCE: Differentiable Accelerator/Network Co-ExplorationKanghyun Choi, Deokki Hong, Hojae Yoon, Joonsang Yu 等DAC 2021 · 被引用 49 次
