Warehouse-scale video acceleration: co-design and deployment in the wild
Parthasarathy Ranganathan, Daniel Stodolsky, Jeff Calow, Jeremy Dorfman, Marisabel Guevara, Clinton Wills Smullen IV, Aki Kuusela, Raghu Balasubramanian, Sandeep Bhatia, Prakash Chauhan, Anna Cheung, In Suk Chong
Abstract
Video sharing (e.g., YouTube, Vimeo, Facebook, TikTok) accounts for the majority of internet traffic, and video processing is also foundational to several other key workloads (video conferencing, virtual/augmented reality, cloud gaming, video in Internet-of-Things devices, etc.). The importance of these workloads motivates larger video processing infrastructures and -with the slowing of Moore's law -specialized hardware accelerators to deliver more computing at higher efficiencies. This paper describes the design and deployment, at scale, of a new accelerator targeted at warehouse-scale video transcoding. We present our hardware design including a new accelerator building block -the video coding unit (VCU) -and discuss key design trade-offs for balanced systems at data center scale and co-designing accelerators with large-scale distributed software systems. We evaluate these accelerators "in the wild" serving live data center jobs, demonstrating 20-33x improved efficiency over our prior well-tuned non-accelerated baseline. Our design also enables effective adaptation to changing bottlenecks and improved failure Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83346fe8-cc28-4295-8cc8-6eed6ef3efd2Cited by top-tier papers15
- BTS: an accelerator for bootstrappable fully homomorphic encryptionSangpyo Kim, Jongmin Kim, Michael Jaemin Kim, Wonkyung Jung et al.ISCA 2022 · 184 citations
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu et al.ASPLOS 2022 · 48 citations
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao et al.MICRO 2021 · 42 citations
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- Profiling Hyperscale Big Data ProcessingAbraham Gonzalez, Aasheesh Kolli, Samira Manabi Khan, Sihang Liu et al.ISCA 2023 · 30 citations
Builds on2
- Interstellar: Using Halide's Scheduling Language to Analyze DNN AcceleratorsXuan Yang, Mingyu Gao, Qiaoyi Liu, Jeff Setter et al.ASPLOS 2020 · 237 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
Related papers
- PacketGame: Multi-Stream Packet Gating for Concurrent Video Inference at ScaleMu Yuan, Lan Zhang, Xuanke You, Xiang-Yang LiSIGCOMM 2023 · 18 citations
- ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling InsightsTao Lu, Jiapin Wang, Yelin Shan, Xiangping Zhang et al.EuroSys 2026
- F-CAD: A Framework to Explore Hardware Accelerators for Codec Avatar DecodingXiaofan Zhang, Dawei Wang, Pierce Chuang, Shugao Ma et al.DAC 2021 · 9 citations
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual RealityMingzhi Zhu, Ding Shang, Sai Qian ZhangNeurIPS 2025
- eXpressSFU: Toward Super-Scalable Video Conferencing with SmartNICsTuan Tran, S. M. H. Hosseini, Seyeon Kim, Kyunghan Lee et al.NSDI 2026
