CHOPIN: Scalable Graphics Rendering in Multi-GPU Systems via Parallel Image Composition
Xiaowei Ren, Mieszko Lis
Abstract
The appetite for higher and higher 3D graphics quality continues to drive GPU computing requirements. To satisfy these demands, GPU vendors are moving towards new architectures, such as MCM-GPU and multi-GPUs, that connect multiple chip modules or GPUs with high-speed links (e.g., NVLink and XGMI) to provide higher computing capability. Unfortunately, it is not clear how to adequately parallelize the rendering pipeline to take advantage of these resources while maintaining low rendering latencies. Current implementations of Split Frame Rendering (SFR) are bottlenecked by redundant computations and sequential inter-GPU synchronization, and fail to scale as the GPU count increases. In this paper, we propose CHOPIN, a novel SFR scheme for multi-GPU systems that exploits the parallelism available in image composition to eliminate the bottlenecks inherent to existing solutions. CHOPIN composes opaque sub-images out-of order, and leverages the associativity of image composition to compose adjacent sub-images of transparent objects asynchronously. To mitigate load imbalance across GPUs and avoid inter-GPU network congestion, CHOPIN includes two new scheduling mechanisms: a draw-command scheduler and an image composition scheduler. Detailed cycle-level simulations on eight real-world game traces show that, in an 8-GPU system, CHOPIN offers speedups of up to 1.56× (1.25× gmean) compared to the best prior SFR implementation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc2af8df-5ed7-46ac-b554-f53edbc4b5c4Cited by top-tier papers5
- gVulkan: Scalable GPU Pooling for Pixel-Grained Rendering in Ray TracingYicheng Gu, Yun Wang, Yunfan Sun, Yuxin Xiang et al.USENIX ATC 2024 · 4 citations
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones et al.SC 2025 · 3 citations
- Improving Resource and Energy Efficiency for Cloud 3D through Excessive Rendering ReductionTianyi Liu, Jerry Lucas, Sen He, Tongping Liu et al.EuroSys 2024 · 1 citation
- LIBRA: Memory Bandwidth- and Locality-Aware Parallel Tile RenderingAurora Tomás, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio GonzálezMICRO 2024 · 1 citation
- OS Rendering Service Made Parallel with Out-of-Order Execution and In-Order CommitYuanpei Wu, Chao Xu, Yubin Xia, Yang Yu et al.OSDI 2025 · 1 citation
Builds on2
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
- HMG: Extending Cache Coherence Protocols Across Modern Hierarchical Multi-GPU SystemsXiaowei Ren, Daniel Lustig, Evgeny Bolotin, Aamer Jaleel et al.HPCA 2020 · 38 citations
Related papers
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 20 citations
- NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU SystemsAmel Fatima, Yang Yang, Yifan Sun, Rachata Ausavarungnirun et al.ISCA 2025 · 2 citations
- XSched: Preemptive Scheduling for Diverse XPUsWeihang Shen, Mingcong Han, Jialong Liu, Rong Chen et al.OSDI 2025 · 9 citations
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He et al.VLDB 2026
- Towards Energy-Efficient Real-Time Scheduling of Heterogeneous Multi-GPU SystemsYidi Wang, Mohsen Karimi, Hyoseung KimRTSS 2022 · 8 citations
