Ultrafast CPU/GPU Kernels for Density Accumulation in Placement
Zizheng Guo, Jing Mai, Yibo Lin
Abstract
Density accumulation is a widely-used primitive operation in physical design, especially for placement. Iterative invocation in the optimization flow makes it one of the runtime bottlenecks. Accelerating density accumulation is challenging due to data dependency and workload imbalance. In this paper, we propose efficient CPU/GPU kernels for density accumulation by decomposing the problem into two phases: constant-time density collection for each instance and a linear-time prefix sum. We develop CPU and GPU dedicated implementations, and demonstrate promising efficiency benefits on tasks from large-scale placement problems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- VLSI Structure-aware Placement for Convolutional Neural Network Accelerator UnitsYun Chou, Jhih-Wei Hsu, Yao-Wen Chang, Tung-Chieh ChenDAC 2021 · 11 citations
- Xplace: an extremely fast and extensible global placement frameworkLixin Liu, Bangqi Fu, Martin D. F. Wong, Evangeline F. Y. YoungDAC 2022 · 29 citations
- SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUsJianqi Zhao, Yao Wen, Yuchen Luo, Zhou Jin et al.DAC 2021 · 30 citations
- Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingZuocheng Shi, Jie Sun, Ziyu Song, Mo Sun et al.SC 2025 · 3 citations
- BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU SystemsAmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh et al.ISCA 2021 · 15 citations
