Ultrafast CPU/GPU Kernels for Density Accumulation in Placement
Zizheng Guo, Jing Mai, Yibo Lin
2021年份
11被引次数
1顶会引用
摘要
Density accumulation is a widely-used primitive operation in physical design, especially for placement. Iterative invocation in the optimization flow makes it one of the runtime bottlenecks. Accelerating density accumulation is challenging due to data dependency and workload imbalance. In this paper, we propose efficient CPU/GPU kernels for density accumulation by decomposing the problem into two phases: constant-time density collection for each instance and a linear-time prefix sum. We develop CPU and GPU dedicated implementations, and demonstrate promising efficiency benefits on tasks from large-scale placement problems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- VLSI Structure-aware Placement for Convolutional Neural Network Accelerator UnitsYun Chou, Jhih-Wei Hsu, Yao-Wen Chang, Tung-Chieh ChenDAC 2021 · 被引用 11 次
- Xplace: an extremely fast and extensible global placement frameworkLixin Liu, Bangqi Fu, Martin D. F. Wong, Evangeline F. Y. YoungDAC 2022 · 被引用 29 次
- SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUsJianqi Zhao, Yao Wen, Yuchen Luo, Zhou Jin 等DAC 2021 · 被引用 30 次
- Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingZuocheng Shi, Jie Sun, Ziyu Song, Mo Sun 等SC 2025 · 被引用 3 次
- BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU SystemsAmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh 等ISCA 2021 · 被引用 15 次
