HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
Yiming Gao, Chao Jiang, Wesley Piard, Xiangru Chen, Bhavesh Patel, Herman Lam
摘要
Point cloud is an important type of geometric data structure for many embedded applications such as autonomous driving and augmented reality. Current Point Cloud Networks (PCNs) have proven to achieve great success in using inference to perform point cloud analysis, including object part segmentation, shape classification, and so on. However, point cloud applications on the computing edge require more than just the inference step. They require an end-to-end (E2E) processing of the point cloud workloads: pre-processing of raw data, input preparation, and inference to perform point cloud analysis. Current PCN approaches to support end-to-end processing of point cloud workload cannot meet the real-time latency requirement on the edge, i.e., the ability of the AI service to keep up with the speed of raw data generation by 3D sensors.
Latency for end-to-end processing of the point cloud workloads stems from two reasons: memory-intensive down-sampling in the pre-processing phase and the data structuring step for input preparation in the inference phase. In this paper, we present HgPCN, an end-to-end heterogeneous architecture for real-time embedded point cloud applications. In HgPCN, we introduce two novel methodologies based on spatial indexing to address the two identified bottlenecks. In the Pre-processing Engine of HgPCN, an Octree-Indexed-Sampling method is used to optimize the memory-intensive down-sampling bottleneck of the pre-processing phase. In the Inference Engine, HgPCN extends a commercial DLA with a customized Data Structuring Unit which is based on a Voxel-Expanded Gathering method to fundamentally reduce the workload of the data structuring step in the inference phase.
The initial prototype of HgPCN has been implemented on an Intel PAC (Xeon+FPGA) platform. Four commonly available point cloud datasets were used for comparison, running on three baseline devices: Intel Xeon W-2255, Nvidia Xavier NX Jetson GPU, and Nvidia 4060ti GPU. These point cloud datasets were also run on two existing PCN accelerators for comparison: PointACC and Mesorasi. Our results show that for the inference phase, depending on the dataset size, HgPCN achieves speedup from 1.3× to 10.2× vs. PointACC, 2.2× to 16.5× vs. Mesorasi, and 6.4× to 21× vs. Jetson NX GPU. Along with optimization of the memory-intensive down-sampling bottleneck in pre-processing phase, the overall latency shows that HgPCN can reach the real-time requirement by providing end-to-end service with keeping up with the raw data generation rate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- L-PCN: A Point Cloud Accelerator Exploiting Spatial Locality through Octree-Based IslandizationYiming Gao, Jieming Yin, Yuxiang Wang, Xiangru Chen 等ISCA 2026 · 被引用 1 次
- FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud ProcessingYuzhe Fu, Changchun Zhou, Hancheng Ye, Bowen Duan 等HPCA 2026 · 被引用 1 次
它引用的顶会 Paper6
- PointAcc: Efficient Point Cloud AcceleratorYujun Lin, Zhekai Zhang, Haotian Tang, Hanrui Wang 等MICRO 2021 · 被引用 90 次
- QuickNN: Memory and Performance Optimization of k-d Tree Based Nearest Neighbor Search for 3D Point CloudsReid Pinkham, Shuqing Zeng, Zhengya ZhangHPCA 2020 · 被引用 76 次
- Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-AggregationYu Feng, Boyuan Tian, Tiancheng Xu, Paul N. Whatmough 等MICRO 2020 · 被引用 72 次
- Crescent: taming memory irregularities for accelerating deep point cloud analyticsYu Feng, Gunnar Hammonds, Yiming Gan, Yuhao ZhuISCA 2022 · 被引用 44 次
- Grid-GCN for Fast and Scalable Point Cloud LearningQiangeng Xu, Xudong Sun, Cho-Ying Wu, Panqu Wang 等CVPR 2020
相关 Paper
- EdgePC: Efficient Deep Learning Analytics for Point Clouds on Edge DevicesZiyu Ying, Sandeepa Bhuyan, Yan Kang, Yingtian Zhang 等ISCA 2023 · 被引用 28 次
- Point Cloud Acceleration by Exploiting Geometric SimilarityCen Chen, Xiaofeng Zou, Hongen Shao, Yangfan Li 等MICRO 2023 · 被引用 19 次
- MoC: A Morton-Code-Based Fine-Grained Quantization for Accelerating Point Cloud Neural NetworksXueyuan Liu, Zhuoran Song, Hao Chen, Xing Li 等DAC 2024 · 被引用 6 次
- PointShuffler: Accelerating Point Cloud Neural Networks on General-Purpose GPUsYangfan Li, Zhengjie Jin, Yue Tian, Mengquan Li 等EuroSys 2026
- An Efficient Compute-in-Memory based Accelerator for Point-based Point Cloud Neural NetworksXipeng Lin, Cong Wang, Shanshi Huang, Hongwu JiangDAC 2025
