HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
Yiming Gao, Chao Jiang, Wesley Piard, Xiangru Chen, Bhavesh Patel, Herman Lam
Abstract
Point cloud is an important type of geometric data structure for many embedded applications such as autonomous driving and augmented reality. Current Point Cloud Networks (PCNs) have proven to achieve great success in using inference to perform point cloud analysis, including object part segmentation, shape classification, and so on. However, point cloud applications on the computing edge require more than just the inference step. They require an end-to-end (E2E) processing of the point cloud workloads: pre-processing of raw data, input preparation, and inference to perform point cloud analysis. Current PCN approaches to support end-to-end processing of point cloud workload cannot meet the real-time latency requirement on the edge, i.e., the ability of the AI service to keep up with the speed of raw data generation by 3D sensors.
Latency for end-to-end processing of the point cloud workloads stems from two reasons: memory-intensive down-sampling in the pre-processing phase and the data structuring step for input preparation in the inference phase. In this paper, we present HgPCN, an end-to-end heterogeneous architecture for real-time embedded point cloud applications. In HgPCN, we introduce two novel methodologies based on spatial indexing to address the two identified bottlenecks. In the Pre-processing Engine of HgPCN, an Octree-Indexed-Sampling method is used to optimize the memory-intensive down-sampling bottleneck of the pre-processing phase. In the Inference Engine, HgPCN extends a commercial DLA with a customized Data Structuring Unit which is based on a Voxel-Expanded Gathering method to fundamentally reduce the workload of the data structuring step in the inference phase.
The initial prototype of HgPCN has been implemented on an Intel PAC (Xeon+FPGA) platform. Four commonly available point cloud datasets were used for comparison, running on three baseline devices: Intel Xeon W-2255, Nvidia Xavier NX Jetson GPU, and Nvidia 4060ti GPU. These point cloud datasets were also run on two existing PCN accelerators for comparison: PointACC and Mesorasi. Our results show that for the inference phase, depending on the dataset size, HgPCN achieves speedup from 1.3× to 10.2× vs. PointACC, 2.2× to 16.5× vs. Mesorasi, and 6.4× to 21× vs. Jetson NX GPU. Along with optimization of the memory-intensive down-sampling bottleneck in pre-processing phase, the overall latency shows that HgPCN can reach the real-time requirement by providing end-to-end service with keeping up with the raw data generation rate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44fe8617-938e-44bf-b5c5-7985e9f0099fCited by top-tier papers2
- L-PCN: A Point Cloud Accelerator Exploiting Spatial Locality through Octree-Based IslandizationYiming Gao, Jieming Yin, Yuxiang Wang, Xiangru Chen et al.ISCA 2026 · 1 citation
- FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud ProcessingYuzhe Fu, Changchun Zhou, Hancheng Ye, Bowen Duan et al.HPCA 2026 · 1 citation
Builds on6
- PointAcc: Efficient Point Cloud AcceleratorYujun Lin, Zhekai Zhang, Haotian Tang, Hanrui Wang et al.MICRO 2021 · 90 citations
- QuickNN: Memory and Performance Optimization of k-d Tree Based Nearest Neighbor Search for 3D Point CloudsReid Pinkham, Shuqing Zeng, Zhengya ZhangHPCA 2020 · 76 citations
- Mesorasi: Architecture Support for Point Cloud Analytics via Delayed-AggregationYu Feng, Boyuan Tian, Tiancheng Xu, Paul N. Whatmough et al.MICRO 2020 · 72 citations
- Crescent: taming memory irregularities for accelerating deep point cloud analyticsYu Feng, Gunnar Hammonds, Yiming Gan, Yuhao ZhuISCA 2022 · 44 citations
- Grid-GCN for Fast and Scalable Point Cloud LearningQiangeng Xu, Xudong Sun, Cho-Ying Wu, Panqu Wang et al.CVPR 2020
Related papers
- EdgePC: Efficient Deep Learning Analytics for Point Clouds on Edge DevicesZiyu Ying, Sandeepa Bhuyan, Yan Kang, Yingtian Zhang et al.ISCA 2023 · 28 citations
- Point Cloud Acceleration by Exploiting Geometric SimilarityCen Chen, Xiaofeng Zou, Hongen Shao, Yangfan Li et al.MICRO 2023 · 19 citations
- MoC: A Morton-Code-Based Fine-Grained Quantization for Accelerating Point Cloud Neural NetworksXueyuan Liu, Zhuoran Song, Hao Chen, Xing Li et al.DAC 2024 · 6 citations
- PointShuffler: Accelerating Point Cloud Neural Networks on General-Purpose GPUsYangfan Li, Zhengjie Jin, Yue Tian, Mengquan Li et al.EuroSys 2026
- An Efficient Compute-in-Memory based Accelerator for Point-based Point Cloud Neural NetworksXipeng Lin, Cong Wang, Shanshi Huang, Hongwu JiangDAC 2025
