FSHNet: Fully Sparse Hybrid Network for 3D Object Detection
Shuai Liu, Mingyue Cui, Boyang Li, Quanmin Liang, Tinghe Hong, Yunxiao Shan, Kai Huang
Abstract
Fully sparse 3D detectors have recently gained significant attention due to their efficiency in long-range detection. However, sparse 3D detectors extract features only from non-empty voxels, which impairs long-range interactions and causes the center feature missing. The former weakens the feature extraction capability, while the latter hinders network optimization. To address these challenges, we introduce the Fully Sparse Hybrid Network (FSHNet). FSHNet incorporates a proposed SlotFormer block to enhance the long-range feature extraction capability of existing sparse encoders. The SlotFormer divides sparse voxels using a slot partition approach, which, compared to traditional window partition, provides a larger receptive field. Additionally, we propose a dynamic sparse label assignment strategy to deeply optimize the network by providing more high-quality positive samples. To further enhance performance, we introduce a sparse upsampling module to refine downsampled voxels, preserving fine-grained details crucial for detecting small objects. Extensive experiments on the Waymo, nuScenes, and Argoverse2 benchmarks demonstrate the effectiveness of FSHNet. The code is available at https://github.com/Say2L/FSHNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac35ad44-c5d8-4a91-a347-15bd62c2b909Cited by top-tier papers6
- Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent DiffusionWentao Qu, Guofeng Mei, Jing Wang, Yujiao Wu et al.AAAI 2026 · 6 citations
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene GenerationWentao Qu, Guofeng Mei, Yang Wu, Yongshun Gong et al.CVPR 2026 · 4 citations
- Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D DetectionXiang Li, Zhangchi Hu, Xu Xiao, Bin KongCVPR 2026 · 3 citations
- 3D MeanFlow: One-Step Point Cloud Completion and Generation via Average-Velocity TransportHaowen Zhong, Jiujun Cheng, Haowen Wang, Chao Wei et al.ICML 2026
- Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object DetectionZhiwei Ning, Xuanang Gao, Jiaxi Cao, Runze Yang et al.ICLR 2026
Builds on30
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
Related papers
- SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object DetectionGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li et al.CVPR 2024 · 56 citations
- HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point CloudsGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li et al.NeurIPS 2023 · 95 citations
- Fully Sparse 3D Object DetectionLue Fan, Feng Wang, Naiyan Wang, Zhaoxiang ZhangNeurIPS 2022 · 168 citations
- SparseFormer: Detecting Objects in HRW Shots via Sparse Vision TransformerWenxi Li, Yuchen Guo, Jilai Zheng, Haozhe Lin et al.ACM MM 2024 · 3 citations
- Focal Sparse Convolutional Networks for 3D Object DetectionYukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun et al.CVPR 2022 · 293 citations
