State Space Model Meets Transformer: A New Paradigm for 3D Object Detection
Chuxin Wang, Wenfei Yang, Xiang Liu, Tianzhu Zhang
Abstract
DETR-based methods, which use multi-layer transformer decoders to refine object queries iteratively, have shown promising performance in 3D indoor object detection. However, the scene point features in the transformer decoder remain fixed, leading to minimal contributions from later decoder layers, thereby limiting performance improvement. Recently, State Space Models (SSM) have shown efficient context modeling ability with linear complexity through iterative interactions between system states and inputs. Inspired by SSMs, we propose a new 3D object DEtection paradigm with an interactive STate space model (DEST). In the interactive SSM, we design a novel state-dependent SSM parameterization method that enables system states to effectively serve as queries in 3D indoor detection tasks. In addition, we introduce four key designs tailored to the characteristics of point cloud and SSM: The serialization and bidirectional scanning strategies enable bidirectional feature interaction among scene points within the SSM. The inter-state attention mechanism models the relationships between state points, while the gated feed-forward network enhances inter-channel correlations. To the best of our knowledge, this is the first method to model queries as system states and scene points as system inputs, which can simultaneously update scene point features and query features with linear complexity. Extensive experiments on two challenging datasets demonstrate the effectiveness of our DESTbased method. Our method improves the GroupFree baseline in terms of AP 50 on ScanNet V2 (+5.3) and SUN RGB-D (+3.2) datasets. Based on the VDETR baseline, Our method sets a new SOTA on the ScanNetV2 and SUN RGB-D datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7eb4150a-7e2b-43fc-bc94-1018e59a4ce1Cited by top-tier papers8
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?Tuan Anh Tran, Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz et al.NeurIPS 2025 · 5 citations
- GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic SegmentationXujing Tao, Chuxin Wang, Yubo Ai, Zhixin Cheng et al.CVPR 2026 · 3 citations
- FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated DepthZhixin Cheng, Yujia Chen, Xujing Tao, Bohao Liao et al.ICML 2026 · 2 citations
- StruMamba3D: Exploring Structural Mamba for Self-Supervised Point Cloud Representation LearningChuxin Wang, Yixin Zha, Wenfei Yang, Tianzhu ZhangICCV 2025 · 2 citations
- Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency AlignmentZhixin Cheng, Bohao Liao, Jiacheng Deng, Xiaotian Yin et al.CVPR 2026 · 2 citations
Builds on24
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- 3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object DetectionMingsheng Li, Jiakang Yuan, Sijin Chen, Lin Zhang et al.NeurIPS 2024 · 5 citations
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object DetectionYang Cao, Feize Wu, Dave Chen, Yingji Zhong et al.CVPR 2026 · 6 citations
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich et al.AAAI 2025 · 7 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
