BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence
Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie, Zhizhong Su
摘要
In embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments. Previous algorithms primarily rely on point cloud, which, despite offering precise geometric information, still constrain perception performance due to inherent sparsity, noise, and data scarcity. In this work, we introduce a novel image-centric 3D perception model, BIP3D, which leverages expressive image features with explicit 3D position encoding to overcome the limitations of point-centric methods. Specifically, we leverage pre-trained 2D vision foundation models to enhance semantic understanding, and introduce a spatial enhancer module to improve spatial understanding. Together, these modules enable BIP3D to achieve multi-view, multi-modal feature fusion and end-to-end 3D perception. In our experiments, BIP3D outperforms current state-of-the-art results on the EmbodiedScan benchmark, achieving improvements of 5.69% in the 3D detection task and 15.25% in the 3D visual grounding task. Code has been released at https: //github.com/HorizonRobotics/BIP3D .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu 等CVPR 2024 · 被引用 53 次
- TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIPFan Li, Zanyi Wang, Zeyi Huang, Guang Dai 等ACM MM 2025
- ODIN: A Single Model for 2D and 3D SegmentationAyush Jain, Pushkal Katara, Nikolaos Gkanatsios, Adam W. Harley 等CVPR 2024
- Situational Awareness Matters in 3D Vision Language ReasoningYunze Man, Liang-Yan Gui, Yu-Xiong WangCVPR 2024 · 被引用 9 次
- Learning Multi-View Aggregation In the Wild for Large-Scale 3D Semantic SegmentationDamien Robert, Bruno Vallet, Loïc LandrieuCVPR 2022 · 被引用 84 次
