RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured Environments
Haisheng Su, Feixiang Song, Cong Ma, Wei Wu, Junchi Yan
Abstract
Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes an important research topic in the areas of egocentric perceptual tasks related to navigation in both crowded and unstructured environments. Due to the complexity of environmental conditions and difficulty of surrounding obstacles owing to truncation and occlusion, the perception capability under this circumstance is still inferior. To further enhance the intelligence of mobile robots, in this paper, we setup an egocentric multisensor data collection platform based on 3 main types of sensors (Camera, LiDAR and Fisheye), which supports flexible sensor configurations to enable dynamic sight of view from ego-perspective, capturing either near or farther areas. Meanwhile, a large-scale multimodal dataset is constructed, named RoboSense, to facilitate egocentric robot perception. Specifically, RoboSense contains more than 133K synchronized data with 1.4M 3D bounding box and IDs annotated in the full 360 • view, forming 216K trajectories across 7.6K temporal sequences. It has 270× and 18× as many annotations of surrounding obstacles within near ranges as the previous datasets collected for autonomous driving scenarios such as KITTI and nuScenes. Moreover, we define a novel matching criterion for near-field 3D perception and prediction metrics. Based on RoboSense, we formulate 6 popular tasks to facilitate the future research development, where the detailed analysis as well as benchmarks are also provided accordingly. Data desensitization measures have been conducted for privacy protection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c692755-de30-4fa5-b97d-73b48eb7ef38Cited by top-tier papers4
- DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous DrivingHaisheng Su, Wei Wu, Feixiang Song, Junjie Zhang et al.ICLR 2026 · 11 citations
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based TransformerXin Jin, Haisheng Su, Cong Ma, Kai Liu et al.ICCV 2025 · 2 citations
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang et al.CVPR 2026 · 1 citation
- UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object DetectionXin Jin, Haisheng Su, Kai Liu, Cong Ma et al.CVPR 2025
Builds on12
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- Scene as OccupancyWenwen Tong, Chonghao Sima, Tai Wang, Li Chen et al.ICCV 2023 · 251 citations
- SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionPin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng et al.CVPR 2024 · 37 citations
- Forecasting from LiDAR via Future Object DetectionNeehar Peri, Jonathon Luiten, Mengtian Li, Aljosa Osep et al.CVPR 2022 · 33 citations
Related papers
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu et al.CVPR 2024 · 53 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and TrackingEdward Vendrow, Duy-Tho Le, Jianfei Cai, Hamid RezatofighiCVPR 2023
- Human-centric Scene Understanding for 3D Large-scale ScenariosYiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen et al.ICCV 2023 · 34 citations
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object UnderstandingChenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei et al.ICCV 2023 · 45 citations
