RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured Environments
Haisheng Su, Feixiang Song, Cong Ma, Wei Wu, Junchi Yan
摘要
Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes an important research topic in the areas of egocentric perceptual tasks related to navigation in both crowded and unstructured environments. Due to the complexity of environmental conditions and difficulty of surrounding obstacles owing to truncation and occlusion, the perception capability under this circumstance is still inferior. To further enhance the intelligence of mobile robots, in this paper, we setup an egocentric multisensor data collection platform based on 3 main types of sensors (Camera, LiDAR and Fisheye), which supports flexible sensor configurations to enable dynamic sight of view from ego-perspective, capturing either near or farther areas. Meanwhile, a large-scale multimodal dataset is constructed, named RoboSense, to facilitate egocentric robot perception. Specifically, RoboSense contains more than 133K synchronized data with 1.4M 3D bounding box and IDs annotated in the full 360 • view, forming 216K trajectories across 7.6K temporal sequences. It has 270× and 18× as many annotations of surrounding obstacles within near ranges as the previous datasets collected for autonomous driving scenarios such as KITTI and nuScenes. Moreover, we define a novel matching criterion for near-field 3D perception and prediction metrics. Based on RoboSense, we formulate 6 popular tasks to facilitate the future research development, where the detailed analysis as well as benchmarks are also provided accordingly. Data desensitization measures have been conducted for privacy protection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous DrivingHaisheng Su, Wei Wu, Feixiang Song, Junjie Zhang 等ICLR 2026 · 被引用 11 次
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based TransformerXin Jin, Haisheng Su, Cong Ma, Kai Liu 等ICCV 2025 · 被引用 2 次
- EnergyAction: Unimanual to Bimanual Composition with Energy-Based ModelsMingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang 等CVPR 2026 · 被引用 1 次
- UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object DetectionXin Jin, Haisheng Su, Kai Liu, Cong Ma 等CVPR 2025
它引用的顶会 Paper12
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
- Scene as OccupancyWenwen Tong, Chonghao Sima, Tai Wang, Li Chen 等ICCV 2023 · 被引用 251 次
- SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionPin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng 等CVPR 2024 · 被引用 37 次
- Forecasting from LiDAR via Future Object DetectionNeehar Peri, Jonathon Luiten, Mengtian Li, Aljosa Osep 等CVPR 2022 · 被引用 33 次
相关 Paper
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu 等CVPR 2024 · 被引用 53 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
- JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and TrackingEdward Vendrow, Duy-Tho Le, Jianfei Cai, Hamid RezatofighiCVPR 2023
- Human-centric Scene Understanding for 3D Large-scale ScenariosYiteng Xu, Peishan Cong, Yichen Yao, Runnan Chen 等ICCV 2023 · 被引用 34 次
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object UnderstandingChenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei 等ICCV 2023 · 被引用 45 次
