Embodied-DETR: End-to-End Temporal 3D Object Detection in Egocentric Views
Ziheng Ding, Xiaze Zhang, Yuejie Zhang, lifeng chen, Rui Feng
摘要
Embodied 3D object detection is a fundamental perceptual capability for embodied agents, in which observations are partial, heavily occluded, and sequential, requiring modeling of temporal continuity. However, existing benchmarks and methods are primarily designed for fully reconstructed global scenes and fail to capture temporal observation context and instance evolution in first-person perception. We introduce Embodied-Det , a new benchmark for embodied 3D object detection that evaluates detection accuracy, temporal stability, and consistency under egocentric sequential views. Building on this benchmark, we propose Embodied-DETR , an end-to-end temporal detection framework that models scene-level context and instance-level consistency through two complementary temporal modules, Scene-aware Feature Aggregation and Instance-aware Query Embedding . Experiments on Embodied-Det show that existing methods suffer substantial performance degradation in egocentric temporal settings, while Embodied-DETR achieves superior accuracy and temporal consistency, demonstrating the effectiveness of temporal modeling for embodied 3D perception. Codes are available at https://github.com/UniPerceptor/UniPerceptor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 被引用 602 次
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li 等NeurIPS 2022 · 被引用 401 次
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu 等ICCV 2021 · 被引用 368 次
- RIO: 3D Object Instance Re-Localization in Changing Indoor EnvironmentsJohanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari 等ICCV 2019 · 被引用 233 次
相关 Paper
- DetAny4D: Detect Anything 4D Temporally in a Streaming RGB VideoJiawei Hou, Shenghao Zhang, Can Wang, Zheng Gu 等CVPR 2026
- Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint FramesSahithya Ravi, Gabriel Herbert Sarch, Vibhav Vineet, Andrew D. Wilson 等EMNLP 2025 · 被引用 1 次
- EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-Based Online Scene UnderstandingYuqi Wu, Wenzhao Zheng, Sicheng Zuo, Yuanhui Huang 等ICCV 2025 · 被引用 4 次
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AITai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu 等CVPR 2024 · 被引用 53 次
- Online Segment Any 3D Thing as Instance TrackingHanshi Wang, Zijian Cai, Jin Gao, Yiwei Zhang 等NeurIPS 2025 · 被引用 6 次
