Learn How to See: Collaborative Embodied Learning for Object Detection and Camera Adjusting
Lingdong Shen, Chunlei Huo, Nuo Xu, Chaowei Han, Zichen Wang
Abstract
Passive object detectors, trained on large-scale static datasets, often overlook the feedback from object detection to image acquisition. Embodied vision and active detection mitigate this issue by interacting with the environment. Nevertheless, the materialization of activeness hinges on resource-intensive data collection and annotation. To tackle these challenges, we propose a collaborative student-teacher framework. Technically, a replay buffer is built based on the trajectory data to encapsulate the relationship of state, action, and reward. In addition, the student network diverges from reinforcement learning by redefining sequential decision pathways using a GPT structure enriched with causal self-attention. Moreover, the teacher network establishes a subtle state-reward mapping based on adjacent benefit differences, providing reliable rewards for student adaptively self-tuning with the vast unlabeled replay buffer data. Additionally, an innovative yet straightforward benefit reference value is proposed within the teacher network, adding to its effectiveness and simplicity. Leveraging a flexible replay buffer and embodied collaboration between teacher and student, the framework learns to see before detection with shallower features and shorter inference steps. Experiments highlight significant advantages of our algorithm over state-of-the-art detectors. The code is released at https://github.com/lydonShen/STF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89587536-877d-4156-9627-e9e9bc3ac778Cited by top-tier papers1
Ask how each one uses itBuilds on16
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
Related papers
- Efficient Action Recognition via Dynamic Knowledge PropagationHanul Kim, Mihir Jain, Jun-Tae Lee, Sungrack Yun et al.ICCV 2021 · 29 citations
- Cycle Self-Training for Semi-Supervised Object Detection with Distribution Consistency ReweightingHao Liu, Bin Chen, Bo Wang, Chunpeng Wu et al.ACM MM 2022 · 8 citations
- Active Teacher for Semi-Supervised Object DetectionPeng Mi, Jianghang Lin, Yiyi Zhou, Yunhang Shen et al.CVPR 2022 · 83 citations
- Masked Retraining Teacher-Student Framework for Domain Adaptive Object DetectionZijing Zhao, Sitong Wei, Qingchao Chen, Dehui Li et al.ICCV 2023 · 54 citations
- Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake DetectionZhanhe Lei, Zhongyuan Wang, Jikang Cheng, Baojin Huang et al.CVPR 2026
