Learn How to See: Collaborative Embodied Learning for Object Detection and Camera Adjusting
Lingdong Shen, Chunlei Huo, Nuo Xu, Chaowei Han, Zichen Wang
摘要
Passive object detectors, trained on large-scale static datasets, often overlook the feedback from object detection to image acquisition. Embodied vision and active detection mitigate this issue by interacting with the environment. Nevertheless, the materialization of activeness hinges on resource-intensive data collection and annotation. To tackle these challenges, we propose a collaborative student-teacher framework. Technically, a replay buffer is built based on the trajectory data to encapsulate the relationship of state, action, and reward. In addition, the student network diverges from reinforcement learning by redefining sequential decision pathways using a GPT structure enriched with causal self-attention. Moreover, the teacher network establishes a subtle state-reward mapping based on adjacent benefit differences, providing reliable rewards for student adaptively self-tuning with the vast unlabeled replay buffer data. Additionally, an innovative yet straightforward benefit reference value is proposed within the teacher network, adding to its effectiveness and simplicity. Leveraging a flexible replay buffer and embodied collaboration between teacher and student, the framework learns to see before detection with shallower features and shorter inference steps. Experiments highlight significant advantages of our algorithm over state-of-the-art detectors. The code is released at https://github.com/lydonShen/STF.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
相关 Paper
- Efficient Action Recognition via Dynamic Knowledge PropagationHanul Kim, Mihir Jain, Jun-Tae Lee, Sungrack Yun 等ICCV 2021 · 被引用 29 次
- Cycle Self-Training for Semi-Supervised Object Detection with Distribution Consistency ReweightingHao Liu, Bin Chen, Bo Wang, Chunpeng Wu 等ACM MM 2022 · 被引用 8 次
- Active Teacher for Semi-Supervised Object DetectionPeng Mi, Jianghang Lin, Yiyi Zhou, Yunhang Shen 等CVPR 2022 · 被引用 83 次
- Masked Retraining Teacher-Student Framework for Domain Adaptive Object DetectionZijing Zhao, Sitong Wei, Qingchao Chen, Dehui Li 等ICCV 2023 · 被引用 54 次
- Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake DetectionZhanhe Lei, Zhongyuan Wang, Jikang Cheng, Baojin Huang 等CVPR 2026
