Interactron: Embodied Adaptive Object Detection
Klemen Kotar, Roozbeh Mottaghi
Abstract
Over the years various methods have been proposed for the problem of object detection. Recently, we have wit-nessed great strides in this domain owing to the emergence of powerful deep neural networks. However, there are typically two main assumptions common among these approaches. First, the model is trained on a fixed training set and is evaluated on a pre-recorded test set. Second, the model is kept frozen after the training phase, so no further updates are performed after the training is finished. These two assumptions limit the applicability of these methods to real-world settings. In this paper, we propose Interactron, a method for adaptive object detection in an interactive setting, where the goal is to perform object detection in images observed by an embodied agent navigating in different environments. Our idea is to continue training during inference and adapt the model at test time without any explicit supervision via interacting with the environment. Our adaptive object detection model provides a 11.8 point improvement in AP (and 19.1 points in <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> ) over DETR [5]. a recent, high-performance object detector. Moreover, we show that our object detection model adapts to environments with completely different appearance characteristics, and its performance is on par with a model trained with full supervision for those environments. The code is available at: https://github.com/allenai/interactron.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 909b28c0-f20e-4e86-9343-91b35a41e10eCited by top-tier papers12
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- Ask4Help: Learning to Leverage an Expert for Embodied TasksKunal Pratap Singh, Luca Weihs, Alvaro Herrasti, Jonghyun Choi et al.NeurIPS 2022 · 33 citations
- Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsLiang Qin, Min Wang, Peiwei Li, Wengang Zhou et al.ICCV 2025 · 6 citations
- Embodied Active Defense: Leveraging Recurrent Feedback to Counter Adversarial PatchesLingxuan Wu, Xiao Yang, Yinpeng Dong, Liuwei Xie et al.ICLR 2024 · 6 citations
- Learn How to See: Collaborative Embodied Learning for Object Detection and Camera AdjustingLingdong Shen, Chunlei Huo, Nuo Xu, Chaowei Han et al.AAAI 2024 · 4 citations
Builds on9
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet et al.ICLR 2022 · 435 citations
- Embodied Amodal Recognition: Learning to Move to Perceive ObjectsJianwei Yang, Zhile Ren, Mingze Xu, Xinlei Chen et al.ICCV 2019 · 70 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- Learning About Objects by Learning to Interact with ThemMartin Lohmann, Jordi Salvador, Aniruddha Kembhavi, Roozbeh MottaghiNeurIPS 2020 · 19 citations
Related papers
- Embodied-DETR: End-to-End Temporal 3D Object Detection in Egocentric ViewsZiheng Ding, Xiaze Zhang, Yuejie Zhang, lifeng chen et al.ICML 2026
- Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild With Pose AnnotationsAdel Ahmadyan, Liangkai Zhang, Artsiom Ablavatski, Jianing Wei et al.CVPR 2021
- Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic ScenariosDeng Li, Aming Wu, Yang Li, Yaowei Wang et al.ICCV 2025 · 1 citation
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingYanchao Yang, Brian Lai, Stefano SoattoCVPR 2021
- EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformerMunish Monga, Vishal Chudasama, Pankaj Wasnik, C.V. JawaharCVPR 2026 · 1 citation
