Learning Active Camera for Multi-Object Navigation
Peihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu, Wenbing Huang, Thomas H. Li, Mingkui Tan, Chuang Gan
Abstract
Getting robots to navigate to multiple objects autonomously is essential yet difficult in robot applications. One of the key challenges is how to explore environments efficiently with camera sensors only. Existing navigation methods mainly focus on fixed cameras and few attempts have been made to navigate with active cameras. As a result, the agent may take a very long time to perceive the environment due to limited camera scope. In contrast, humans typically gain a larger field of view by looking around for a better perception of the environment. How to make robots perceive the environment as efficiently as humans is a fundamental problem in robotics. In this paper, we consider navigating to multiple objects more efficiently with active cameras. Specifically, we cast moving camera to a Markov Decision Process and reformulate the active camera problem as a reinforcement learning problem. However, we have to address two new challenges: 1) how to learn a good camera policy in complex environments and 2) how to coordinate it with the navigation policy. To address these, we carefully design a reward function to encourage the agent to explore more areas by moving camera actively. Moreover, we exploit human experience to infer a rule-based camera action to guide the learning process. Last, to better coordinate two kinds of policies, the camera policy takes navigation actions into account when making camera moving decisions. Experimental results show our camera policy consistently improves the performance of multi-object navigation over four baselines on two datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61d15396-3ea8-435c-a22d-3a338c06fd80Cited by top-tier papers7
- ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object NavigationKaiwen Zhou, Kaizhi Zheng, Connor Pryor, Yilin Shen et al.ICML 2023 · 221 citations
- FGPrompt: Fine-grained Goal Prompting for Image-goal NavigationXinyu Sun, Peihao Chen, Jugang Fan, Jian Chen et al.NeurIPS 2023 · 41 citations
- MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven NavigationHongcheng Wang, Peiqi Liu, Wenzhe Cai, Mingdong Wu et al.NeurIPS 2024 · 12 citations
- Continual Knowledge Adaptation for Reinforcement LearningJinwu Hu, Zihao Lian, Zhiquan Wen, Chenghao Li et al.NeurIPS 2025 · 8 citations
- 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language NavigationJianzhe Gao, Rui Liu, Wenguan WangICCV 2025 · 5 citations
Builds on21
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- MultiON: Benchmarking Semantic Map Memory using Multi-Object NavigationSaim Wani, Shivansh Patel, Unnat Jain, Angel X. Chang et al.NeurIPS 2020 · 156 citations
Related papers
- Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex EnvironmentsLiang Qin, Min Wang, Peiwei Li, Wengang Zhou et al.ICCV 2025 · 6 citations
- Active Vision for Early Recognition of Human ActionsBoyu Wang, Lihan Huang, Minh HoaiCVPR 2020
- Proactive Multi-Camera Collaboration for 3D Human Pose EstimationHai Ci, Mickel Liu, Xuehai Pan, Fangwei Zhong et al.ICLR 2023 · 6 citations
- Pose-Assisted Multi-Camera Collaboration for Active Object TrackingJing Li, Jing Xu, Fangwei Zhong, Xiangyu Kong et al.AAAI 2020 · 55 citations
- Distributional Active InferenceAbdullah Akgül, Gulcin Baykal, Manuel Haussmann, Mustafa Mert Çelikok et al.ICML 2026
