See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
TIANCI TANG, Tielong Cai, Hongwei Wang, Gaoang Wang
Abstract
Pre-trained perception models excel in generic image domains but degrade significantly in novel environments like indoor scenes. The conventional remedy is fine-tuning on downstream data which incurs catastrophic forgetting of prior knowledge and demands costly, scene-specific annotations. We propose a paradigm shift through Sea (e, ct, dapt): rather than adapting the perception modules themselves, we adapt how they are deployed through an intelligent pose-control agent. Sea keeps all perception modules frozen, requiring no downstream labels during training, and uses only scalar perceptual feedback to navigate the agent toward informative viewpoints. Specially, we transform a vision-language model (VLM) into a low-level pose controller through a two-stage training pipeline: first fine-tuning it on rule-based exploration trajectories that systematically probe indoor scenes, and then refining the policy via unsupervised reinforcement learning that constructs rewards from the perception module’s outputs and confidence. Unlike prior active perception methods that couple exploration with specific models or collect data for retraining them, Sea directly leverages off-the-shelf perception models for various tasks without the need for retraining. We conducted experiments on three visual perception tasks, including visual grounding, segmentation and 3D box estimation, with performance improvements of 13.54%, 15.92% and 27.68% respectively on dataset ReplicaCAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c015a55b-7874-4013-a3d9-c84e41f63d41Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- Perception in ReflectionYana Wei, Liang Zhao, Kangheng Lin, En Yu et al.ICML 2025
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language ModelsJuntian Zhang, Song Jin, Chuanqi Cheng, Yuhan Liu et al.ICLR 2026 · 7 citations
- SEAL: Self-supervised Embodied Active Learning using Exploration and 3D ConsistencyDevendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta, Jitendra Malik et al.NeurIPS 2021 · 100 citations
