A Category Agnostic Model for Visual Rearrangment
Yuyi Liu, Xinhang Song, Weijie Li, Xiaohan Wang, Shuqiang Jiang
Abstract
This paper presents a novel category agnostic model for visual rearrangement task, which can help an embodied agent to physically recover the shuffled scene configuration without any category concepts to the goal configuration. Previous methods usually follow a similar architecture, completing the rearrangement task by aligning the scene changes of the goal and shuffled configuration, according to the semantic scene graphs. However, constructing scene graphs requires the inference of category labels, which not only causes the accuracy drop of the entire task but also limits the application in real world scenario. In this paper, we delve deep into the essence of visual rearrangement task and focus on the two most essential issues, scene change detection and scene change matching. We utilize the movement and the protrusion of point cloud to accurately identify the scene changes and match these changes depending on the similarity of category agnostic appearance feature. Moreover, to assist the agent to explore the environment more efficiently and comprehensively, we propose a closer-aligned-retrace exploration policy, aiming to observe more details of the scene at a closer distance. We conduct extensive experiments on AI2THOR Rearrangement Challenge based on RoomR dataset and a new multiroom multi-instance dataset MrMiR collected by us. The experimental results demonstrate the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc86774a-671b-4873-b524-3dfdd29f53a6Cited by top-tier papers2
- Trial-Oriented Visual RearrangementYuyi Liu, Xinhang Song, Tianliang Qi, Shuqiang JiangICCV 2025 · 1 citation
- Rethinking Visual Rearrangement from A Diffusion PerspectiveTianliang Qi, Xinhang Song, Yuyi Liu, Shuqiang JiangCVPR 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk et al.ICLR 2022 · 189 citations
- SEAL: Self-supervised Embodied Active Learning using Exploration and 3D ConsistencyDevendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta, Jitendra Malik et al.NeurIPS 2021 · 100 citations
Related papers
- A Simple Approach for Visual Room Rearrangement: 3D Mapping and Semantic SearchBrandon Trabucco, Gunnar A. Sigurdsson, Robinson Piramuthu, Gaurav S. Sukhatme et al.ICLR 2023
- SGAligner: 3D Scene Alignment with Scene GraphsSayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath et al.ICCV 2023 · 27 citations
- Hierarchical Abstraction for Combinatorial Generalization in Object RearrangementMichael Chang, Alyssa L. Dayan, Franziska Meier, Thomas L. Griffiths et al.ICLR 2023
- Continuous Scene Representations for Embodied AISamir Yitzhak Gadre, Kiana Ehsani, Shuran Song, Roozbeh MottaghiCVPR 2022 · 40 citations
- Task Planning for Object Rearrangement in Multi-Room EnvironmentsKaran Mirakhor, Sourav Ghosh, Dipanjan Das, Brojeshwar BhowmickAAAI 2024 · 2 citations
