Interactive Cross-modal Learning for Text-3D Scene Retrieval
Yanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin, Dezhong Peng, Peng Hu
Abstract
Text-3D Scene Retrieval (T3SR) aims to retrieve relevant scenes using linguistic queries. Although traditional T3SR methods have made significant progress in capturing fine-grained associations, they implicitly assume that query descriptions are information-complete. In practical deployments, however, limited by the capabilities of users and models, it is difficult or even impossible to directly obtain a perfect textual query suiting the entire scene and model, thereby leading to performance degradation. To address this issue, we propose a novel Interactive Text-3D Scene Retrieval Method (IDeal), which promotes the enhancement of the alignment between texts and 3D scenes through continuous interaction. To achieve this, we present an Interactive Retrieval Refinement framework (IRR), which employs a questioner to pose contextually relevant questions to an answerer in successive rounds that either promote detailed probing or encourage exploratory divergence within scenes. Upon the iterative responses received from the answerer, IRR adopts a retriever to perform both feature-level and semantic-level information fusion, facilitating scene-level interaction and understanding for more precise re-rankings. To bridge the domain gap between queries and interactive texts, we propose an Interaction Adaptation Tuning strategy (IAT). IAT mitigates the discriminability and diversity risks among augmented text features that approximate the interaction text domain, achieving contrastive domain adaptation for our retriever. Extensive experimental results on three datasets demonstrate the superiority of IDeal. Code is available at https://github.com/Yangl1nFeng/IDeal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 452385f8-44cd-4fed-a17e-ce74e6a599beCited by top-tier papers7
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical DecouplingXianjie Liu, Yiman Hu, Yixiong Zou, Liang Wu et al.ICML 2026 · 14 citations
- E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMsXianjie Liu, Yiman Hu, Liang Wu, Ping Hu et al.ICML 2026 · 1 citation
- EXOTIC: External Vision-driven Incomplete Multi-view ClassificationShilin Xu, Dezhong Peng, Zhenwen Ren, Yuan SunCVPR 2026
- Multimodal Nested Learning for Decoupled and Coordinated OptimizationYanglin Feng, Yang Qin, Dezhong Peng, Rui Wang et al.ICML 2026
- RLSF-V: Mitigating Hallucinations in MLLMs via Fuzzy Semantic Self-FeedbackChanghao He, ShuhaoYan, Shuxian Li, Xi Peng et al.ICML 2026
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang et al.ICML 2024 · 303 citations
Related papers
- Conversational Composed Retrieval with Iterative Sequence RefinementHao Wei, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2023 · 4 citations
- Struct-Align: Zero-Shot Text-to-3D Scene Retrieval via Locality-Aware Structural AlignmentXiong Li, Yikang Yan, Zhenyu Wen, Jie Su et al.SIGIR 2026
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo et al.CVPR 2026
- Learning to Retrieve Videos by Asking QuestionsAvinash Madasu, Junier Oliva, Gedas BertasiusACM MM 2022 · 17 citations
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu et al.ICML 2026
