Interactive Cross-modal Learning for Text-3D Scene Retrieval
Yanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin, Dezhong Peng, Peng Hu
摘要
Text-3D Scene Retrieval (T3SR) aims to retrieve relevant scenes using linguistic queries. Although traditional T3SR methods have made significant progress in capturing fine-grained associations, they implicitly assume that query descriptions are information-complete. In practical deployments, however, limited by the capabilities of users and models, it is difficult or even impossible to directly obtain a perfect textual query suiting the entire scene and model, thereby leading to performance degradation. To address this issue, we propose a novel Interactive Text-3D Scene Retrieval Method (IDeal), which promotes the enhancement of the alignment between texts and 3D scenes through continuous interaction. To achieve this, we present an Interactive Retrieval Refinement framework (IRR), which employs a questioner to pose contextually relevant questions to an answerer in successive rounds that either promote detailed probing or encourage exploratory divergence within scenes. Upon the iterative responses received from the answerer, IRR adopts a retriever to perform both feature-level and semantic-level information fusion, facilitating scene-level interaction and understanding for more precise re-rankings. To bridge the domain gap between queries and interactive texts, we propose an Interaction Adaptation Tuning strategy (IAT). IAT mitigates the discriminability and diversity risks among augmented text features that approximate the interaction text domain, achieving contrastive domain adaptation for our retriever. Extensive experimental results on three datasets demonstrate the superiority of IDeal. Code is available at https://github.com/Yangl1nFeng/IDeal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical DecouplingXianjie Liu, Yiman Hu, Yixiong Zou, Liang Wu 等ICML 2026 · 被引用 14 次
- E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMsXianjie Liu, Yiman Hu, Liang Wu, Ping Hu 等ICML 2026 · 被引用 1 次
- EXOTIC: External Vision-driven Incomplete Multi-view ClassificationShilin Xu, Dezhong Peng, Zhenwen Ren, Yuan SunCVPR 2026
- Multimodal Nested Learning for Decoupled and Coordinated OptimizationYanglin Feng, Yang Qin, Dezhong Peng, Rui Wang 等ICML 2026
- RLSF-V: Mitigating Hallucinations in MLLMs via Fuzzy Semantic Self-FeedbackChanghao He, ShuhaoYan, Shuxian Li, Xi Peng 等ICML 2026
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu 等ICML 2024 · 被引用 361 次
- 3D-VLA: A 3D Vision-Language-Action Generative World ModelHaoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang 等ICML 2024 · 被引用 303 次
相关 Paper
- Conversational Composed Retrieval with Iterative Sequence RefinementHao Wei, Shuhui Wang, Zhe Xue, Shengbo Chen 等ACM MM 2023 · 被引用 4 次
- Struct-Align: Zero-Shot Text-to-3D Scene Retrieval via Locality-Aware Structural AlignmentXiong Li, Yikang Yan, Zhenyu Wen, Jie Su 等SIGIR 2026
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo 等CVPR 2026
- Learning to Retrieve Videos by Asking QuestionsAvinash Madasu, Junier Oliva, Gedas BertasiusACM MM 2022 · 被引用 17 次
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu 等ICML 2026
