Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
Sicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun, You He, Jiankang Deng, Hang Zhang, Jifei Song, Zhensong Zhang
摘要
The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs) struggle to resolve these multimodal ambiguous inputs, often failing silently or hallucinating responses. To address these ambiguities, we introduce the Plug-and-Play Clarifier, a zero-shot and modular framework that decomposes the problem into discrete, solvable sub-tasks. Specifically, our framework consists of three synergistic modules: (1) a text clarifier that uses dialogue-driven reasoning to interactively disambiguate linguistic intent, (2) a vision clarifier that delivers real-time guidance feedback, instructing users to adjust their positioning for improved capture quality, and (3) a cross-modal clarifier with grounding mechanism that robustly interprets 3D pointing gestures and identifies the specific objects users are pointing to. Extensive experiments demonstrate that our framework improves the intent clarification performance of small language models (4-8B) by approximately 30%, making them competitive with significantly larger counterparts. We also observe consistent gains when applying our framework to these larger models. Furthermore, our vision clarifier increases corrective guidance accuracy by over 20%, and our cross-modal clarifier improves semantic answer accuracy for referential grounding by 5%. Overall, our method provides a plug-and-play framework that effectively resolves multimodal ambiguity and significantly enhances user experience in egocentric interaction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Building and Evaluating Open-Domain Dialogue Corpora with Clarifying QuestionsMohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton 等EMNLP 2021 · 被引用 61 次
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision TasksRahul Ramachandran, Ali Garjani, Roman Bachmann, Andrei Atanov 等ICLR 2026 · 被引用 21 次
- EgoPCA: A New Framework for Egocentric Hand-Object Interaction UnderstandingYue Xu, Yong-Lu Li, Zhemin Huang, Michael Xu Liu 等ICCV 2023 · 被引用 15 次
相关 Paper
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu 等ICML 2026
- Visual Intention Grounding for Egocentric AssistantsPengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 等ICCV 2025 · 被引用 2 次
- ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionYuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang 等CVPR 2025
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung 等ICCV 2025 · 被引用 1 次
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
