MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai, Xinyue Lou, Yunghwei Lai, Ziyue Wang, Yawen Wang, Kaiyu Huang, Yile Wang, Peng Li, Yang Liu
摘要
Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text pairs with clear and explicit meanings. However, resolving the inherent ambiguities present in real-world language and visual contexts remains a challenge. Existing multimodal benchmarks typically overlook linguistic and visual ambiguities, relying mainly on unimodal context for disambiguation and thus failing to exploit the mutual clarification potential between modalities. To bridge this gap, we introduce MUCAR, a novel and challenging benchmark designed explicitly for evaluating multimodal ambiguity resolution across multilingual and cross-modal scenarios. MU-CAR includes: (1) a multilingual dataset where ambiguous textual expressions are uniquely resolved by corresponding visual contexts, and (2) a dual-ambiguity dataset that systematically pairs ambiguous images with ambiguous textual contexts, with each combination carefully constructed to yield a single, clear interpretation through mutual disambiguation. Extensive evaluations involving 19 state-of-the-art multimodal models-encompassing both opensource and proprietary architectures-reveal substantial gaps compared to human-level performance, highlighting the need for future research into more sophisticated cross-modal ambiguity comprehension methods, further pushing the boundaries of multimodal reasoning. * Equal contribution. Corresponding authors. C: "I'm going to the bank." "Go fishing? Good Luck!" Q: Is there a misunderstanding between them? Answer: Yes. Explanation: The "Bank" in this sentence means a bank where money can be deposited and withdrawn. Answer: No. Explanation: The "Bank" in this sentence means the riverbank. Homonymy C: "I'm standing on the shoulders of giants now." Q: Does this sentence have a metaphor? Answer: Yes. Explanation: Each generation innovates and develops on the basis of the predecessors. Answer: No. Explanation: This is a real scene from "Gulliver's Travels". Polysemy C: The chicken is ready to eat. Q: What is the subject in the sentence going to eat? Answer: Chicken feed. Explanation: The chicken itself is hungry and ready to eat something. Answer: Chicken. Explanation: The chicken is cooked and prepared, so it is ready for someone to eat. Semantics C: 我的门没有锁。 Q: 上文的"锁"是动词还是名词?(Is "锁" above a verb or noun?) Answer: 名词。(Noun.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language ModelsJian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao 等ACL 2026 · 被引用 1 次
- VFA: Empowering Multilingual MLLMs via Vision-Free AdaptationYixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang 等ACL 2026
它引用的顶会 Paper11
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 等ICML 2024 · 被引用 184 次
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- Controllable Image Captioning via PromptingNing Wang, Jiahao Xie, Jihao Wu, Mingbo Jia 等AAAI 2023 · 被引用 43 次
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee 等EMNLP 2024 · 被引用 8 次
- Answering Ambiguous Questions via Iterative PromptingWeiwei Sun, Hengyi Cai, Hongshen Chen, Pengjie Ren 等ACL 2023 · 被引用 3 次
相关 Paper
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung 等ICCV 2025 · 被引用 1 次
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 被引用 7 次
- Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading ScenariosYunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou 等EMNLP 2025 · 被引用 1 次
- CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language ModelsFuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang 等ACL 2024 · 被引用 3 次
- MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi 等EMNLP 2024 · 被引用 7 次
