MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai, Xinyue Lou, Yunghwei Lai, Ziyue Wang, Yawen Wang, Kaiyu Huang, Yile Wang, Peng Li, Yang Liu
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text pairs with clear and explicit meanings. However, resolving the inherent ambiguities present in real-world language and visual contexts remains a challenge. Existing multimodal benchmarks typically overlook linguistic and visual ambiguities, relying mainly on unimodal context for disambiguation and thus failing to exploit the mutual clarification potential between modalities. To bridge this gap, we introduce MUCAR, a novel and challenging benchmark designed explicitly for evaluating multimodal ambiguity resolution across multilingual and cross-modal scenarios. MU-CAR includes: (1) a multilingual dataset where ambiguous textual expressions are uniquely resolved by corresponding visual contexts, and (2) a dual-ambiguity dataset that systematically pairs ambiguous images with ambiguous textual contexts, with each combination carefully constructed to yield a single, clear interpretation through mutual disambiguation. Extensive evaluations involving 19 state-of-the-art multimodal models-encompassing both opensource and proprietary architectures-reveal substantial gaps compared to human-level performance, highlighting the need for future research into more sophisticated cross-modal ambiguity comprehension methods, further pushing the boundaries of multimodal reasoning. * Equal contribution. Corresponding authors. C: "I'm going to the bank." "Go fishing? Good Luck!" Q: Is there a misunderstanding between them? Answer: Yes. Explanation: The "Bank" in this sentence means a bank where money can be deposited and withdrawn. Answer: No. Explanation: The "Bank" in this sentence means the riverbank. Homonymy C: "I'm standing on the shoulders of giants now." Q: Does this sentence have a metaphor? Answer: Yes. Explanation: Each generation innovates and develops on the basis of the predecessors. Answer: No. Explanation: This is a real scene from "Gulliver's Travels". Polysemy C: The chicken is ready to eat. Q: What is the subject in the sentence going to eat? Answer: Chicken feed. Explanation: The chicken itself is hungry and ready to eat something. Answer: Chicken. Explanation: The chicken is cooked and prepared, so it is ready for someone to eat. Semantics C: 我的门没有锁。 Q: 上文的"锁"是动词还是名词?(Is "锁" above a verb or noun?) Answer: 名词。(Noun.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c40f05b4-18eb-4261-88da-4bb86ed55e5aCited by top-tier papers2
- LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language ModelsJian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao et al.ACL 2026 · 1 citation
- VFA: Empowering Multilingual MLLMs via Vision-Free AdaptationYixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang et al.ACL 2026
Builds on11
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li et al.ICML 2024 · 184 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Controllable Image Captioning via PromptingNing Wang, Jiahao Xie, Jihao Wu, Mingbo Jia et al.AAAI 2023 · 43 citations
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee et al.EMNLP 2024 · 8 citations
- Answering Ambiguous Questions via Iterative PromptingWeiwei Sun, Hengyi Cai, Hongshen Chen, Pengjie Ren et al.ACL 2023 · 3 citations
Related papers
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung et al.ICCV 2025 · 1 citation
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 7 citations
- Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading ScenariosYunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou et al.EMNLP 2025 · 1 citation
- CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language ModelsFuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang et al.ACL 2024 · 3 citations
- MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi et al.EMNLP 2024 · 7 citations
