Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Qing Yu, Go Irie, Yixuan Li, Hai Helen Li, Ziwei Liu, Kiyoharu Aizawa
摘要
This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed Unsolvable Problem Detection (UPD). Multiplechoice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM's ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. What color is Standard Q: What kind of weather is depicted in the picture? (a) Base (b) Option (c) Instruction
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Large Language Models Struggle with Unreasonability in Math ProblemsJingyuan Ma, Damai Dai, Zihang Yuan, Rui Li 等AAAI 2026 · 被引用 10 次
- Outlier Synthesis via Hamiltonian Monte Carlo for Out-of-Distribution DetectionHengzhuang Li, Teng ZhangICLR 2025
- MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingFei Wang, Xingyu Fu, James Y. Huang, Zekun Li 等ICLR 2025
它引用的顶会 Paper7
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- VOS: Learning What You Don't Know by Virtual Outlier SynthesisXuefeng Du, Zhaoning Wang, Mu Cai, Yixuan LiICLR 2022 · 被引用 417 次
- Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIPSepideh Esmaeilpour, Bing Liu, Eric Robertson, Lei ShuAAAI 2022 · 被引用 219 次
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 被引用 78 次
相关 Paper
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li 等CVPR 2024
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li 等AAAI 2026 · 被引用 2 次
- MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual QuestionsYanxu Zhu, Shitong Duan, Xiangxu Zhang, Jitao Sang 等AAAI 2026 · 被引用 2 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video UnderstandingXueqing Yu, Bohan Li, Yan Li, Zhenheng YangCVPR 2026 · 被引用 2 次
