Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz
Abstract
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present SYMBAL, which utilizes a structured, dual-stage setup with off-the-shelf foundation models to identify systematic misalignments and summarize results in natural language. As our second key contribution, we introduce SYMBAL-BENCH, a benchmark designed to evaluate automated methods on our proposed task. SYMBAL-BENCH consists of 1.7 million image-text pairs from two domains (natural and medical images), organized into 420 vision-language datasets with annotated systematic misalignments. SYMBAL exhibits strong performance on this benchmark, correctly identifying systematic misalignments in 63.8% of datasets, a nearly 4x improvement over the closest baseline. We supplement our evaluations on SYMBALBENCH with real-world evaluations, showing that (1) SYMBAL can accurately surface systematic misalignments in captions generated by four MLLMs and (2) SYMBAL is a powerful tool for auditing off-the-shelf image-caption datasets. Ultimately, our novel task, method, and benchmark can aid users with auditing MLLMgenerated captions and identifying critical errors, without requiring access to the underlying MLLM. Code is available at https://github.com/ Stanford-AIMI/Symbal .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de61567a-0d89-4144-bb1d-8675832d692fBuilds on18
- Video Action DifferencingJames Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau et al.ICLR 2025 · 1,149 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
Related papers
- Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and CoverageSaehyung Lee, Seunghyun Yoon, Trung Bui, Jing Shi et al.ICML 2025
- CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language ModelsFuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang et al.ACL 2024 · 3 citations
- CaReBench: A Fine-grained Benchmark for Video Captioning and RetrievalYifan Xu, Xinhao Li, Yichun Yang, Desen Meng et al.ICLR 2026 · 10 citations
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsKazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki et al.EMNLP 2025
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li et al.AAAI 2026 · 2 citations
