Automated testing of image captioning systems
Boxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao, Yuancheng Wang, Pinjia He
Abstract
Image captioning (IC) systems, which automatically generate a text description of the salient objects in an image (real or synthetic), have seen great progress over the past few years due to the development of deep neural networks. IC plays an indispensable role in human society, for example, labeling massive photos for scientific studies and assisting visually-impaired people in perceiving the world. However, even the top-notch IC systems, such as Microsoft Azure Cognitive Services and IBM Image Caption Generator, may return incorrect results, leading to the omission of important objects, deep misunderstanding, and threats to personal safety. To address this problem, we propose MetaIC, the first metamorphic testing approach to validate IC systems. Our core idea is that the object names should exhibit directional changes after object insertion. Specifically, MetaIC (1) extracts objects from existing images to construct an object corpus; (2) inserts an object into an image via novel object resizing and location tuning algorithms; and (3) reports image pairs whose captions do not exhibit differences in an expected way. In our evaluation, we use MetaIC to test one widely-adopted image captioning API and five state-of-theart (SOTA) image captioning models. Using 1,000 seeds, MetaIC successfully reports 16,825 erroneous issues with high precision (84.9%-98.4%). There are three kinds of errors: misclassification, omission, and incorrect quantity. We visualize the errors reported by MetaIC, which shows that flexible overlapping setting facilitates IC testing by increasing and diversifying the reported errors. In addition, MetaIC can be further generalized to detect label errors in the training dataset, which has successfully detected 151 incorrect labels in MS COCO Caption, a standard dataset in image captioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88f5c56a-5efb-48c2-b6f0-147bc0ac3419Cited by top-tier papers7
- Testing Graph Database Systems via Equivalent Query RewritingQiuyang Mang, Aoyang Fang, Boxi Yu, Hanfei Chen et al.ICSE 2024 · 12 citations
- Automated Testing and Improvement of Named Entity Recognition SystemsBoxi Yu, Yiyan Hu, Qiuyang Mang, Wenhan Hu et al.FSE 2023 · 8 citations
- FedSlice: Protecting Federated Learning Models from Malicious Participants with Model SlicingZiqi Zhang, Yuanchun Li, Bingyan Liu, Yifeng Cai et al.ICSE 2023 · 8 citations
- VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic ManipulationZhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang et al.FSE 2025 · 6 citations
- ROME: Testing Image Captioning Systems via Recursive Object MeltingBoxi Yu, Zhiqing Zhong, Jiaqi Li, Yixing Yang et al.ISSTA 2023 · 3 citations
Builds on13
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
Related papers
- ATOM: Automated Black-Box Testing of Multi-Label Image Classification SystemsShengyou Hu, Huayao Wu, Peng Wang, Jing Chang et al.ASE 2023 · 4 citations
- Metamorphic Object Insertion for Testing Object Detection SystemsShuai Wang, Zhendong SuASE 2020 · 69 citations
- Rethinking the Reference-based Distinctive Image CaptioningYangjun Mao, Long Chen, Zhihong Jiang, Dong Zhang et al.ACM MM 2022 · 22 citations
- Mitigating Gender Bias in Captioning SystemsRuixiang Tang, Mengnan Du, Yuening Li, Zirui Liu et al.WWW 2021 · 77 citations
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 8 citations
