ROME: Testing Image Captioning Systems via Recursive Object Melting
Boxi Yu, Zhiqing Zhong, Jiaqi Li, Yixing Yang, Shilin He, Pinjia He
Abstract
Image captioning (IC) systems aim to generate a text description of the salient objects in an image. In recent years, IC systems have been increasingly integrated into our daily lives, such as assistance for visually-impaired people and description generation in Microsoft Powerpoint. However, even the cutting-edge IC systems (e.g., Microsoft Azure Cognitive Services) and algorithms (e.g., OFA) could produce erroneous captions, leading to incorrect captioning of important objects, misunderstanding, and threats to personal safety. The existing testing approaches either fail to handle the complex form of IC system output (i.e., sentences in natural language) or generate unnatural images as test cases. To address these problems, we introduce Recursive Object MElting (ROME), a novel metamorphic testing approach for validating IC systems. Different from existing approaches that generate test cases by inserting objects, which easily make the generated images unnatural, ROME melts (i.e., remove and inpaint) objects. ROME assumes that the object set in the caption of an image includes the object set in the caption of a generated image after object melting. Given an image, ROME can recursively remove its objects to generate different pairs of images. We use ROME to test one widely-adopted image captioning API and four state-of-the-art (SOTA) algorithms. The results show that the test cases generated by ROME look much more natural than the SOTA IC testing approach and they achieve comparable naturalness to the original images. Meanwhile, by generating test pairs using 226 seed images, ROME reports a total of 9,121 erroneous issues with high precision (86.47%-92.17%). In addition, we further utilize the test cases generated by ROME to retrain the Oscar, which improves its performance across multiple evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe8cc5ce-1b01-41d3-8b7e-1fda836d430bBuilds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- Towards Adversarially Robust Object DetectionHaichao Zhang, Jianyu WangICCV 2019 · 152 citations
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 114 citations
Related papers
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao et al.ISSTA 2022 · 24 citations
- Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic TestingYuanyuan Yuan, Shuai Wang, Mingyue Jiang, Tsong Yueh ChenCVPR 2021
- Validation on machine reading comprehension software without annotated labels: a property-based methodSongqiang Chen, Shuo Jin, Xiaoyuan XieFSE 2021 · 27 citations
- ASRTest: automated testing for deep-neural-network-driven speech recognition systemsPin Ji, Yang Feng, Jia Liu, Zhihong Zhao et al.ISSTA 2022 · 22 citations
- Aesthetically Relevant Image CaptioningZhipeng Zhong, Fei Zhou, Guoping QiuAAAI 2023 · 16 citations
