ChatGPT Incorrectness Detection in Software Reviews
Minaoar Hossain Tanzil, Junaed Younus Khan, Gias Uddin
Abstract
We conducted a survey of 135 software engineering (SE) practitioners to understand how they use Generative AI-based chatbots like ChatGPT for SE tasks. We find that they want to use ChatGPT for SE tasks like software library selection but often worry about the truthfulness of ChatGPT responses. We developed a suite of techniques and a tool called CID (ChatGPT Incorrectness Detector) to automatically test and detect the incorrectness in ChatGPT responses. CID is based on the iterative prompting to ChatGPT by asking it contextually similar but textually divergent questions (using an approach that utilizes metamorphic relationships in texts). The underlying principle in CID is that for a given question, a response that is different from other responses (across multiple incarnations of the question) is likely an incorrect response. In a benchmark study of library selection, we show that CID can detect incorrect responses from ChatGPT with an F1-score of 0.74 -0.75. CCS CONCEPTS • Computing methodologies → Natural language processing; • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d83a6f2-ed05-41b4-a8a2-4b62b088fe38Cited by top-tier papers5
- Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering StudiesFlorian Angermeir, Maximilian Amougou, Mark Kreitz, Andreas Bauer et al.ICSE 2026 · 1 citation
- “I need to learn better searching tactics for privacy policy laws.” Investigating Software Developers’ Behavior When Using Sources on Privacy IssuesStefan Albert Horstmann, Sandy Hong, Maziar Niazian, Cristiana Santos et al.ICSE 2026 · 1 citation
- HFuzzer: Testing Large Language Models for Package Hallucinations via Phrase-based FuzzingYukai Zhao, Menghan Wu, Xing Hu, Xin XiaASE 2025 · 1 citation
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li et al.FSE 2026
- Beyond Correctness: Exposing LLM-generated Logical Flaws in Reasoning via Multi-step Automated Theorem ProvingXinyi Zheng, Ningke Li, Xiaokun Luan, Wang Kailong et al.ICSE 2026
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
Related papers
- Chatgpt-Based Test Generation for Refactoring Engines Enhanced by Feature Analysis on ExamplesChunhao Dong, Yanjie Jiang, Yuxia Zhang, Yang Zhang et al.ICSE 2025 · 4 citations
- Chatgpt Inaccuracy Mitigation During Technical Report Understanding: Are we There Yet?Salma Begum Tamanna, Gias Uddin, Song Wang, Lan Xia et al.ICSE 2025 · 2 citations
- Demystifying and Detecting Misuses of Deep Learning APIsMoshi Wei, Nima Shiri Harzevili, Yuekai Huang, Jinqiu Yang et al.ICSE 2024 · 13 citations
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 4 citations
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
