When to Say What: Learning to Find Condition-Message Inconsistencies
Islem Bouzenia, Michael Pradel
Abstract
Programs often emit natural language messages, e.g., in logging statements or exceptions raised on unexpected paths. To be meaningful to users and developers, the message, i.e., what to say, must be consistent with the condition under which it gets triggered, i.e., when to say it. However, checking for inconsistencies between conditions and messages is challenging because the conditions are expressed in the logic of the programming language, while messages are informally expressed in natural language. This paper presents CMI-Finder, an approach for detecting condition-message inconsistencies. CMI-Finder is based on a neural model that takes a condition and a message as its input and then predicts whether the two are consistent. To address the problem of obtaining realistic, diverse, and large-scale training data, we present six techniques to generate large numbers of inconsistent examples to learn from automatically. Moreover, we describe and compare three neural models, which are based on binary classification, triplet loss, and fine-tuning, respectively. Our evaluation applies the approach to 300K condition-message statements extracted from 42 million lines of Python code. The best model achieves a precision of 78% at a recall of 72% on a dataset of past bug fixes. Applying the approach to the newest versions of popular open-source projects reveals 50 previously unknown bugs, 19 of which have been confirmed by the developers so far.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- exLong: Generating Exceptional Behavior Tests with Large Language ModelsJiyang Zhang, Yu Liu, Pengyu Nie, Junyi Jessy Li et al.ICSE 2025 · 2 citations
- DyLin: A Dynamic Linter for PythonAryaz Eghbali, Felix Burk, Michael PradelFSE 2025 · 1 citation
Related papers
- Nalin: learning from Runtime Behavior to Find Name-Value Inconsistencies in Jupyter NotebooksJibesh Patra, Michael PradelICSE 2022 · 14 citations
- Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?Madeline Endres, Sarah Fakhoury, Saikat Chakraborty, Shuvendu K. LahiriFSE 2024 · 27 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Code Comment Inconsistency Detection and Rectification Using a Large Language ModelGuoping Rong, Yongda Yu, Song Liu, Xin Tan et al.ICSE 2025 · 4 citations
- DocPrism: Multi-lingual Detection of Incorrectness Inconsistencies between Code and DocumentationXiaomeng Xu, Zahin Wahab, Reid Holmes, Caroline LemieuxISSTA 2026
