Consistency Analysis of ChatGPT
Myeongjun Jang, Thomas Lukasiewicz
Abstract
ChatGPT has gained a huge popularity since its introduction. Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields. Others, however, doubt its reliability and trustworthiness. This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency. Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions. We also ascertain via experiments that prompt designing, few-shot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c58bb21a-26ce-45e0-bb31-60d2215fb9acCited by top-tier papers16
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu et al.WWW 2024 · 126 citations
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff et al.NeurIPS 2025 · 51 citations
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationMinzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan et al.EMNLP 2023 · 33 citations
- ChatGPT Incorrectness Detection in Software ReviewsMinaoar Hossain Tanzil, Junaed Younus Khan, Gias UddinICSE 2024 · 10 citations
- Improving Large Language Models in Event Relation Logical PredictionMeiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao et al.ACL 2024 · 7 citations
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- Maieutic Prompting: Logically Consistent Reasoning with Recursive ExplanationsJaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman et al.EMNLP 2022 · 72 citations
Related papers
- The Lawyer That Never Thinks: Consistency and Fairness as Keys to Reliable AIDana R. Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Cosmo Yang Wu et al.ACL 2025
- Benchmarking and Improving Generator-Validator Consistency of Language ModelsXiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto et al.ICLR 2024 · 45 citations
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 5 citations
- Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsYinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi et al.ICML 2025
- Can Large Language Model Agents Simulate Human Trust Behavior?Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye et al.NeurIPS 2024 · 183 citations
