Consistency Analysis of ChatGPT
Myeongjun Jang, Thomas Lukasiewicz
摘要
ChatGPT has gained a huge popularity since its introduction. Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields. Others, however, doubt its reliability and trustworthiness. This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency. Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions. We also ascertain via experiments that prompt designing, few-shot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu 等WWW 2024 · 被引用 126 次
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff 等NeurIPS 2025 · 被引用 51 次
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationMinzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan 等EMNLP 2023 · 被引用 33 次
- ChatGPT Incorrectness Detection in Software ReviewsMinaoar Hossain Tanzil, Junaed Younus Khan, Gias UddinICSE 2024 · 被引用 10 次
- Improving Large Language Models in Event Relation Logical PredictionMeiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao 等ACL 2024 · 被引用 7 次
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 被引用 94 次
- Maieutic Prompting: Logically Consistent Reasoning with Recursive ExplanationsJaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman 等EMNLP 2022 · 被引用 72 次
相关 Paper
- The Lawyer That Never Thinks: Consistency and Fairness as Keys to Reliable AIDana R. Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Cosmo Yang Wu 等ACL 2025
- Benchmarking and Improving Generator-Validator Consistency of Language ModelsXiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto 等ICLR 2024 · 被引用 45 次
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
- Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsYinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi 等ICML 2025
- Can Large Language Model Agents Simulate Human Trust Behavior?Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye 等NeurIPS 2024 · 被引用 183 次
