Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
Jiayi Zeng, Yizhe Feng, Mengliang He, Wenhui Lei, Wei Zhang, Zeming Liu, Xiaoming Shi, Aimin Zhou
Abstract
Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit errorhandling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. The dataset will be publicly available. Ortegal is a comarca in the province of A Coruña, in Galicia, Spain. The three municipalities in the Ortegal comarca with a total population of 12,238 in 2018 are: 1. Cerdido 2.Mañón 3. Val de San Vicente. These three make up the municipalities in Ortegal. Ortegal has three municipalities with a total population of 12,238 in 2018, one of which was called Cerdido. What are the names of the other two municipalities? GPT-4o GPT-4o Passive Error Handling Proactive Error Handling You will be given a question. First, check if there are any mistakes in the question. If the question is correct, provide the answer. If there is an error, first tell the user that there is an error in the question, identify what is wrong, and correct the question if necessary. Then, guide the user to rephrase the question. Here is the question: Ortegal has three municipalities with a total population of 12,238 in 2018, one of which was called Cerdido. What are the names of the other two municipalities? There is an error in your original question. You mentioned that Ortegal has three municipalities, but it should be that Ortegal has four municipalities. The correct question should be: Ortegal has four municipalities, with a total population of 12,238 in 2018, one of which is Cerdido. What are the names of the other three municipalities? The answer is: Besides Cerdido, the other three municipalities are: Ortigueira, Cariño, and Mañón.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9a3260a-e841-48f8-ad9d-682a98ae975aBuilds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
- Towards Conversational Recommendation over Multi-Type DialogsZeming Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu et al.ACL 2020 · 157 citations
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri et al.EMNLP 2023 · 29 citations
Related papers
- Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsQianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao et al.EMNLP 2025
- CityGPT: Empowering Urban Spatial Cognition of Large Language ModelsJie Feng, Tianhui Liu, Yuwei Du, Siqi Guo et al.KDD 2025 · 9 citations
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi et al.AAAI 2023 · 112 citations
- I am a Strange Dataset: Metalinguistic Tests for Language ModelsTristan Thrush, Jared Moore, Miguel Monares, Christopher Potts et al.ACL 2024
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 10 citations
