Reflective Unit Test Generation for Precise Type Error Detection with Large Language Models
Chen Yang, Ziqi Wang, Yanjie Jiang, Lin Yang, Yuteng Zheng, Jianyi Zhou, Junjie Chen
Abstract
Type errors in Python often lead to runtime failures, posing significant challenges to software reliability and developer productivity. Existing static analysis tools aim to detect such errors without execution but frequently suffer from high false positive rates. Recently, unit test generation techniques offer great promise in achieving high test coverage, but they often struggle to produce bug-revealing tests without tailored guidance. To address these limitations, we present rTED, a novel type-aware test generation technique for automatically detecting Python type errors. Specifically, rTED combines step-by-step type constraint analysis with reflective validation to guide the test generation process and effectively suppress false positives. We evaluated rTED on two widely-used benchmarks, BugsInPy and TypeBugs. Experimental results show that rTED can detect 22 ∼ 29 more benchmarked type errors than four state-of-the-art techniques. rTED is also capable of producing fewer false positives, achieving an improvement of 173.9%∼245.9% in precision. Furthermore, we applied rTED to six real-world open-source Python projects, and successfully discovered 12 previously unknown type errors, demonstrating rTED’s practical value.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c42f5c9-32c7-4e02-8202-5cb95d45d84eCited by top-tier papers4
- Clarifying Semantics of In-Context Examples for Unit Test GenerationChen Yang, Lin Yang, Ziqi Wang, Dong Wang et al.ASE 2025 · 1 citation
- Characterizing and Mitigating False-Positive Bug Reports in the Linux KernelJiashuo Tian, Dong Wang, Chen Yang, Haichi Wang et al.FSE 2026
- Context Matters: Improving the Practical Reliability of LLM-Based Unit Test Generation (Experience Paper)Junjie Chen, Ziqi Wang, Lin Yang, Chen Yang et al.ISSTA 2026
- Uncovering Business Logic Bugs via Semantics-Driven Unit Test Generation (Experience Paper)Chen Yang, Junjie ChenISSTA 2026
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
Related papers
- Towards Effective Static Type-Error Detection for PythonWonseok Oh, Hakjoo OhASE 2024 · 1 citation
- PyTER: effective program repair for Python type errorsWonseok Oh, Hakjoo OhFSE 2022 · 22 citations
- Names Are All You Need: Effective and Safe Regression Test Selection for PythonYou Wang, Michael Pradel, Zhongxin LiuISSTA 2026
- The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed LanguagesBoqi Chen, José Antonio Hernández López, Gunter Mussbacher, Dániel VarróICSE 2025
- Domain Knowledge Matters: Improving Prompts with Fix Templates for Repairing Python Type ErrorsYun Peng, Shuzheng Gao, Cuiyun Gao, Yintong Huo et al.ICSE 2024 · 39 citations
