Generative Type Inference for Python
Yun Peng, Chaozheng Wang, Wenxuan Wang, Cuiyun Gao, Michael R. Lyu
摘要
Python is a popular dynamic programming language, evidenced by its ranking as the second most commonly used language on GitHub. However, its dynamic type system can lead to potential type errors, leading researchers to explore automatic type inference approaches for Python programs. Existing type inference approaches can be generally grouped into three categories, i.e., rule-based, supervised, and cloze- style approaches. The rule-based type inference approaches can ensure the accuracy of predicted variable types, but they suffer from low coverage problems caused by dynamic features and external calls. Supervised type inference approaches, while feature-agnostic and able to mitigate the low coverage problem, require large, high- quality annotated datasets and are limited to pre-defined types. As zero-shot approaches, the cloze-style approaches reformulate the type inference problem into a fill-in-the-blank problem by leveraging the general knowledge in powerful pre-trained code models. However, their performance is limited since they ignore the domain knowledge from static typing rules which reflect the inference logic. What is more, their predictions are not interpretable, hindering developers' understanding and verification of the results. This paper introduces Typegen, a few-shot generative type inference approach that incorporates static domain knowledge from static analysis. Typegen creates chain-of-thought (COT) prompts by translating the type inference steps of static analysis into prompts based on the type dependency graphs (TDGs), enabling language models to learn from how static analysis infers types. By combining COT prompts with code slices and type hints, TypegEnconstructs example prompts from human annotations. Typeg Enonly requires very few annotated examples to teach language models to generate similar COT prompts via in-context learning. Moreover, Typeg Enenhances the interpretability of results through the use of the input- explanation-output strategy, which generates both explanations and type predictions in COT prompts. Experiments show that Typegen outperforms the best baseline Type4Py by 10.0% for argument type prediction and 22.5 % in return value type prediction in terms of top-l Exact Match by using only five examples. Furthermore, Typeg Enachieves substantial improvements of 27 % to 84 % compared to the zero-shot performance of large language models with parameter sizes ranging from 1.3B to 175B in terms of top-I Exact Match.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- LILAC: Log Parsing using LLMs with Adaptive Parsing CacheZhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li 等FSE 2024 · 被引用 85 次
- Face It Yourselves: An LLM-Based Two-Stage Strategy to Localize Configuration Errors via LogsShiwen Shan, Yintong Huo, Yuxin Su, Yichen Li 等ISSTA 2024 · 被引用 18 次
- Go Static: Contextualized Logging Statement GenerationYichen Li, Yintong Huo, Renyi Zhong, Zhihan Jiang 等FSE 2024 · 被引用 17 次
- SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionXin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao 等ISSTA 2024 · 被引用 17 次
- Search-Based LLMs for Code OptimizationShuzheng Gao, Cuiyun Gao, Wenchao Gu, Michael R. LyuICSE 2025 · 被引用 11 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
相关 Paper
- Static Inference Meets Deep learning: A Hybrid Type Inference Approach for PythonYun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao 等ICSE 2022 · 被引用 48 次
- TypeT5: Seq2seq Type Inference using Static AnalysisJiayi Wei, Greg Durrett, Isil DilligICLR 2023 · 被引用 4 次
- Concrete Type Inference for Code Optimization using Machine Learning with SMT SolvingFangke Ye, Jisheng Zhao, Jun Shirako, Vivek SarkarOOPSLA 2023 · 被引用 5 次
- TypePro: Boosting LLM-Based Type Inference via Inter-Procedural SlicingTeyu Lin, Minghao Fan, Huaxun Huang, Zhirong Shen 等FSE 2026
- TIGER: A Generating-Then-Ranking Framework for Practical Python Type InferenceChong Wang, Jian Zhang, Yiling Lou, Mingwei Liu 等ICSE 2025 · 被引用 1 次
