Prompt-tuned Code Language Model as a Neural Knowledge Base for Type Inference in Statically-Typed Partial Code
Qing Huang, Zhiqiang Yuan, Zhenchang Xing, Xiwei Xu, Liming Zhu, Qinghua Lu
摘要
Partial code usually involves non-fully-qualified type names (non-FQNs) and undeclared receiving objects. Resolving the FQNs of these non-FQN types and undeclared receiving objects (referred to as type inference) is the prerequisite to effective search and reuse of partial code. Existing dictionary-lookup based methods build a symbolic knowledge base of API names and code contexts, which involve significant compilation overhead and are sensitive to unseen API names and code context variations. In this paper, we formulate type inference as a cloze-style fill-in-blank language task. Built on source code naturalness, our approach fine-tunes a code masked language model (MLM) as a neural knowledge base of code elements with a novel “pre-train, prompt and predict” paradigm from raw source code. Our approach is lightweight and has minimum requirements on code compilation. Unlike existing symbolic name and context matching for type inference, our prompt-tuned code MLM packs FQN syntax and usage in its parameters and supports fuzzy neural type inference. We systematically evaluate our approach on a large amount of source code from GitHub and Stack Overflow. Our results confirm the effectiveness of our approach design and the practicality for partial code type inference. As the first of its kind, our neural type inference method opens the door to many innovative ways of using partial code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 被引用 143 次
- Let's Chat to Find the APIs: Connecting Human, LLM and Knowledge Graph through AI ChainQing Huang, Zhenyu Wan, Zhenchang Xing, Changjing Wang 等ASE 2023 · 被引用 15 次
- Pre-training by Predicting Program Dependencies for Vulnerability Analysis TasksZhongxin Liu, Zhijie Tang, Junwei Zhang, Xin Xia 等ICSE 2024 · 被引用 15 次
- Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language ModelsZejun Zhang, Zhenchang Xing, Xiaoxue Ren, Qinghua Lu 等FSE 2024 · 被引用 11 次
- Investigating Documented Privacy Changes in Android OSChuan Yan, Mark Huasong Meng, Fuman Xie, Guangdong BaiFSE 2024 · 被引用 6 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- Code Completion by Modeling Flattened Abstract Syntax Trees as GraphsYanlin Wang, Hui LiAAAI 2021 · 被引用 96 次
相关 Paper
- Searching a Database of Source Codes Using Contextualized Code SearchRohan Mukherjee, Chris Jermaine, Swarat ChaudhuriVLDB 2020 · 被引用 11 次
- TypeT5: Seq2seq Type Inference using Static AnalysisJiayi Wei, Greg Durrett, Isil DilligICLR 2023 · 被引用 4 次
- Probing Linguistic Information for Logical Inference in Pre-trained Language ModelsZeming Chen, Qiyue GaoAAAI 2022 · 被引用 11 次
- Scitix: Scalable Constraint-Based Type Inference for Code Snippets with Missing TypesYiwen Dong, Zhenyang Xu, Yongqiang Tian, Edward Lee 等ISSTA 2026
- CKTyper: Enhancing Type Inference for Java Code Snippets by Leveraging Crowdsourcing Knowledge in Stack OverflowAnji Li, Neng Zhang, Ying Zou, Zhixiang Chen 等FSE 2025 · 被引用 1 次
