Prompt-tuned Code Language Model as a Neural Knowledge Base for Type Inference in Statically-Typed Partial Code
Qing Huang, Zhiqiang Yuan, Zhenchang Xing, Xiwei Xu, Liming Zhu, Qinghua Lu
Abstract
Partial code usually involves non-fully-qualified type names (non-FQNs) and undeclared receiving objects. Resolving the FQNs of these non-FQN types and undeclared receiving objects (referred to as type inference) is the prerequisite to effective search and reuse of partial code. Existing dictionary-lookup based methods build a symbolic knowledge base of API names and code contexts, which involve significant compilation overhead and are sensitive to unseen API names and code context variations. In this paper, we formulate type inference as a cloze-style fill-in-blank language task. Built on source code naturalness, our approach fine-tunes a code masked language model (MLM) as a neural knowledge base of code elements with a novel “pre-train, prompt and predict” paradigm from raw source code. Our approach is lightweight and has minimum requirements on code compilation. Unlike existing symbolic name and context matching for type inference, our prompt-tuned code MLM packs FQN syntax and usage in its parameters and supports fuzzy neural type inference. We systematically evaluate our approach on a large amount of source code from GitHub and Stack Overflow. Our results confirm the effectiveness of our approach design and the practicality for partial code type inference. As the first of its kind, our neural type inference method opens the door to many innovative ways of using partial code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acde2cb6-8e81-44b1-abe4-0f3f55d4dddbCited by top-tier papers13
- Prompting Is All You Need: Automated Android Bug Replay with Large Language ModelsSidong Feng, Chunyang ChenICSE 2024 · 143 citations
- Let's Chat to Find the APIs: Connecting Human, LLM and Knowledge Graph through AI ChainQing Huang, Zhenyu Wan, Zhenchang Xing, Changjing Wang et al.ASE 2023 · 15 citations
- Pre-training by Predicting Program Dependencies for Vulnerability Analysis TasksZhongxin Liu, Zhijie Tang, Junwei Zhang, Xin Xia et al.ICSE 2024 · 15 citations
- Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language ModelsZejun Zhang, Zhenchang Xing, Xiaoxue Ren, Qinghua Lu et al.FSE 2024 · 11 citations
- Investigating Documented Privacy Changes in Android OSChuan Yan, Mark Huasong Meng, Fuman Xie, Guangdong BaiFSE 2024 · 6 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- Code Completion by Modeling Flattened Abstract Syntax Trees as GraphsYanlin Wang, Hui LiAAAI 2021 · 96 citations
Related papers
- Searching a Database of Source Codes Using Contextualized Code SearchRohan Mukherjee, Chris Jermaine, Swarat ChaudhuriVLDB 2020 · 11 citations
- TypeT5: Seq2seq Type Inference using Static AnalysisJiayi Wei, Greg Durrett, Isil DilligICLR 2023 · 4 citations
- Probing Linguistic Information for Logical Inference in Pre-trained Language ModelsZeming Chen, Qiyue GaoAAAI 2022 · 11 citations
- Scitix: Scalable Constraint-Based Type Inference for Code Snippets with Missing TypesYiwen Dong, Zhenyang Xu, Yongqiang Tian, Edward Lee et al.ISSTA 2026
- CKTyper: Enhancing Type Inference for Java Code Snippets by Leveraging Crowdsourcing Knowledge in Stack OverflowAnji Li, Neng Zhang, Ying Zou, Zhixiang Chen et al.FSE 2025 · 1 citation
