CKTyper: Enhancing Type Inference for Java Code Snippets by Leveraging Crowdsourcing Knowledge in Stack Overflow
Anji Li, Neng Zhang, Ying Zou, Zhixiang Chen, Jian Wang, Zibin Zheng
Abstract
Code snippets are widely used in technical forums to demonstrate solutions to programming problems. They can be leveraged by developers to accelerate problem-solving. However, code snippets often lack concrete types of the APIs used in them, which impedes their understanding and resue. To enhance the description of a code snippet, a number of approaches are proposed to infer the types of APIs. Although existing approaches can achieve good performance, their performance is limited by ignoring other information outside the input code snippet (e.g., the descriptions of similar code snippets) that could potentially improve the performance.
In this paper, we propose a novel type inference approach, named CKTyper, by leveraging crowdsourcing knowledge in technical posts. The key idea is to generate a relevant context for a target code snippet from the posts containing similar code snippets and then employ the context to promote the type inference with large language models (e.g., ChatGPT). More specifically, we build a crowdsourcing knowledge base (CKB) by extracting code snippets from a large set of posts and index the CKB using Lucene. An API type dictionary is also built from a set of API libraries. Given a code snippet to be inferred, we first retrieve a list of similar code snippets from the indexed CKB. Then, we generate a crowdsourcing knowledge context (CKC) by extracting and summarizing useful content (e.g., API-related sentences) in the posts that contain the similar code snippets. The CKC is subsequently used to improve the type inference of ChatGPT on the input code snippet. The hallucination of ChatGPT is eliminated by employing the API type dictionary. Evaluation results on two open-source datasets demonstrate the effectiveness and efficiency of CKTyper. CKTyper achieves the optimal precision/recall of 97.80% and 95.54% on both datasets, respectively, significantly outperforming three state-of-the-art baselines and ChatGPT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9ba36e7-c087-4955-bc5c-b2f29f5bc71aBuilds on8
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 119 citations
- Static Inference Meets Deep learning: A Hybrid Type Inference Approach for PythonYun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao et al.ICSE 2022 · 48 citations
- Prompt-tuned Code Language Model as a Neural Knowledge Base for Type Inference in Statically-Typed Partial CodeQing Huang, Zhiqiang Yuan, Zhenchang Xing, Xiwei Xu et al.ASE 2022 · 40 citations
- CrystalBLEU: Precisely and Efficiently Measuring the Similarity of CodeAryaz Eghbali, Michael PradelASE 2022 · 33 citations
Related papers
- SnR: Constraint-Based Type Inference for Incomplete Java Code SnippetsYiwen Dong, Tianxiao Gu, Yongqiang Tian, Chengnian SunICSE 2022 · 10 citations
- Scitix: Scalable Constraint-Based Type Inference for Code Snippets with Missing TypesYiwen Dong, Zhenyang Xu, Yongqiang Tian, Edward Lee et al.ISSTA 2026
- Are Human Rules Necessary? Generating Reusable APIs with CoT Reasoning and In-Context LearningYubo Mai, Zhipeng Gao, Xing Hu, Lingfeng Bao et al.FSE 2024 · 4 citations
- Automating API Documentation from Crowdsourced KnowledgeBonan Kou, Zijie Zhou, Muhao Chen, Tianyi ZhangICSE 2026
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and BeyondWenchao Gu, Juntao Chen, Yanlin Wang, Tianyue Jiang et al.ICSE 2026 · 1 citation
