CKTyper: Enhancing Type Inference for Java Code Snippets by Leveraging Crowdsourcing Knowledge in Stack Overflow
Anji Li, Neng Zhang, Ying Zou, Zhixiang Chen, Jian Wang, Zibin Zheng
摘要
Code snippets are widely used in technical forums to demonstrate solutions to programming problems. They can be leveraged by developers to accelerate problem-solving. However, code snippets often lack concrete types of the APIs used in them, which impedes their understanding and resue. To enhance the description of a code snippet, a number of approaches are proposed to infer the types of APIs. Although existing approaches can achieve good performance, their performance is limited by ignoring other information outside the input code snippet (e.g., the descriptions of similar code snippets) that could potentially improve the performance.
In this paper, we propose a novel type inference approach, named CKTyper, by leveraging crowdsourcing knowledge in technical posts. The key idea is to generate a relevant context for a target code snippet from the posts containing similar code snippets and then employ the context to promote the type inference with large language models (e.g., ChatGPT). More specifically, we build a crowdsourcing knowledge base (CKB) by extracting code snippets from a large set of posts and index the CKB using Lucene. An API type dictionary is also built from a set of API libraries. Given a code snippet to be inferred, we first retrieve a list of similar code snippets from the indexed CKB. Then, we generate a crowdsourcing knowledge context (CKC) by extracting and summarizing useful content (e.g., API-related sentences) in the posts that contain the similar code snippets. The CKC is subsequently used to improve the type inference of ChatGPT on the input code snippet. The hallucination of ChatGPT is eliminated by employing the API type dictionary. Evaluation results on two open-source datasets demonstrate the effectiveness and efficiency of CKTyper. CKTyper achieves the optimal precision/recall of 97.80% and 95.54% on both datasets, respectively, significantly outperforming three state-of-the-art baselines and ChatGPT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 被引用 119 次
- Static Inference Meets Deep learning: A Hybrid Type Inference Approach for PythonYun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao 等ICSE 2022 · 被引用 48 次
- Prompt-tuned Code Language Model as a Neural Knowledge Base for Type Inference in Statically-Typed Partial CodeQing Huang, Zhiqiang Yuan, Zhenchang Xing, Xiwei Xu 等ASE 2022 · 被引用 40 次
- CrystalBLEU: Precisely and Efficiently Measuring the Similarity of CodeAryaz Eghbali, Michael PradelASE 2022 · 被引用 33 次
相关 Paper
- SnR: Constraint-Based Type Inference for Incomplete Java Code SnippetsYiwen Dong, Tianxiao Gu, Yongqiang Tian, Chengnian SunICSE 2022 · 被引用 10 次
- Scitix: Scalable Constraint-Based Type Inference for Code Snippets with Missing TypesYiwen Dong, Zhenyang Xu, Yongqiang Tian, Edward Lee 等ISSTA 2026
- Are Human Rules Necessary? Generating Reusable APIs with CoT Reasoning and In-Context LearningYubo Mai, Zhipeng Gao, Xing Hu, Lingfeng Bao 等FSE 2024 · 被引用 4 次
- Automating API Documentation from Crowdsourced KnowledgeBonan Kou, Zijie Zhou, Muhao Chen, Tianyi ZhangICSE 2026
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and BeyondWenchao Gu, Juntao Chen, Yanlin Wang, Tianyue Jiang 等ICSE 2026 · 被引用 1 次
