Streamlining Java Programming: Uncovering Well-Formed Idioms with IdioMine
Yanming Yang, Xing Hu, Xin Xia, David Lo, Xiaohu Yang
Abstract
Code idioms are commonly used patterns, techniques, or practices that aid in solving particular problems or specific tasks across multiple software projects. They can improve code quality, performance, and maintainability, and also promote program standardization and reuse across projects. However, identifying code idioms is significantly challenging, as existing studies have still suffered from three main limitations. First, it is difficult to recognize idioms that span non-contiguous code lines. Second, identifying idioms with intricate data flow and code structures can be challenging. Moreover, they only extract dataset-specific idioms, so common idioms or well-established code/design patterns that are rarely found in datasets cannot be identified.
To overcome these limitations, we propose a novel approach, named IdioMine, to automatically extract generic and specific idioms from both Java projects and libraries. We perform program analysis on Java functions to transform them into concise PDGs, for integrating the data flow and control flow of code fragments. We then develop a novel chain structure, Data-driven Control Chain (DCC), to extract sub-idioms that possess contiguous semantic meanings from PDGs. After that, we utilize GraphCodeBERT to generate code embeddings of these sub-idioms and perform densitybased clustering to obtain frequent sub-idioms. We use heuristic rules to identify interrelated sub-idioms among the frequent ones. Finally, we employ ChatGPT to synthesize interrelated sub-idioms into potential code idioms and infer real idioms from them.
We conduct well-designed experiments and a user study to evaluate IdioMine's correctness and the practical value of the extracted idioms. Our experimental results show that IdioMine effectively extracts more idioms with better performance in most metrics. We
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c48e1db-876f-47eb-bfc8-87b80dc529ffCited by top-tier papers2
- Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient ShorthandZhensu Sun, Chengran Yang, Xiaoning Du, Zhou Yang et al.ASE 2025 · 2 citations
- Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code GenerationYu Yu, Zhihong Sun, Jia Li, Yao Wan et al.FSE 2026
Builds on1
Related papers
- Learning to find naming issues with big code and small supervisionJingxuan He, Cheng-Chun Lee, Veselin Raychev, Martin T. VechevPLDI 2021 · 9 citations
- Hard to Read and Understand Pythonic Idioms? DeIdiom and Explain Them in Non-Idiomatic Equivalent CodeZejun Zhang, Zhenchang Xing, Dehai Zhao, Qinghua Lu et al.ICSE 2024 · 7 citations
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 7 citations
- Making Python code idiomatic by automatic refactoring non-idiomatic Python code with pythonic idiomsZejun Zhang, Zhenchang Xing, Xin Xia, Xiwei Xu et al.FSE 2022 · 35 citations
- From Misuse to Mastery: Enhancing Code Generation with Knowledge-Driven AI ChainingXiaoxue Ren, Xinyuan Ye, Dehai Zhao, Zhenchang Xing et al.ASE 2023 · 28 citations
