Detect Hidden Dependency to Untangle Commits
Mengdan Fan, Wei Zhang, Haiyan Zhao, Guangtai Liang, Zhi Jin
摘要
In collaborative software development, developers generally make code changes and commit the changes to the repositories. Among others, "making small, single-purpose commits" is considered the best practice for making commits, allowing the team to quickly understand the code changes. Rather than following best practices, developers often make tangled commits, which wrap code changes that implement different purposes. Such commits make it difficult for other developers to understand the code changes when conducting subsequent development. Early works on untangling code changes rely on human-specified heuristic rules or features, do not consider context, and are labor intensive. Recent works model the local context of code changes as a graph at the statement level, with statements as nodes and code dependencies as edges, and then cluster the changed statements. However, recent works ignore the hidden dependencies in the global context, e.g. a pair of tangled code changes may have no code dependency, and a pair of untangled code changes may have obvious code dependency. To solve this problem, we focus on detecting hidden dependencies among code changes. We model the global context of code changes as graphs at finer-grained, hierarchical levels, i.e., at both entity and statement levels. Then we propose a Heterogeneous Directed Graph Neural Network (HD-GNN) to detect hidden dependencies among code changes by aggregating the global context in both connected or disconnected entity-level subgraphs that intersected with the code changes. Evaluation of common C # and Java datasets with 1,612 and 14k tangled commits and manually validated datasets (MVD) with 600 commits shows that HD-GNN achieves an average enhancement of effectiveness of 25% and 19.2% compared to existing approaches and far superior to existing approaches in MVD, without sacrificing time efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- EvoClaw: Evaluating AI Agents on Continuous Software EvolutionGangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan 等ICML 2026 · 被引用 6 次
- Atomizer: An LLM-based Collaborative Multi-Agent Framework for Intent-Driven Commit UntanglingKangchen Zhu, Zhiliang Tian, Shangwen Wang, Mingyue Leng 等ICSE 2026
相关 Paper
- UTANGO: untangling commits with context-aware, graph-based, code change clustering learning modelYi Li, Shaohua Wang, Tien N. NguyenFSE 2022 · 被引用 16 次
- SmartCommit: a graph-based interactive assistant for activity-oriented commitsBo Shen, Wei Zhang, Christian Kästner, Haiyan Zhao 等FSE 2021 · 被引用 20 次
- Flexeme: untangling commits using lexical flowsProfir-Petru Pârtachi, Santanu Kumar Dash, Miltiadis Allamanis, Earl T. BarrFSE 2020 · 被引用 19 次
- GraphSPD: Graph-Based Security Patch Detection with Enriched Code SemanticsShu Wang, Xinda Wang, Kun Sun, Sushil Jajodia 等S&P 2023
- Learning semantic program embeddings with graph interval neural networkYu Wang, Ke Wang, Fengjuan Gao, Linzhang WangOOPSLA 2020 · 被引用 61 次
