UTANGO: untangling commits with context-aware, graph-based, code change clustering learning model
Yi Li, Shaohua Wang, Tien N. Nguyen
摘要
During software evolution, developers make several changes and commit them into the repositories. Unfortunately, many of them tangle different purposes, both hampering program comprehension and reducing separation of concerns. Automated approaches with deterministic solutions have been proposed to untangle commits. However, specifying an effective clustering criteria on the changes in a commit for untangling is challenging for those approaches. In this work, we present UTango, a machine learning (ML)-based approach that learns to untangle the changes in a commit. We develop a novel code change clustering learning model that learns to cluster the code changes, represented by the embeddings, into different groups with different concerns. We adapt the agglomerative clustering algorithm into a supervised-learning clustering model operating on the learned code change embeddings via trainable parameters and a loss function in comparing the predicted clusters and the correct ones during training. To facilitate our clustering learning model, we develop a context-aware, graph-based, code change representation learning model, leveraging Label, Graph-based Convolution Network to produce the contextualized embeddings for code changes, that integrates program dependencies and the surrounding contexts of the changes. The contexts and cloned code are also explicitly represented, helping UTango distinguish the concerns. Our empirical evaluation on C# and Java datasets with 1,612 and 14k tangled commits show that it achieves the accuracy of 28.6%– 462.5% and 13.3%–100.0% relatively higher than the state-of-the-art commit-untangling approaches for C# and Java, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language ModelsKunpeng Zhang, Shuai Wang, Jitao Han, Xiaogang Zhu 等ICSE 2025 · 被引用 6 次
- Testora: Using Natural Language Intent to Detect Behavioral RegressionsMichael PradelICSE 2026 · 被引用 1 次
- Depradar: Agentic Coordination for Context-Aware Defect Impact Analysis in Deep Learning LibrariesYi Gao, Xing Hu, Tongtong Xu, Jiali Zhao 等ICSE 2026
- DISPATCH: Unraveling Security Patches from Entangled Code ChangesShiyu Sun, Yunlong Xing, Xinda Wang, Shu Wang 等USENIX Security 2025
- Atomizer: An LLM-based Collaborative Multi-Agent Framework for Intent-Driven Commit UntanglingKangchen Zhu, Zhiliang Tian, Shangwen Wang, Mingyue Leng 等ICSE 2026
相关 Paper
- Detect Hidden Dependency to Untangle CommitsMengdan Fan, Wei Zhang, Haiyan Zhao, Guangtai Liang 等ASE 2024 · 被引用 6 次
- Flexeme: untangling commits using lexical flowsProfir-Petru Pârtachi, Santanu Kumar Dash, Miltiadis Allamanis, Earl T. BarrFSE 2020 · 被引用 19 次
- MuDelta: Delta-Oriented Mutation Testing at Commit TimeWei Ma, Thierry Titcheu Chekam, Mike Papadakis, Mark HarmanICSE 2021 · 被引用 10 次
- COME: Commit Message Generation with Modification EmbeddingYichen He, Liran Wang, Kaiyi Wang, Yupeng Zhang 等ISSTA 2023 · 被引用 14 次
- Commit-Level, Neural Vulnerability Detection and AssessmentYi Li, Aashish Yadavally, Jiaxing Zhang, Shaohua Wang 等FSE 2023 · 被引用 12 次
