Atomizer: An LLM-based Collaborative Multi-Agent Framework for Intent-Driven Commit Untangling
Kangchen Zhu, Zhiliang Tian, Shangwen Wang, Mingyue Leng, Xiaoguang Mao
摘要
Composite commits, which entangle multiple unrelated concerns, are prevalent in software development and significantly hinder program comprehension and maintenance. Existing automated untangling methods, particularly state-of-the-art graph clustering-based approaches, are fundamentally limited by two issues. (1) They over-rely on structural information, failing to grasp the crucial semantic intent behind changes, and (2) they operate as “single-pass” algorithms, lacking a mechanism for the critical reflection and refinement inherent in human review processes. To overcome these challenges, we introduce Atomizer, a novel collaborative multi-agent framework for composite commit untangling. To address the semantic deficit, Atomizer employs an Intent-Oriented Chain-of-Thought (IO-CoT) strategy, which prompts large language models (LLMs) to infer the intent of each code change according to both the structure and the semantic information of code. To overcome the limitations of “single-pass” grouping, we employ two agents to establish a grouper-reviewer collaborative refinement loop, which mirrors human review practices by iteratively refining groupings until all changes in a cluster share the same underlying semantic intent. Extensive experiments on two benchmark C# and Java datasets demonstrate that Atomizer significantly outperforms several representative baselines. On average, it surpasses the state-of-the-art graph-based methods by over 6.0% on the C# dataset and 5.5% on the Java dataset. This superiority is particularly pronounced on complex commits, where Atomizer’s performance advantage widens to over 16%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang 等ICSE 2024 · 被引用 124 次
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu 等FSE 2024 · 被引用 49 次
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang 等ICSE 2023 · 被引用 49 次
- Developer-Intent Driven Code Comment GenerationFangwen Mu, Xiao Chen, Lin Shi, Song Wang 等ICSE 2023 · 被引用 25 次
相关 Paper
- UTANGO: untangling commits with context-aware, graph-based, code change clustering learning modelYi Li, Shaohua Wang, Tien N. NguyenFSE 2022 · 被引用 16 次
- LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link RecoveryArshia Akhavan, Alireza Hoseinpour, Abbas Heydarnoori, Hamid Bagheri 等FSE 2026
- SmartCommit: a graph-based interactive assistant for activity-oriented commitsBo Shen, Wei Zhang, Christian Kästner, Haiyan Zhao 等FSE 2021 · 被引用 20 次
- GPTSwarm: Language Agents as Optimizable GraphsMingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio 等ICML 2024 · 被引用 45 次
- An LLM-Based Agent-Oriented Approach for Automated Code Design Issue LocalizationFraol Batole, David O'Brien, Tien N. Nguyen, Robert Dyer 等ICSE 2025 · 被引用 7 次
