Semantic-Enhanced Automatic Refinement of Architecture Recovery Results Using LLMs
Yiran Zhang, Chengwei Liu, Yuqiang Sun, Zhengzi Xu, Weisong Sun, Wenke Li, Wuxia Jin, Yang Liu
摘要
Understanding the architecture is crucial for effectively maintaining and managing large software systems. However, discrepancies often exist between the designed and implemented architectures, which can pose significant risks. To identify these discrepancies, architects need to extract the architecture from the system implementation, which is both time-consuming and error-prone. To simplify this procedure, many automatic architecture recovery techniques have been developed. Yet, their accuracy is often limited. Architects must still invest significant effort in refining recovery results to ensure they accurately reflect the implemented architecture.
To reduce such manual effort, we introduce SemRef, a framework that combines LLMs with dependency analysis to automatically refine architectures recovered by existing architecture recovery tools. By leveraging the LLM's semantic understanding capabilities and integrating structural dependencies, SemRef enhances both the accuracy and the comprehension of recovered architectures. To evaluate SemRef, we tested on 9 projects with published ground-truth architectures and 10 state-of-the-art architecture recovery tools. 5 commonly used metrics are adopted to evaluate the effectiveness of SemRef. The results show that SemRef improves accuracy across various metrics, with normalized gains ranges from 17.72% to 43.35%. Specifically, for MoJoFM and 𝑎2𝑎 𝑎𝑑 𝑗 metrics, SemRef achieves relative improvements of 118.57% and 100.41%, respectively. Moreover, SemRef is highly scalable. It maintains stable performance across projects ranging from thousands to trillions of lines of code with the cost scale linearly with project size. Further, we test SemRef on various LLMs to demonstrate its generalizability
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- SSAR: A Novel Software Architecture Recovery Approach Enhancing Accuracy and ScalabilityWei Ding, Ran Mo, Chaochao Wu, Haopeng SongICSE 2026
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationDominik Fuchß, Tobias Hey, Jan Keim, Haoyu Liu 等ICSE 2025 · 被引用 8 次
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software RefactoringKhouloud Oueslati, Maxime Lamothe, Foutse KhomhICSE 2026 · 被引用 1 次
- RepLLM: Toward Automatically Reproducing Network Research ResultsYining Jiang, Yunxin Xu, Wenyun Xu, Yufan Zhu 等SIGCOMM 2026
- Semantic-Enhanced Indirect Call Analysis with Large Language ModelsBaijun Cheng, Cen Zhang, Kailong Wang, Ling Shi 等ASE 2024 · 被引用 4 次
