Improving Code Localization with Repository Memory
Boshi Wang, Weijian Xu, Yunsheng Li, Xuemei Gao, Yujia Xie, Huan Sun, Dongdong Chen
摘要
Code localization is a fundamental challenge in repository-level software engineering tasks such as bug fixing. While existing methods equip language agents with comprehensive tools/interfaces to fetch information from the repository, they overlook the critical aspect of memory, where each instance is typically handled from scratch assuming no prior repository knowledge. In contrast, human developers naturally build long-term repository memory, such as the functionality of key modules and associations between various bug types and their likely fix locations. In this work, we augment language agents with such memory by leveraging a repository's commit history - a rich yet underutilized resource that chronicles the codebase's evolution. We introduce tools that allow the agent to retrieve from a non-parametric memory encompassing recent historical commits and linked issues, as well as functionality summaries of actively evolving parts of the codebase identified via commit patterns. We demonstrate that augmenting such a memory can significantly improve LocAgent, a state-of-the-art localization framework, on both SWE-bench-verified and the more recent SWE-bench-live benchmarks. Our research contributes towards developing agents that can accumulate and leverage past experience for long-horizon tasks, more closely emulating the expertise of human developers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Structurally Aligned Subtask-Level Memory for Software Engineering AgentsKangning Shen, Jingyuan Zhang, Chenxi Sun, Wencong Zeng 等ICML 2026 · 被引用 9 次
- Context-Aware Reasoner: Enhancing Contextual Reasoning in Multimodal Large Language ModelsZhe Zheng, Wenqi Zhang, Xiaohe Zhou, Guiyang Hou 等ICML 2026
- Closing the Loop: Universal Repository Representation with RPG-EncoderJane Luo, Chengyu Yin, Xin Zhang, Qingtao Li 等ICML 2026
- ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic MemoryWeihao Bo, Shan Zhang, Yanpeng Sun, Jingjing Wu 等CVPR 2026
它引用的顶会 Paper20
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
- RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and GenerationFengji Zhang, Bei Chen, Yue Zhang, Jacky Keung 等EMNLP 2023 · 被引用 110 次
相关 Paper
- LocAgent: Graph-Guided LLM Agents for Code LocalizationZhaoling Chen, Robert Tang, Gangda Deng, Fang Wu 等ACL 2025
- Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided ScalingXu Yang, Jiayuan Zhou, Michael Pacheco, Wenhan Zhu 等ISSTA 2026
- OrcaLoca: An LLM Agent Framework for Software Issue LocalizationZhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang 等ICML 2025
- SWERank: Software Issue Localization with Code RankingRevanth Gangi Reddy, Tarun Suresh, JaeHyeok Doo, Ye Liu 等ICLR 2026 · 被引用 28 次
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
