CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
Alex Thillen, Niels Mündler, Veselin Raychev, Martin Vechev
Abstract
LLM coding agents can generate working code, but their solutions often accumulate complexity, duplication, and architectural debt. Human developers address such issues through refactoring: behavior-preserving program transformations that improve structure and maintainability. We investigate whether agents (i) can execute refactorings reliably and (ii) identify the refactorings that human developers actually chose in real codebases. To this end, we construct CODE-TASTE, a benchmark mined from large multifile open-source refactorings. To score solutions, we combine repository test suites that measure functional correctness with tailored static checks that verify removal of undesired and introduction of desired code patterns using dataflow reasoning. Our results show a clear gap: agents perform well at implementing refactorings that are specified in detail, but often fail to discover the human refactoring choices when given a focus area for changes. A propose-then-implement decomposition improves alignment, and selecting the best-aligned proposal before implementation can yield further gains. CODETASTE provides an evaluation target and a potential preference signal for aligning coding agents with human refactoring decisions in realistic codebases. We release the benchmark, leaderboard, and code 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 804463f4-bbd6-4e73-99b7-41a9241781bdBuilds on5
- FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature ImplementationWei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao et al.ACL 2025 · 40 citations
- SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?Xinyi He, Qian Liu, Mingzhe Du, Lin Yan et al.ICML 2026 · 31 citations
- Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language ModelsZejun Zhang, Zhenchang Xing, Xiaoxue Ren, Qinghua Lu et al.FSE 2024 · 11 citations
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 9 citations
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li et al.ICLR 2025
Related papers
- RefactorBench: Evaluating Stateful Reasoning in Language Agents Through CodeDhruv Gautam, Spandan Garg, Jinu Jang, Neel Sundaresan et al.ICLR 2025
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software RefactoringKhouloud Oueslati, Maxime Lamothe, Foutse KhomhICSE 2026 · 1 citation
- OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic CodingDeming Ding, Shichun Liu, Enhui Yang, Jiahang Lin et al.ACL 2026 · 10 citations
- CodeClash: Benchmarking Goal-Oriented Software EngineeringJohn Yang, Kilian Lieret, Joyce Yang, Carlos Jimenez et al.ICML 2026 · 5 citations
- CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang et al.ICML 2026 · 2 citations
