EditFlow: Benchmarking and Optimizing Code Edit Recommendation Systems via Reconstruction of Developer Flows
Chenyan Liu, Yun Lin, Jiaxin Chang, Jiawei Liu, Binhang Qi, Bo Jiang, Zhiyong Huang, Jin Song Dong
Abstract
Large language models (LLMs) for code editing have achieved remarkable progress, yet recent empirical studies reveal a fundamental disconnect between technical accuracy and developer productivity . Despite their strong benchmark performance, developers complete tasks 19% slower when using AI assistance, with over 68.81% of recommendations disrupting their mental flow. This misalignment stems from the use of static commit snapshots that lack temporal information, causing models to optimize for end results rather than the incremental, context-sensitive steps that align with developers’ natural reasoning process. To bridge this gap, we present EditFlow , which benchmarks and optimizes subsequent code edit recommendation systems through the reconstruction of developer editing flows. EditFlow addresses three key challenges. First, collecting edit-order data that reflects developers’ flow is inherently difficult: manual annotation introduces prohibitive overhead, while development logs capture only single trajectories instead of all plausible editing flows. Second, benchmarking recommendation performance against developers’ ongoing editing flow requires a digital-twin-like simulation that can faithfully simulate the editing process. Third, existing heterogeneous systems vary drastically in scale and architecture, posing challenges for developing a unified optimization strategy that endows all models with mental-flow awareness regardless of design or capability. To overcome these challenges, we propose three tightly coupled components: (1) a prompt auto-tuning mechanism that learns an optimized prompt for inferring the relative order between two edits, (2) a digital twin that replays reconstructed edit sequences to simulate developers’ editing process, and (3) EditFlow , a unified optimization strategy that optimizes the flow continuity of subsequent edit suggestions based on developers’ ongoing flow. Evaluations across diverse benchmarks, including manually annotated commits, real-world industrial code, and open-source repositories, show that EditFlow improves order reconstruction accuracy by 63.81%, reduces flow violations by over 75%, and boosts recommendation precision by 66.99%. A user study with 32 developers further demonstrates 25.11% faster task completion and significantly higher perceived recommendation quality. To the best of our knowledge, EditFlow is the first to evaluate and optimize code edit recommendation systems from the perspective of developers’ mental flow, establishing flow-awareness as a new dimension for advancing human-AI code collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8355673-a317-44d0-9d27-bc19d32ef894Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng et al.ICML 2024 · 669 citations
- A syntax-guided edit decoder for neural program repairQihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang et al.FSE 2021 · 214 citations
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li et al.ASE 2022 · 81 citations
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu et al.FSE 2023 · 69 citations
- CodePlan: Repository-Level Coding using LLMs and PlanningRamakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D. C. et al.FSE 2024 · 67 citations
Related papers
- Grace: Language Models Meet Code EditsPriyanshu Gupta, Avishree Khare, Yasharth Bajpai, Saikat Chakraborty et al.FSE 2023 · 15 citations
- EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative DecodingPeiding Wang, Li Zhang, Fang Liu, Yinghao Zhu et al.ASE 2025 · 3 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringTanghaoran Zhang, Yue Yu, Xinjun Mao, Shangwen Wang et al.ICSE 2025 · 3 citations
- Learning Project-wise Subsequent Code Edits via Interleaving Neural-based Induction and Tool-based DeductionChenyan Liu, Yun Lin, Yuhuan Huang, Jiaxin Chang et al.ASE 2025 · 1 citation
