On the Road to Personalized Code Intelligence: Portraiting and Assisting Developers Based on Their In-IDE Behaviors
Yuhong Liu, Yunhe Su, Zhipeng Peng, Zhiwen Luo, Lin Shi, Zhi Jin, Li Zhang
Abstract
With the advent of powerful large language models (LLMs), research in automated software engineering has increasingly focused on leveraging these models to achieve a deeper semantic understanding of code or to engineer sophisticated agent-based processes. The predominant goal of these efforts is to enhance developer productivity through automated assistance. However, this research trajectory has largely overlooked a critical factor: the developers themselves. Programming is a deeply human and individualized activity; developers exhibit significant variation in their coding styles, tool-chain preferences, domain-specific expertise, and problem-solving strategies. Consequently, the current paradigm of one-size-fits-all code intelligence systems struggles to accommodate the unique characteristics and needs of individual developers. To address this gap, we introduce VirtualME, a novel IDE-embedded data infrastructure designed to model the developer by continuously capturing and interpreting their dynamic programming behaviors and preferences. VirtualME contains three components. (1) Log-level Behavior Extraction: it captures and extracts developers' log-level behaviors (edits, navigations, etc.) from IDE. (2) Task-level Behavior Recognition: it aggregates log-level behaviors into task-level behaviors ("skimming API docs", "iterative debugging", etc.) via a multi-agent pipeline.
(3) Developer-persona Measurement: it builds a rule engine to distill a four-dimensional developer persona: Core Technical Foundation, Practical Development Efficiency, Personal Development Norms, and Technical Adaptability. On top of VirtualME, we propose a solution for personalized repository-level knowledge Q&A by integrating the developer persona into a Chain-of-Thought (CoT) guided agent. We evaluated VirtualME by building a multi-repository benchmark with real-world developer trajectories, balancing correctness and personalization. Experimental results show that VirtualME-enhanced answers outperform generic baselines on five dimensions: correctness, cognitive-level fit, technology-stack relevance, behavioral-pattern alignment, and stylistic preference, yielding an average 33.80% improvement. Our results demonstrate that abundant, continuous developer-behavior data can unlock Personalized Code Intelligence. By integrating this personalized understanding into the code intelligence loop, our approach paves the new way for adaptive and personalized code intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu et al.FSE 2024 · 49 citations
Related papers
- On Behavioral Alignment of Model-Code and Human-Code Understandability via Behavioral ProxiesXiaokai Rong, Aashish Yadavally, Hridya Dhulipala, Anh H. N. Nguyen et al.ISSTA 2026
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 96 citations
- How Do Developers Interact with AI? An Exploratory Study on Modeling Developer Programming BehaviorYinan Wu, Ze Shi Li, Kathryn Thomasset Stolee, Bowen XuFSE 2026
- CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level TasksLingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan et al.ACL 2026
- Enhanced Prompting Framework for Code Summarization with Large Language ModelsMinying Fang, Xing Yuan, Yuying Li, Haojie Li et al.ISSTA 2025 · 3 citations
