From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
Nathanaël Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber, Alberto Testoni, Marzieh Fadaee, Sandro Pezzelle, Marco Del Tredici
Abstract
Large Language Models (LLMs) are increasingly used in working environments for a wide range of tasks, excelling at solving individual problems in isolation. However, are they also able to effectively collaborate over longterm interactions? To investigate this, we introduce MEMORYCODE, a synthetic multisession dataset designed to test LLMs' ability to track and execute simple coding instructions amid irrelevant information, simulating a realistic setting. While all the models we tested handle isolated instructions well, even stateof-the-art models like GPT-4o and the reasoning model DeepSeek-R1 show degraded performance when instructions are spread across sessions. Our analysis suggests this is due to their failure to retrieve and integrate information over long instruction chains. Our results highlight a fundamental limitation of current LLMs, restricting their ability to collaborate effectively in long interactions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e311c20-1093-4bc2-975c-674f6bb726b3Cited by top-tier papers3
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMsMohammad Tavakoli, Alireza Salemi, Carrie Ye, Mohamed Abdalla et al.ICLR 2026 · 56 citations
- One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving FrameworkQi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang et al.ACL 2026 · 3 citations
- Do LLMs Forget What They Should? Evaluating In-Context Forgetting in Large Language ModelsYuli Qian, Zechuan Yang, Wenbiao Ding, Hongzhi Li et al.ICLR 2026
Builds on15
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
- Chain-of-Thought Reasoning Without PromptingXuezhi Wang, Denny ZhouNeurIPS 2024 · 305 citations
- MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical ReasoningKe Wang, Houxing Ren, Aojun Zhou, Zimu Lu et al.ICLR 2024 · 188 citations
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong et al.EMNLP 2023 · 109 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
Related papers
- RECODE-H: A Benchmark for Research Code Development with Interactive Human FeedbackChunyu Miao, Henry Peng Zou, Yangning Li, Yankai Chen et al.ICLR 2026 · 25 citations
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsTerry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu et al.ICLR 2025
- Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction TuningJian Yang, Wei Zhang, Yibo Miao, Shanghaoran Quan et al.ACL 2025 · 4 citations
- Exploring Distributional Shifts in Large Language Models for Code AnalysisShushan Arakelyan, Rocktim Jyoti Das, Yi Mao, Xiang RenEMNLP 2023 · 14 citations
- WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction TuningZhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang et al.ACL 2024
