An Empirical Study of Python Library Migration Using Large Language Models
Mohayeminul Islam, Ajay Kumar Jha, May Mahmoud, Ildar Akhmetov, Sarah Nadi
Abstract
Library migration is the process of replacing one library with another library that provides similar functionality. Manual library migration is time consuming and error prone, as it requires developers to understand the APIs of both libraries, map them, and perform the necessary code transformations. Large Language Models (LLMs) are shown to be effective at generating and transforming code as well as finding similar code, which are necessary upstream tasks for library migration. Such capabilities suggest that LLMs may be suitable for library migration. Accordingly, this paper investigates the effectiveness of LLMs for migration between Python libraries. We evaluate three LLMs, LLama 3.1, GPT-4o mini, and GPT-4o on PyMigBench, where we migrate 321 real-world library migrations that include 2,989 migration-related code changes. To measure correctness, we (1) compare the LLM’s migrated code with the developers’ migrated code in the benchmark and (2) run the unit tests available in the client repositories. We find that LLama 3.1, GPT-4o mini, and GPT-4o correctly migrate 89%, 89%, and 94% of the migration-related code changes, respectively. We also find that 36%, 52% and 64% of the LLama 3.1, GPT-4o mini, and GPT-4o migrations pass the same tests that passed in the developer’s migration. To ensure the LLMs are not reciting the migrations, we also evaluate them on 10 new repositories where the migration never happened. Overall, our results suggest that LLMs can be effective in migrating code between libraries, but we also identify some open challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aab9176d-1635-4b0d-aabe-5adeb5d331cbBuilds on6
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- Keep me Updated: An Empirical Study of Third-Party Library Updatability on AndroidErik Derr, Sven Bugiel, Sascha Fahl, Yasemin Acar et al.CCS 2017 · 196 citations
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program RepairYuxiang Wei, Chunqiu Steven Xia, Lingming ZhangFSE 2023 · 111 citations
- SOAR: A Synthesis Approach for Data Science API RefactoringAnsong Ni, Daniel Ramos, Aidan Z. H. Yang, Inês Lynce et al.ICSE 2021 · 27 citations
Related papers
- Pig: Leveraging Large Language Models for Python Library MigrationsMiryeong Kang, Wonseok Oh, Gabin An, Hakjoo OhFSE 2026
- Characterizing Python Library MigrationsMohayeminul Islam, Ajay Kumar Jha, Ildar Akhmetov, Sarah NadiFSE 2024 · 3 citations
- LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-Based Code CompletionChong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng et al.ICSE 2025 · 3 citations
- TransLibEval: Demystify Large Language Models' Capability in Third-Party Library-Targeted Code TranslationPengyu Xue, Kunwu Zheng, Zhen Yang, Yifei Pei et al.FSE 2026 · 1 citation
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun et al.AAAI 2026 · 10 citations
