Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
Md Nakhla Rafi, Dong Jae Kim, Tse-Hsun (Peter) Chen, Shaowei Wang
Abstract
Large Language Models (LLMs) show great promise in software engineering tasks like Fault Localization (FL) and Automatic Program Repair (APR). This study investigates the impact of input order and context size on LLM performance in FL, a crucial step for many downstream software engineering tasks. We test different orders for methods using Kendall Tau distances, including "perfect" (where ground truths come first) and "worst" (where ground truths come last), using two benchmarks that consist of both Java and Python projects. Our results indicate a significant bias in order; Top-1 FL accuracy in Java projects drops from 57% to 20%, while in Python projects, it decreases from 38% to approximately 3% when we reverse the code order. Breaking down inputs into smaller contexts helps reduce this bias, narrowing the performance gap in FL from 22% and 6% to just 1% on both benchmarks. We then investigated whether the bias in order was caused by data leakage by renaming the method names with more meaningful alternatives. Our findings indicated that the trend remained consistent, suggesting that the bias was not due to data leakage. We also look at ordering methods based on traditional FL techniques and metrics. Ordering using DepGraph's ranking achieves 48% Top-1 accuracy, which is better than more straightforward ordering approaches like CallGraph DFS . These findings underscore the importance of how we structure inputs, manage contexts, and choose ordering methods to improve LLM performance in FL and other software engineering tasks.
• Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8aad6a6e-1ce8-4f03-a20e-ab49cb88c96bBuilds on9
- Boosting coverage-based fault localization via graph-based representation learningYiling Lou, Qihao Zhu, Jinhao Dong, Xia Li et al.FSE 2021 · 157 citations
- Large Language Models for Test-Free Fault LocalizationAidan Z. H. Yang, Claire Le Goues, Ruben Martins, Vincent J. HellendoornICSE 2024 · 98 citations
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 96 citations
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault LocalizationSungmin Kang, Gabin An, Shin YooFSE 2024 · 69 citations
- Premise Order Matters in Reasoning with Large Language ModelsXinyun Chen, Ryan A. Chi, Xuezhi Wang, Denny ZhouICML 2024 · 59 citations
Related papers
- Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language ModelsAidan Z. H. Yang, Sophia Kolak, Vincent J. Hellendoorn, Ruben Martins et al.ICSE 2025 · 2 citations
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li et al.FSE 2024 · 60 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Repair Ingredients Are All You Need: Improving Large Language Model-Based Program Repair via Repair Ingredients SearchJiayi Zhang, Kai Huang, Jian Zhang, Yang Liu et al.ICSE 2026
- Input Reduction Enhanced LLM-based Program RepairBoyang Yang, Luyao Ren, Xin Yin, Jiadong Ren et al.ICSE 2026 · 1 citation
