POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge Distillation
Yifei Wang, Feng Xiong, Yong Wang, Linjing Li, Xiangxiang Chu, Daniel Dajun Zeng
Abstract
Positional bias (PB), manifesting as non-uniform sensitivity across different contextual locations, significantly impairs long-context comprehension and processing capabilities. Previous studies have addressed PB either by modifying the underlying architectures or by employing extensive contextual awareness training. However, the former approach fails to effectively eliminate the substantial performance disparities, while the latter imposes significant data and computational overhead. To address PB effectively, we introduce Pos2Distill, a position to position knowledge distillation framework. Pos2Distill transfers the superior capabilities from advantageous positions to less favorable ones, thereby reducing the huge performance gaps. The conceptual principle is to leverage the inherent, position-induced disparity to counteract the PB itself. We identify distinct manifestations of PB under retrieval and reasoning paradigms, thereby designing two specialized instantiations: Pos2Distill-R1 and Pos2Distill-R2 respectively, both grounded in this core principle. By employing the Pos2Distill approach, we achieve enhanced uniformity and significant performance gains across all contextual positions in long-context retrieval and reasoning tasks. Crucially, both specialized systems exhibit strong cross-task generalization mutually, while achieving superior performance on their respective tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6918feb5-0575-4f50-bb08-2efcb7cb080fCited by top-tier papers5
- FASA: FREQUENCY-AWARE SPARSE ATTENTIONYifei Wang, Yueqi Wang, Zhenrui Yue, Huimin Zeng et al.ICLR 2026 · 7 citations
- Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented GenerationQianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng et al.ACL 2026 · 3 citations
- Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPOJinquan Zheng, Jia Yuan, Jiacheng Yao, Chenyang Gu et al.ACL 2026 · 1 citation
- AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical RevisitingRenda Li, Hailang Huang, Fei Wei, Feng Xiong et al.AAAI 2026 · 1 citation
- Seeing Is Coding: On the Effectiveness of Vision Language Models in Code UnderstandingYuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen et al.ISSTA 2026 · 1 citation
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Towards Long-window Anchoring in Vision-Language Model DistillationHaoyi Zhou, Shuo Li, Tianyu Chen, Qi Song et al.AAAI 2026 · 1 citation
- Bridging the Tokenizer Gap: Semantics and Distribution-aware Knowledge Transfer for Unbiased Cross-Tokenizer DistillationHuazheng Wang, Yongcheng Jing, Haifeng Sun, Jingyu Wang et al.AAAI 2026
- Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented GenerationJiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei et al.EMNLP 2025
- LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration DistillationZican Dong, Junyi Li, Jinhao Jiang, Mingyu Xu et al.ACL 2025
- Context-aware Biases for Length ExtrapolationAli Veisi, Hamidreza Amirzadeh, Amir MansourianEMNLP 2025 · 2 citations
