Understanding Synthetic Context Extension via Retrieval Heads
Xinyu Zhao, Fangcong Yin, Greg Durrett
Abstract
Long-context LLMs are increasingly in demand for applications such as retrieval-augmented generation. To defray the cost of pretraining LLMs over long contexts, recent work takes an approach of synthetic context extension: fine-tuning LLMs with synthetically-generated long-context data. However, it remains unclear how and why this synthetic context extension imparts abilities for downstream long-context tasks. In this paper, we investigate fine-tuning on synthetic data for three long-context tasks that require retrieval and reasoning. We vary the realism of "needle" concepts to be retrieved and diversity of the surrounding "haystack" context, from using LLMs to construct synthetic documents to using templated relations and creating symbolic datasets. Although models trained on synthetic data underperform models trained on the real data, the impacts of both training settings can be understood via a shared feature of the attention computation, retrieval heads (Wu et al., 2025) . The retrieval heads learned from synthetic data have high overlap with retrieval heads learned on real data. Furthermore, there is a strong correlation between the recall of heads learned and the downstream performance of a model, allowing us to interpret and predict the performance of models trained in different settings. Our results shed light on how to interpret synthetic data fine-tuning performance and how to approach creating better data for learning realworld LLM capabilities over long contexts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a46a6ac6-636e-4ed9-ba32-4e7d24110d61Cited by top-tier papers4
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė et al.ICLR 2026 · 6 citations
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu et al.ICLR 2026 · 5 citations
- Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-rankingWuwei Zhang, Fangcong Yin, Howard Yen, Danqi Chen et al.EMNLP 2025
- Revealing Long-context Potential of Attention Heads via Frequency KernelsSenyu Han, Yilu Cao, Kai Yu, Lu ChenICML 2026
Builds on18
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng et al.NeurIPS 2024 · 212 citations
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 207 citations
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue et al.ICML 2024 · 204 citations
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov et al.ICLR 2024 · 113 citations
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 100 citations
Related papers
- From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic DataZheyang Xiong, Vasilis Papageorgiou, Kangwook Lee, Dimitris PapailiopoulosICLR 2025
- Training with "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-Context TasksYijiong Yu, Yongfeng Huang, Zhixiao Qi, Zhe ZhouAAAI 2025 · 5 citations
- Scaling Instruction-tuned LLMs to Million-token Contexts via Hierarchical Synthetic Data GenerationLinda He, Jue Wang, Maurice Weber, Shang Zhu et al.ICLR 2025
- Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training EfficiencyEhsan Doostmohammadi, Marco KuhlmannEMNLP 2025
- SEAL: Scaling to Emphasize Attention for Long-Context RetrievalChanghun Lee, Minsang Seok, Jungyu Jin, Younghyun Cho et al.ACL 2025
