Understanding Synthetic Context Extension via Retrieval Heads
Xinyu Zhao, Fangcong Yin, Greg Durrett
摘要
Long-context LLMs are increasingly in demand for applications such as retrieval-augmented generation. To defray the cost of pretraining LLMs over long contexts, recent work takes an approach of synthetic context extension: fine-tuning LLMs with synthetically-generated long-context data. However, it remains unclear how and why this synthetic context extension imparts abilities for downstream long-context tasks. In this paper, we investigate fine-tuning on synthetic data for three long-context tasks that require retrieval and reasoning. We vary the realism of "needle" concepts to be retrieved and diversity of the surrounding "haystack" context, from using LLMs to construct synthetic documents to using templated relations and creating symbolic datasets. Although models trained on synthetic data underperform models trained on the real data, the impacts of both training settings can be understood via a shared feature of the attention computation, retrieval heads (Wu et al., 2025) . The retrieval heads learned from synthetic data have high overlap with retrieval heads learned on real data. Furthermore, there is a strong correlation between the recall of heads learned and the downstream performance of a model, allowing us to interpret and predict the performance of models trained in different settings. Our results shed light on how to interpret synthetic data fine-tuning performance and how to approach creating better data for learning realworld LLM capabilities over long contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė 等ICLR 2026 · 被引用 6 次
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu 等ICLR 2026 · 被引用 5 次
- Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-rankingWuwei Zhang, Fangcong Yin, Howard Yen, Danqi Chen 等EMNLP 2025
- Revealing Long-context Potential of Attention Heads via Frequency KernelsSenyu Han, Yilu Cao, Kai Yu, Lu ChenICML 2026
它引用的顶会 Paper18
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 等NeurIPS 2024 · 被引用 212 次
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 被引用 207 次
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue 等ICML 2024 · 被引用 204 次
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 被引用 100 次
相关 Paper
- From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic DataZheyang Xiong, Vasilis Papageorgiou, Kangwook Lee, Dimitris PapailiopoulosICLR 2025
- Training with "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-Context TasksYijiong Yu, Yongfeng Huang, Zhixiao Qi, Zhe ZhouAAAI 2025 · 被引用 5 次
- Scaling Instruction-tuned LLMs to Million-token Contexts via Hierarchical Synthetic Data GenerationLinda He, Jue Wang, Maurice Weber, Shang Zhu 等ICLR 2025
- Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training EfficiencyEhsan Doostmohammadi, Marco KuhlmannEMNLP 2025
- SEAL: Scaling to Emphasize Attention for Long-Context RetrievalChanghun Lee, Minsang Seok, Jungyu Jin, Younghyun Cho 等ACL 2025
