LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMs
Jianghao Chen, Junhong Wu, Yangyifan Xu, Jiajun Zhang
Abstract
Long-context modeling has drawn more and more attention in the area of Large Language Models (LLMs). Continual training with long-context data becomes the de-facto method to equip LLMs with the ability to process long inputs. However, it still remains an open challenge to measure the quality of long-context training data. To address this issue, we propose a Long-context data selection framework with Attention-based Dependency Measurement (LADM), which can efficiently identify high-quality long-context data from a large-scale, multi-domain pre-training corpus. LADM leverages the retrieval capabilities of the attention mechanism to capture contextual dependencies, ensuring a comprehensive quality measurement of long-context data. Experimental results show that our LADM framework significantly boosts the performance of LLMs on multiple long-context tasks with only 1B tokens for continual training. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4140a5d-58d6-465c-bb7b-d767c75abd0bCited by top-tier papers7
- An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical ReasoningWei Sun, Qianlong Du, Fuwei Cui, Jiajun ZhangACL 2025 · 15 citations
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction TuningManish Nagaraj, Sakshi Choudhary, Utkarsh Saxena, Deepak Ravikumar et al.ICML 2026 · 2 citations
- CGMIS: Concept-Graph Based Multi-Hop Instructions Synthesis for Enhancing Long-Context ReasoningZechen Sun, Zecheng Tang, Juntao Li, Wenpeng Hu et al.AAAI 2026
- SHIFT: Selected Helpful Informative Frame for Video-guided Machine TranslationBoyu Guan, Chuang Han, Yining Zhang, Yupu Liang et al.EMNLP 2025
- CROP: Contextual Region-Oriented Visual Token PruningJiawei Guo, Feifei Zhai, Pu Jian, Qianrun Wei et al.EMNLP 2025
Builds on21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining DataHaoran Deng, Yingyu Lin, Zhenghao Lin, Xiao Liu et al.ICLR 2026 · 6 citations
- Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language ModelsLongze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng et al.ACL 2024 · 4 citations
- Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language ModelsBoheng Sheng, Jiacheng Yao, Meicong Zhang, Guoxiu HeACL 2025 · 9 citations
- SEAL: Scaling to Emphasize Attention for Long-Context RetrievalChanghun Lee, Minsang Seok, Jungyu Jin, Younghyun Cho et al.ACL 2025
- How to Train Long-Context Language Models (Effectively)Tianyu Gao, Alexander Wettig, Howard Yen, Danqi ChenACL 2025
