Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
Longze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng, Hao Sun, Yunshui Li, Run Luo, Min Yang
摘要
Long-context modeling capabilities are important for large language models (LLMs) in various applications. However, directly training LLMs with long context windows is insufficient to enhance this capability since some training samples do not exhibit strong semantic dependencies across long contexts. In this study, we propose a data mining framework ProLong 1 that can assign each training sample with a long dependency score, which can be used to rank and filter samples that are more advantageous for enhancing long-context modeling abilities in LLM training. Specifically, we first use delta perplexity scores to measure the Dependency Strength between text segments in a given document. Then, we refine this metric based on the Dependency Distance of these segments to incorporate spatial relationships across long contexts. Final results are calibrated with a Dependency Specificity metric to prevent trivial dependencies introduced by repetitive patterns. Moreover, a random sampling approach is proposed to optimize the computational efficiency of ProLong. Comprehensive experiments on multiple benchmarks indicate that ProLong effectively identifies documents that carry long dependencies, and LLMs trained on these documents exhibit significantly enhanced long-context modeling capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MDCure: A Scalable Pipeline for Multi-Document Instruction-FollowingGabrielle Kaili-May Liu, Bowen Shi, Avi Caciularu, Idan Szpektor 等ACL 2025 · 被引用 13 次
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess 等ICLR 2026 · 被引用 9 次
- Re³Syn: A Dependency-Based Data Synthesis Framework for Long-Context Post-trainingZhiyang Zhang, Ziqiang Liu, Huiming Wang, Renke Shan 等ACL 2025 · 被引用 4 次
- Flora: Effortless Context Construction to Arbitrary Length and ScaleTianxiang Chen, Zhentao Tan, Xiaofan Bo, Yue Wu 等AAAI 2026 · 被引用 2 次
- LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMsYushi Bai, Jiajie Zhang, Xin Lv, Linzhi Zheng 等ICLR 2025
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
相关 Paper
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining DataHaoran Deng, Yingyu Lin, Zhenghao Lin, Xiao Liu 等ICLR 2026 · 被引用 6 次
- LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMsJianghao Chen, Junhong Wu, Yangyifan Xu, Jiajun ZhangACL 2025
- NExtLong: Toward Effective Long-Context Training without Long DocumentsChaochen Gao, Xing Wu, Zijia Lin, Debing Zhang 等ICML 2025
- EntropyLong: Effective Long-Context Training via Predictive UncertaintyJunlong Jia, Ziyang Chen, Xing Wu, Chaochen Gao 等ICLR 2026 · 被引用 6 次
- Probing How Scalable Table Data Enhances General Long-Context ReasoningHuaibing Xie, Guoliang Zhao, Yang Liu, Shihan Dou 等ICML 2026
