NExtLong: Toward Effective Long-Context Training without Long Documents
Chaochen Gao, Xing Wu, Zijia Lin, Debing Zhang, Songlin Hu
摘要
Large language models (LLMs) with extended context windows have made significant strides yet remain a challenge due to the scarcity of long documents. Existing methods tend to synthesize long-context data but lack a clear mechanism to reinforce the long-range dependency modeling. To address this limitation, we propose NExtLong, a novel framework for synthesizing long-context data through Negative document Extension. NExtLong decomposes a document into multiple meta-chunks and extends the context by interleaving hard negative distractors retrieved from pretraining corpora. This approach compels the model to discriminate long-range dependent context from distracting content, enhancing its ability to model long-range dependencies. Extensive experiments demonstrate that NExtLong achieves significant performance improvements on the HELMET and RULER benchmarks compared to existing long-context synthesis approaches and leading models, which are trained on non-synthetic long documents. These findings highlight NExtLong's ability to reduce reliance on non-synthetic long documents, making it an effective framework for developing advanced long-context LLMs. Our code is available in https://github.com/caskcsg/ longcontext/tree/main/NExtLong .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory AgentHongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen 等ICLR 2026 · 被引用 231 次
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context InstructionsChaochen Gao, Xing Wu, Zijia Lin, Debing Zhang 等NeurIPS 2025 · 被引用 8 次
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu 等ICLR 2026 · 被引用 5 次
- Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free MethodXinhao Xu, Jiaxin Li, Hui Chen, Zijia Lin 等ACL 2025 · 被引用 1 次
- LiteraryQA: Towards Effective Evaluation of Long-document Narrative QATommaso Bonomo, Luca Gioffré, Roberto NavigliEMNLP 2025 · 被引用 1 次
它引用的顶会 Paper40
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining DataHaoran Deng, Yingyu Lin, Zhenghao Lin, Xiao Liu 等ICLR 2026 · 被引用 6 次
- Re³Syn: A Dependency-Based Data Synthesis Framework for Long-Context Post-trainingZhiyang Zhang, Ziqiang Liu, Huiming Wang, Renke Shan 等ACL 2025 · 被引用 4 次
- Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language ModelsLongze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng 等ACL 2024 · 被引用 4 次
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMsJunlong Jia, Xing Wu, Chaochen Gao, Ziyang Chen 等AAAI 2026
- EntropyLong: Effective Long-Context Training via Predictive UncertaintyJunlong Jia, Ziyang Chen, Xing Wu, Chaochen Gao 等ICLR 2026 · 被引用 6 次
