Revisiting Long-context Modeling from Context Denoising Perspective
Zecheng Tang, Baibei Ji, Juntao Li, Lijun Wu, Haijia Gui, Min Zhang
Abstract
Long-context models (LCMs) have demonstrated great potential in processing long sequences, facilitating many real-world applications. The success of LCMs can be attributed to their ability to locate implicit critical information within the context for further prediction. However, recent research reveals that LCMs are often susceptible to contextual noise, i.e., irrelevant tokens, that can mislead model attention. In this paper, we conduct a fine-grained analysis of the context noise and propose an effective metric, the Integrated Gradient (IG) score, to detect and quantify the noise information within the context. Our findings reveal that even simple mitigation of detected context noise can substantially boost the model's attention on critical tokens and benefit subsequent predictions. Building on this insight, we propose Context Denoising Training (CDT), a straightforward yet effective training strategy that improves attention on critical tokens while reinforcing their influence on model predictions. Extensive experiments across four tasks, under both context window scaling and long-context alignment settings, demonstrate the superiority of CDT. Notably, when trained with CDT, an open-source 8B model can achieve performance (50.92) comparable to GPT-4o (51.00).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on31
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
Related papers
- Integral Transformer: Denoising Attention, Not Too Much Not Too LittleIvan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing ChenEMNLP 2025
- LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMsJianghao Chen, Junhong Wu, Yangyifan Xu, Jiajun ZhangACL 2025
- LOGO - Long cOntext aliGnment via efficient preference OptimizationZecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu et al.ICML 2025
- Efficient OpAmp Adaptation for Zoom Attention to Golden ContextsHaoyuan Wu, Rui Ming, Haisheng Zheng, Zhuolun He et al.ACL 2025 · 1 citation
- A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context CompressionChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li et al.ACL 2025
