Let's (not) just put things in Context: Test-time Training for Long-context LLMs
Rachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan, Sai Surya Duvvuri, Fnu Devvrit, David Brandfonbrener, David Alvarez-Melis, Prajjwal Bhargava, Mihir Kale, Samy Jelassi
Abstract
Progress on training and architecture strategies have enabled LLMs with millions of tokens in context length. However, empirical evidence suggests that such long-context LLMs can consume far more text than they can reliably use. On the other hand, it has been shown that inference-time compute can be used to scale performance of LLMs, often by generating thinking tokens, on challenging tasks involving multi-step reasoning. Through controlled experiments on sandbox long-context tasks, we find that such inference-time strategies show rapid diminishing returns, and fail at long context. We attribute these failures to score dilution, a phenomenon inherent to static self-attention. Further, we show that current inference-time strategies cannot retrieve relevant long-context signals under certain conditions. We propose query-only test-time-training (qTTT) that, through targeted gradients updates on the given context, provably overcomes limitations of static self-attention. We find that this simple shift in how inference-time compute is spent leads to consistently large performance improvements across models and long-context benchmarks. qTTT leads to massive 12.6% and 14.1% points improvements for Qwen3-4B on average across subsets of LongBench-v2 and ZeroScrolls benchmarks. The takeaway is practical: for long context, a small amount of context-specific training is a better use of inference compute than current inference-time scaling strategies like producing more thinking tokens.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfc474e4-2bfb-43fe-a97b-ad16f6bb0489Cited by top-tier papers2
- DeltaEvolve: Accelerating Scientific Discovery through Momentum-Driven EvolutionJiachen Jiang, Tianyu Ding, Zhihui ZhuICML 2026 · 17 citations
- Gated Differentiable Working Memory for Long-Context Language ModelingLingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge et al.ACL 2026 · 4 citations
Builds on23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Understanding the Role of Training Data in Test-Time ScalingAdel Javanmard, Baharan Mirzasoleiman, Vahab MirrokniICLR 2026 · 5 citations
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensWei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao et al.ICML 2026 · 20 citations
- Towards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningWenkai Yang, Shuming Ma, Yankai Lin, Furu WeiNeurIPS 2025 · 141 citations
- Inference-Time Hyper-Scaling with KV Cache CompressionAdrian Lancucki, Konrad Staniszewski, Piotr Nawrot, Edoardo Maria PontiNeurIPS 2025 · 36 citations
- Inference Scaling for Long-Context Retrieval Augmented GenerationZhenrui Yue, Honglei Zhuang, Aijun Bai, Kai Hui et al.ICLR 2025
