Lune

ACL2026Top-tier venue

Bridging the Memorization-Utilization Gap: Near-Lossless Context Compression via Reinforcement Learning

Yujan Ting, Xu Tang, Terrence Chen, Weijing Huang

2026Year

Abstract

Despite recent progress in context compression, we identify a fundamental memorization-utilization gap where models can compress context with near-perfect fidelity yet fail to effectively utilize these compressed representations for downstream tasks. We address this with a holistic training paradigm spanning pre-training, instruction tuning, and reinforcement learning, built upon an average pooling compression. Our key innovation uses outcome-based RL to enable implicit expansion: the model learns to adaptively unfold task-relevant details during generation, interleaving reconstruction with reasoning. We achieve near-lossless 16 × context compression ( ≈ 5.3 × de-coder sequence-length reduction in our current implementation) across 7B and 32B models, recovering over 98% of full-context QA performance and outperforming prior methods by 11 points. Our 32B model demonstrates strong out-of-distribution and length generalization, robustly scaling to 120k-token contexts despite training on no more than 4k tokens, matching full-context performance on NIAH, Long-Bench v2, and multi-hop reasoning. We verify the implicit expansion behavior in experiments.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a4a9248e-8b18-44a0-84ee-dfc5894af108

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines