L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression
Junxuan Zhang, Zhengxue Cheng, Yan Zhao, Shihao Wang, Dajiang Zhou, Guo Lu, Li Song
摘要
Learning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity design while maintaining compression performance. We introduce a novel Learned Lossless Low-complexity Text Compression method (L3TC). Specifically, we conduct extensive experiments demonstrating that RWKV models achieve the fastest decoding speed with a moderate compression ratio, making it the most suitable backbone for our method. Second, we propose an outlier-aware tokenizer that uses a limited vocabulary to cover frequent tokens while allowing outliers to bypass the prediction and encoding. Third, we propose a novel high-rank reparameterization strategy that enhances the learning capability during training without increasing complexity during inference. Experimental results validate that our method achieves 48% bit saving compared to gzip compressor. Besides, L3TC offers compression performance comparable to other learned compressors, with a 50× reduction in model parameters. More importantly, L3TC is the fastest among all learned compressors, providing real-time decoding speeds up to megabytes per second. Our code is available at https://github.com/alipay/L3TC-leveraging-rwkv-for- learned-lossless-low-complexity-text-compression.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt 等ICLR 2024 · 被引用 243 次
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
- RepVGG: Making VGG-Style ConvNets Great AgainXiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han 等CVPR 2021
相关 Paper
- Finite-State Autoregressive Entropy Coding for Efficient Learned Lossless CompressionYufeng Zhang, Hang Yu, Jianguo Li, Weiyao LinICLR 2024 · 被引用 5 次
- 500xCompressor: Generalized Prompt Compression for Large Language ModelsZongqian Li, Yixuan Su, Nigel CollierACL 2025 · 被引用 35 次
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard 等NeurIPS 2025 · 被引用 5 次
- PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics DatabaseHui Sun, Yanfeng Ding, Liping Yi, Huidong Ma 等KDD 2025
- RazorAttention: Efficient KV Cache Compression Through Retrieval HeadsHanlin Tang, Yang Lin, Jing Lin, Qingsen Han 等ICLR 2025
