Homomorphic Compression: Making Text Processing on Compression Unlimited
Jiawei Guan, Feng Zhang, Siqi Ma, Kuangyu Chen, Yihua Hu, Yuxing Chen, Anqun Pan, Xiaoyong Du
Abstract
Lossless data compression is an effective way to handle the huge transmission and storage overhead of massive text data. Its utility is even more significant today when data volumes are skyrocketing. The concept of operating on compressed data infuses new blood into efficient text management by enabling mainly access-oriented text processing tasks to be done directly on compressed data without decompression. Facing limitations of the existing compressed text processing schemes such as limited types of operations supported, low efficiency, and high space occupation, we address these problems by proposing a homomorphic compression theory. It enables the generalization and characterization of algorithms with compression processing capabilities. On this basis, we develop HOCO, an efficient text data management engine that supports a variety of processing tasks on compressed text. We select three representative compression schemes and implement them combined with homomorphism in HOCO. HOCO supports the extension of homomorphic compression schemes through a modular and object-oriented design and has convenient interfaces for text processing tasks. We evaluate HOCO on six real-world datasets. The three schemes implemented in HOCO show trade-offs in terms of compression ratio, supported operation types, and efficiency. Experiments also show that HOCO can achieve higher throughput in random access and modification operations (averagely 9.18× than the state-of-the-art) and lower latency in text analytic tasks (averagely 7.16× than processing on uncompressed text) without compromising compression efficacy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e9258e47-9ce0-4972-b6b4-1ccb35e83470Cited by top-tier papers5
- Tribase: A Vector Data Query Engine for Reliable and Lossless Pruning Compression using Triangle InequalitiesQian Xu, Juan Yang, Feng Zhang, Junda Pan et al.SIGMOD 2025 · 14 citations
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian et al.VLDB 2025 · 2 citations
- GPU-Accelerated OLTP: An in-Depth Analysis of Concurrency Control SchemesZihan Sun, Yuyu Luo, Yong Zhang, Chao Li et al.ICDE 2026
- A Systematic Study on Early Stopping Metrics in HPO and the Implications of UncertaintyJiawei Guan, Feng Zhang, Jiesong Liu, Xiaoyong Du et al.VLDB 2025
- Enabling Homomorphic Analytical Operations on Compressed Scientific Data with Multi-Stage DecompressionXuan Wu, Sheng Di, Tripti Agarwal, Kai Zhao et al.ICDE 2026
Related papers
- A Unified Framework for Compressed and Encrypted Text Direct ProcessingYani Liu, Feng Zhang, Yu Zhang, Siqi Ma et al.ICDE 2026 · 1 citation
- Enabling Efficient Random Access to Hierarchically-Compressed DataFeng Zhang, Jidong Zhai, Xipeng Shen, Onur Mutlu et al.ICDE 2020 · 22 citations
- hZCCL: Accelerating Collective Communication with Co-Designed Homomorphic CompressionJiajun Huang, Sheng Di, Xiaodong Yu, Yujia Zhai et al.SC 2024 · 13 citations
- F-TADOC: FPGA-Based Text Analytics Directly on Compression with HLSYanliang Zhou, Feng Zhang, Tuo Lin, Yuanjie Huang et al.ICDE 2024 · 1 citation
- Client-optimized algorithms and acceleration for encrypted compute offloadingMcKenzie van der Hagen, Brandon LuciaASPLOS 2022 · 18 citations
