Sim-Piece: Highly Accurate Piecewise Linear Approximation through Similar Segment Merging
Xenophon Kitsios, Panagiotis Liakos, Katia Papakonstantinopoulou, Yannis Kotidis
摘要
Approximating series of timestamped data points using a sequence of line segments with a maximum error guarantee is a fundamental data compression problem, termed as piecewise linear approximation (PLA). Due to the increasing need to analyze massive collections of time-series data in diverse domains, the problem has recently received significant attention, and recent PLA algorithms that have emerged do help us handle the overwhelming amount of information, at the cost of some precision loss. More specifically, these algorithms entail a trade-off between the maximum precision loss and the space savings achieved. However, advances in the area of lossless compression are undercutting the offerings of PLA techniques in real datasets. In this work, we propose Sim-Piece, a novel lossy compression algorithm for time-series data that optimizes the space requirements of representing PLA line segments, by finding the minimum number of groups we can organize these segments into, to represent them jointly. Our experimental evaluation demonstrates that our approach readily outperforms competing techniques, attaining compression ratios with more than twofold improvement on average over what PLA algorithms can offer. This allows for providing significantly higher accuracy with equivalent space requirements. Moreover, our algorithm, due to the simplicity of its merging phase, imposes little overhead while compacting the PLA description, offering a significantly improved trade-off between space and running time. The aforementioned benefits of our approach significantly improve the efficiency in which we can store time-series data, while allowing a tight maximum error in the representation of their values.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Serf: Streaming Error-Bounded Floating-Point CompressionRuiyuan Li, Zechao Chen, Ruyun Lu, Xiaolong Xu 等SIGMOD 2025 · 被引用 7 次
- Learned Compression of Nonlinear Time Series with Random AccessAndrea Guerra, Giorgio Vinciguerra, Antonio Boffa, Paolo FerraginaICDE 2025 · 被引用 5 次
- Improving Time Series Data Compression in Apache IoTDBYuxin Tang, Feng Zhang, Jiawei Guan, Yuan Tian 等VLDB 2025 · 被引用 2 次
- Largest Triangle Sampling for Visualizing Time Series in DatabaseLei Rui, Xiangdong Huang, Shaoxu Song, Chen Wang 等SIGMOD 2025 · 被引用 1 次
它引用的顶会 Paper3
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- Frequency Domain Data Encoding in Apache IoTDBHaoyu Wang, Shaoxu SongVLDB 2023 · 被引用 17 次
- On Compressing Temporal GraphsPanagiotis Liakos, Katia Papakonstantinopoulou, Theodore Stefou, Alex DelisICDE 2022 · 被引用 8 次
相关 Paper
- CIVET: Exploring Compact Index for Variable-Length Subsequence Matching on Time SeriesHaoran Xiong, Hang Zhang, Zeyu Wang, Zhenying He 等VLDB 2024 · 被引用 4 次
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 被引用 9 次
- REGER: Reordering Time Series Data for Regression EncodingJinzhao Xiao, Wendi He, Shaoxu Song, Xiangdong Huang 等ICDE 2024 · 被引用 1 次
- Fast Min-ϵ Segmented Regression using Constant-Time Segment MergingAnsgar Lößer, Max Schlecht, Florian Schintke, Joel Witzke 等ICML 2025
- Approximate Analytics System over Compressed Time Series with Tight Deterministic Error GuaranteesChunbin Lin, Etienne Boursier, Yannis PapakonstantinouVLDB 2020 · 被引用 505 次
