CLP: Efficient and Scalable Search on Compressed Text Logs
Kirk Rodrigues, Yu Luo, Ding Yuan
摘要
This thesis presents the design and implementation of CLP, a tool capable of losslessly compressing unstructured text logs while enabling fast searches directly on the compressed data. Log search and log archiving, despite being critical problems, are mutually exclusive. Widely used log-search tools like Elasticsearch and Splunk Enterprise index the logs to provide fast search performance, yet the size of the index is within the same order of magnitude as the raw log size. Commonly used log archival and compression tools like Gzip provide high compression ratios, yet searching archived logs is a slow and painful process as it first requires decompressing the logs. In contrast, CLP achieves significantly higher compression ratio than all commonly used compressors, yet delivers fast search performance that is comparable or even better than Elasticsearch and Splunk Enterprise. In addition, CLP outperforms Elasticsearch and Splunk Enterprise's log ingestion performance by over 13×, and we show CLP scales to petabytes of logs. CLP's gains come from using a tuned, domainspecific compression and search algorithm that exploits the significant amount of repetition in text logs. Hence, CLP enables efficient search and analytics on archived logs, something that was impossible without it. This thesis and my entire tenure as a graduate student would not be possible without the contributions of several individuals. Moreover, I would not have reached this point in my life without a true plethora of people. Although I cannot mention you all by name here, I am sincerely grateful for your kindness.
First and foremost, I am grateful to my advisor, Ding Yuan. Thanking your advisor is somewhat of a procedural requirement for the acknowledgements section of a thesis, but for me, it is a personal requirement. Thank you to Ding for taking a chance on me and for constantly seeing more in me than I see in myself. I cannot fathom how lucky I am to have joined a research group with an advisor who is ambitious, yet practical and kind. I do not think I would be a graduate student, or at least a happy one, without Ding.
I would also like to thank a handful of professors at the University of Toronto. Michael Stumm has always stepped in when I have asked him for help and has helped in my development whenever he had the opportunity. Importantly, he has been an invaluable resource for Ding's research group. Ashvin Goel and David Lie have also helped whenever I have asked. The ECE department is incredibly lucky to have such a group of professional, kind, and intelligent professors. Finally, my thanks to Robin Sacks for helping me cross the threshold to graduate school.
I would also like to thank my defense committee members: Vijay Chidambaram, Ashvin Goel, Bianca Schroeder, Michael Stumm, and of course Ding Yuan. Your time and feedback was instrumental to this thesis and is much appreciated.
My time as a graduate student would not be as fun without the graduate students I got to work with. A heartfelt thank you to Xu Zhao for welcoming me onto several interesting projects, and for helping me enter the world of research. Similarly, thank you to Yongle Zhang for our time together on some of your projects, and for helping me with more than just research. Thank you to Xiang Ren for letting me share some of the glory of your Linux paper. Finally, thank you to Xu, Yongle, Xiang, David Lion, Adrian Chiu, and Serhei Makarov for many interesting discussions, fun times, and a lot of help through the years.
Next, my tremendous thanks to several friends who have been instrumental to my graduate studies and more. First, thank you to Yu Luo for working with me for all these years, challenging me, allowing me to challenge you, and for being someone the entire research group can always ask for help. Thank you, also, to Chaim Pressman who was instrumental in my academic performance, but has also been an invaluable friend since our time in undergraduate studies. Similarly, thank you to Denis Neogi iii and Michael Naccarato for our time together on our capstone project, and for being great friends and role models.
Thank you to my family for their unconditional support throughout my life and especially, for valuing an education above all else.
Lastly, thank you to Devesh Agrawal and Uber for enabling CLP's deployment in an impactful setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- LILAC: Log Parsing using LLMs with Adaptive Parsing CacheZhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li 等FSE 2024 · 被引用 85 次
- HARDLOG: Practical Tamper-Proof System Auditing Using a Novel Audit DeviceAdil Ahmad, Sangho Lee, Marcus PeinadoS&P 2022 · 被引用 46 次
- Investigating Managed Language Runtime Performance: Why JavaScript and Python are 8x and 29x slower than C++, yet Java and Go can be Faster?David Lion, Adrian Chiu, Michael Stumm, Ding YuanUSENIX ATC 2022 · 被引用 31 次
- LogShrink: Effective Log Compression by Leveraging Commonality and Variability of Log DataXiaoyun Li, Hongyu Zhang, Van-Hoang Le, Pengfei ChenICSE 2024 · 被引用 22 次
- LogGrep: Fast and Cheap Cloud Log Storage by Exploiting both Static and Runtime PatternsJunyu Wei, Guangyan Zhang, Junchao Chen, Yang Wang 等EuroSys 2023 · 被引用 18 次
它引用的顶会 Paper2
相关 Paper
- μSlope: High Compression and Fast Search on Semi-Structured LogsRui Wang, Devin Gibson, Kirk Rodrigues, Yu Luo 等OSDI 2024 · 被引用 8 次
- LogCrisp: Fast Aggregated Analysis on Large-scale Compressed Logs by Enabling Two-Phase Pattern Extraction and Vectorized QueriesJunyu Wei, Guangyan Zhang, Junchao Chen, Qi ZhouUSENIX ATC 2025 · 被引用 2 次
- Unlocking the Power of Numbers: Log Compression via Numeric Token ParsingSiyu Yu, Yifan Wu, Ying Li, Pinjia HeASE 2024 · 被引用 5 次
- LogCloud: Fast Search of Compressed Logs on Object StorageZiheng Wang, Junyu Wei, Alex Aiken, Guangyan Zhang 等VLDB 2025 · 被引用 1 次
- : Near-Storage Accelerator for High-Performance Log AnalyticsSeongyoung Kang, Jiyoung An, Jinpyo Kim, Sang-Woo JunMICRO 2021 · 被引用 9 次
