Rottnest: Indexing Data Lakes for Search
Ziheng Wang, Sasha Krassovsky, Conor Kennedy, Alex Aiken, Weston Pace, Rain Jiang, Huayi Zhang, Chenyu Jiang, Wei Xu
摘要
Data lakes have become widely popular in managing enterprise data. Their widespread integration with query engines has allowed them to displace specialized data warehouses as the single source of truth for enterprise data. While the columnar storage format and block min-max indices allow query engines to achieve competitive performance on relational data analytics queries, they are not yet suitable for other search-oriented queries like full text and vector nearest neighbor search. We present Rottnest, a general system that builds additional lightweight indices on top of data lakes. We show that our system is more cost efficient compared to un-indexed data lakes or specialized databases across several orders of magnitude of total query loads and operating time horizons.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- An Efficient and Robust Framework for Approximate Nearest Neighbor Search with Attribute ConstraintMengzhao Wang, Lingwei Lv, Xiaoliang Xu, Yuxiang Wang 等NeurIPS 2023 · 被引用 70 次
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo 等VLDB 2024 · 被引用 59 次
- ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured DataLiana Patel, Peter Kraft, Carlos Guestrin, Matei ZahariaSIGMOD 2024 · 被引用 58 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 被引用 45 次
相关 Paper
- Cracking Vector Search IndexesVasilis Mageirakos, Bowen Wu, Gustavo AlonsoVLDB 2025 · 被引用 6 次
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 被引用 1 次
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 被引用 4 次
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 被引用 118 次
- GraphAr: An Efficient Storage Scheme for Graph Data in Data LakesXue Li, Weibin Zeng, Zhibin Wang, Diwen Zhu 等VLDB 2025 · 被引用 5 次
