Scaling Out Schema-free Stream Joins
Damjan Gjurovski, Sebastian Michel
摘要
In this work, we consider computing natural joins over massive streams of JSON documents that do not adhere to a specific schema. We first propose an efficient and scalable partitioning algorithm that uses the main principles of association analysis to identify patterns of co-occurrence of the attributevalue pairs within the documents. Data is then accordingly forwarded to compute nodes and locally joined using a novel FPtree-based join algorithm. By compactly storing the documents and efficiently traversing the FP-tree structure, the proposed join algorithm can operate on large input sizes and provide results in real-time. We discuss data-dependent scalability limitations that are inherent to natural joins over schema-free data and show how to practically circumvent them by artificially expanding the space of possible attribute-value pairs. The proposed algorithms are realized in the Apache Storm stream processing framework. Through extensive experiments with real-world as well as synthetic data, we evaluate the proposed algorithms and show that they outperform competing approaches.
805
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- JSONSki: streaming semi-structured data with bit-parallel fast-forwardingLin Jiang, Zhijia ZhaoASPLOS 2022 · 被引用 13 次
- AJoin: Ad-hoc Stream Joins at ScaleJeyhun Karimov, Tilmann Rabl, Volker MarklVLDB 2020 · 被引用 14 次
- Distributed Streaming Set Similarity JoinJianye Yang, Wenjie Zhang, Xiang Wang, Ying Zhang 等ICDE 2020 · 被引用 14 次
- Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and QualityXilin Tang, Feng Zhang, Shuhao Zhang, Yani Liu 等SIGMOD 2025 · 被引用 1 次
- dsJSON: A Distributed SQL JSON ProcessorMajid Saeedan, Ahmed Eldawy, Zhijia ZhaoSIGMOD 2023 · 被引用 1 次
