DiskJoin: Large-scale Vector Similarity Join with SSD
Yanqi Chen, Xiao Yan, Alexandra Meliou, Eric Lo
摘要
Similarity join-a widely used operation in data science-finds all pairs of items that have distance smaller than a threshold. Prior work has explored distributed computation methods to scale similarity join to large data volumes but these methods require a cluster deployment, and efficiency suffers from expensive inter-machine communication. On the other hand, disk-based solutions are more cost-effective by using a single machine and storing the large dataset on high-performance external storage, such as NVMe SSDs, but in these methods the disk I/O time is a serious bottleneck. In this paper, we propose DiskJoin, the first disk-based similarity join algorithm that can process billion-scale vector datasets efficiently on a single machine. DiskJoin improves disk I/O by tailoring the data access patterns to avoid repetitive accesses and read amplification. It also uses main memory as a dynamic cache and carefully manages cache eviction to improve cache hit rate and reduce disk retrieval time. For further acceleration, we adopt a probabilistic pruning technique that can effectively prune a large number of vector pairs from computation. Our evaluation on real-world, large-scale datasets shows that DiskJoin significantly outperforms alternatives, achieving speedups from 50× to 1000×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li 等NeurIPS 2021 · 被引用 219 次
- SONG: Approximate Nearest Neighbor Search on GPUWeijie Zhao, Shulong Tan, Ping LiICDE 2020 · 被引用 103 次
- An Efficient and Robust Framework for Approximate Nearest Neighbor Search with Attribute ConstraintMengzhao Wang, Lingwei Lv, Xiaoliang Xu, Yuxiang Wang 等NeurIPS 2023 · 被引用 70 次
- Similarity search in the blink of an eye with compressed indicesCecilia Aguerrebere, Ishwar Singh Bhati, Mark Hildebrand, Mariano Tepper 等VLDB 2023 · 被引用 64 次
- Starling: An I/O-Efficient Disk-Resident Graph Index Framework for High-Dimensional Vector Similarity Search on Data SegmentMengzhao Wang, Weizhi Xu, Xiaomeng Yi, Songlin Wu 等SIGMOD 2024 · 被引用 63 次
相关 Paper
- MetricJoin: Leveraging Metric Properties for Robust Exact Set Similarity JoinsManuel Widmoser, Daniel Kocher, Nikolaus Augsten, Willi MannICDE 2023 · 被引用 4 次
- Fast Approximate Similarity Join in Vector DatabasesJiadong Xie, Jeffrey Xu Yu, Yingfan LiuSIGMOD 2025 · 被引用 6 次
- TRIM: Accelerating High-Dimensional Vector Similarity Search with Enhanced Triangle-Inequality-Based PruningYitong Song, Pengcheng Zhang, Chao Gao, Bin Yao 等SIGMOD 2026 · 被引用 1 次
- DISK: A Distributed Framework for Single-Source SimRank with Accuracy GuaranteeYue Wang, Ruiqi Xu, Zonghao Feng, Yulin Che 等VLDB 2021 · 被引用 7 次
- Out-of-Core Parallel Spatial Join Outperforming In-Memory Systems: A BFS-DFS Hybrid ApproachLyuheng Yuan, Da Yan, Akhlaque Ahmad, Jiao Han 等HPDC 2025
