PDX: A Data Layout for Vector Similarity Search
Leonardo Kuffó, Elena Krippner, Peter Boncz
摘要
We propose Partition Dimensions Across (PDX), a data layout for vectors (e.g., embeddings) that, similar to PAX [6], stores multiple vectors in one block, using a vertical layout for the dimensions (Figure 1). PDX accelerates exact and approximate similarity search thanks to its dimension-by-dimension search strategy that operates on multiple-vectors-at-a-time in tight loops. It beats SIMD-optimized distance kernels on standard horizontal vector storage (avg 40% faster), only relying on scalar code that gets auto-vectorized. We combined the PDX layout with recent dimension-pruning algorithms ADSampling [19] and BSA [52] that accelerate approximate vector search. We found that these algorithms on the horizontal vector layout can lose to SIMD-optimized linear scans, even if they are SIMD-optimized. However, when used on PDX, their benefit is restored to 2-7x. We find that search on PDX is especially fast if a limited number of dimensions has to be scanned fully, which is what the dimension-pruning approaches do. We finally introduce PDX-BOND, an even more flexible dimension-pruning strategy, with good performance on exact search and reasonable performance on approximate search. Unlike previous pruning algorithms, it can work on vector data ''as-is'' without preprocessing; making it attractive for vector databases with frequent updates.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 被引用 7 次
- Cracking Vector Search IndexesVasilis Mageirakos, Bowen Wu, Gustavo AlonsoVLDB 2025 · 被引用 6 次
- Exqutor: Extended Query Optimizer for Vector-Augmented Analytical QueriesHyunjoon Kim, Chaerim Lim, Hyeonjun An, Rathijit Sen 等ICDE 2026 · 被引用 1 次
它引用的顶会 Paper10
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 被引用 354 次
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li 等NeurIPS 2021 · 被引用 219 次
- RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor SearchJianyang Gao, Cheng LongSIGMOD 2024 · 被引用 83 次
- High-Dimensional Approximate Nearest Neighbor Search: with Reliable and Efficient Distance Comparison OperationsJianyang Gao, Cheng LongSIGMOD 2023 · 被引用 73 次
相关 Paper
- DistVS: Large-scale Vector Search with Compute-Memory DisaggregationPeiqi Yin, Xiao Yan, Shiyuan Deng, Hui Li 等NSDI 2026 · 被引用 3 次
- ANSMET: Approximate Nearest Neighbor Search with Near-Memory Processing and Hybrid Early TerminationYiwei Li, Yuxin Jin, Boyu Tian, Huanchen Zhang 等ISCA 2025 · 被引用 10 次
- HAP: An Efficient Hamming Space Index Based on Augmented Pigeonhole PrincipleQiyu Liu, Yanyan Shen, Lei ChenSIGMOD 2022 · 被引用 10 次
- Distance Comparison Operations are not Silver Bullets in Vector Similarity Search: A Benchmark Study on their Merits and LimitsZhuanglin Zheng, Yuxiang Zeng, Chenchen Liu, Yunzhen Chi 等ICDE 2026
- GPS: Revisiting the Data Layout for Disk-based High-Dimensional Vector SearchPeiqi Yin, Xiao Yan, Qihui Zhou, Hui Li 等SIGMOD 2026 · 被引用 2 次
