PDX: A Data Layout for Vector Similarity Search
Leonardo Kuffó, Elena Krippner, Peter Boncz
Abstract
We propose Partition Dimensions Across (PDX), a data layout for vectors (e.g., embeddings) that, similar to PAX [6], stores multiple vectors in one block, using a vertical layout for the dimensions (Figure 1). PDX accelerates exact and approximate similarity search thanks to its dimension-by-dimension search strategy that operates on multiple-vectors-at-a-time in tight loops. It beats SIMD-optimized distance kernels on standard horizontal vector storage (avg 40% faster), only relying on scalar code that gets auto-vectorized. We combined the PDX layout with recent dimension-pruning algorithms ADSampling [19] and BSA [52] that accelerate approximate vector search. We found that these algorithms on the horizontal vector layout can lose to SIMD-optimized linear scans, even if they are SIMD-optimized. However, when used on PDX, their benefit is restored to 2-7x. We find that search on PDX is especially fast if a limited number of dimensions has to be scanned fully, which is what the dimension-pruning approaches do. We finally introduce PDX-BOND, an even more flexible dimension-pruning strategy, with good performance on exact search and reasonable performance on approximate search. Unlike previous pruning algorithms, it can work on vector data ''as-is'' without preprocessing; making it attractive for vector databases with frequent updates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 7 citations
- Cracking Vector Search IndexesVasilis Mageirakos, Bowen Wu, Gustavo AlonsoVLDB 2025 · 6 citations
- Exqutor: Extended Query Optimizer for Vector-Augmented Analytical QueriesHyunjoon Kim, Chaerim Lim, Hyeonjun An, Rathijit Sen et al.ICDE 2026 · 1 citation
Builds on10
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 354 citations
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li et al.NeurIPS 2021 · 219 citations
- RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor SearchJianyang Gao, Cheng LongSIGMOD 2024 · 83 citations
- High-Dimensional Approximate Nearest Neighbor Search: with Reliable and Efficient Distance Comparison OperationsJianyang Gao, Cheng LongSIGMOD 2023 · 73 citations
Related papers
- DistVS: Large-scale Vector Search with Compute-Memory DisaggregationPeiqi Yin, Xiao Yan, Shiyuan Deng, Hui Li et al.NSDI 2026 · 3 citations
- ANSMET: Approximate Nearest Neighbor Search with Near-Memory Processing and Hybrid Early TerminationYiwei Li, Yuxin Jin, Boyu Tian, Huanchen Zhang et al.ISCA 2025 · 10 citations
- HAP: An Efficient Hamming Space Index Based on Augmented Pigeonhole PrincipleQiyu Liu, Yanyan Shen, Lei ChenSIGMOD 2022 · 10 citations
- Distance Comparison Operations are not Silver Bullets in Vector Similarity Search: A Benchmark Study on their Merits and LimitsZhuanglin Zheng, Yuxiang Zeng, Chenchen Liu, Yunzhen Chi et al.ICDE 2026
- GPS: Revisiting the Data Layout for Disk-based High-Dimensional Vector SearchPeiqi Yin, Xiao Yan, Qihui Zhou, Hui Li et al.SIGMOD 2026 · 2 citations
