RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
Han Zhang, Dongfang Zhao
Abstract
While high-dimensional embedding vectors are being increasingly employed in various tasks like Retrieval-Augmented Generation and Recommendation Systems, popular dimensionality reduction (DR) methods such as PCA and UMAP have rarely been adopted for accelerating the retrieval process due to their inability of preserving the nearest neighbor (NN) relationship among vectors. Empowered by neural networks' optimization capability and the bounding effect of Rayleigh quotient, we propose a Regularized Auto-Encoder (RAE) for k-NN preserving dimensionality reduction. RAE constrains the network parameter variation through regularization terms, adjusting singular values to control embedding magnitude changes during reduction, thus reasonably preserving k-NN relationships. We provide a theoretical analysis demonstrating that the proposed regularization establishes an upper bound on the norm distortion rate of transformed vectors, thereby offering supportive guarantees for k-NN proximity. With modest training overhead, RAE achieves significantly higher recall rate compared to existing DR approaches while maintaining prompt retrieval latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
Related papers
- Regularized Autoencoders for Isometric Representation LearningYonghyeon Lee, Sangwoong Yoon, Minjun Son, Frank Chongwoo ParkICLR 2022 · 46 citations
- QPAD: Quantile-Preserving Approximate Dimension Reduction for Nearest Neighbors Preservation in High-Dimensional Vector SearchJiuzhou Fu, Dongfang ZhaoICDE 2026
- Graph Geometry-Preserving AutoencodersJungbin Lim, Jihwan Kim, Yonghyeon Lee, Cheongjae Jang et al.ICML 2024 · 10 citations
- It's Enough: Relaxing Diagonal Constraints in Linear Autoencoders for RecommendationJaewan Moon, Hye-young Kim, Jongwuk LeeSIGIR 2023 · 3 citations
- NUDGE: Lightweight Non-Parametric Fine-Tuning of Embeddings for RetrievalSepanta Zeighami, Zac Wellmer, Aditya G. ParameswaranICLR 2025
