RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
Han Zhang, Dongfang Zhao
摘要
While high-dimensional embedding vectors are being increasingly employed in various tasks like Retrieval-Augmented Generation and Recommendation Systems, popular dimensionality reduction (DR) methods such as PCA and UMAP have rarely been adopted for accelerating the retrieval process due to their inability of preserving the nearest neighbor (NN) relationship among vectors. Empowered by neural networks' optimization capability and the bounding effect of Rayleigh quotient, we propose a Regularized Auto-Encoder (RAE) for k-NN preserving dimensionality reduction. RAE constrains the network parameter variation through regularization terms, adjusting singular values to control embedding magnitude changes during reduction, thus reasonably preserving k-NN relationships. We provide a theoretical analysis demonstrating that the proposed regularization establishes an upper bound on the norm distortion rate of transformed vectors, thereby offering supportive guarantees for k-NN proximity. With modest training overhead, RAE achieves significantly higher recall rate compared to existing DR approaches while maintaining prompt retrieval latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
相关 Paper
- Regularized Autoencoders for Isometric Representation LearningYonghyeon Lee, Sangwoong Yoon, Minjun Son, Frank Chongwoo ParkICLR 2022 · 被引用 46 次
- QPAD: Quantile-Preserving Approximate Dimension Reduction for Nearest Neighbors Preservation in High-Dimensional Vector SearchJiuzhou Fu, Dongfang ZhaoICDE 2026
- Graph Geometry-Preserving AutoencodersJungbin Lim, Jihwan Kim, Yonghyeon Lee, Cheongjae Jang 等ICML 2024 · 被引用 10 次
- It's Enough: Relaxing Diagonal Constraints in Linear Autoencoders for RecommendationJaewan Moon, Hye-young Kim, Jongwuk LeeSIGIR 2023 · 被引用 3 次
- NUDGE: Lightweight Non-Parametric Fine-Tuning of Embeddings for RetrievalSepanta Zeighami, Zac Wellmer, Aditya G. ParameswaranICLR 2025
