Energy-Efficient Large-Scale Vector Similarity Search in NAND-Flash via Hybrid Matching
Chih-Yu Hu, Chi-Tse Huang, Hao-Wei Chiang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Andy Wu
摘要
Vector similarity search (VSS) is crucial in many AI applications, such as few-shot learning (FSL) and approximate nearest-neighbor search (ANNS), but it demands significant memory capacity and incurs substantial energy costs for data transfers during large-scale comparisons. Various in-memory search technologies have been developed to improve energy efficiency, with NAND-based multi-bit content-addressable memory (MCAM) standing out as a promising solution for its high density and large capacity. MCAM can operate in exact-search (ES) mode, supporting only perfect matches with low energy cost, or in approximate-search (AS) mode, enabling flexible VSS. However, AS mode incurs significant energy waste when comparing queries with non-target stored vectors. To address this issue, we propose Hybrid-M, a 3D NAND-based in-memory VSS architecture that integrates both modes into a single hybrid matching process, using ES mode as a filter to reduce redundant searches for AS mode. We apply three techniques to optimize this integration: range encoding for multi-level cells (MLC) to enhance filtering, search voltage shifts to mitigate the impact on AS accuracy and reduce matching currents, and a filtering-aware training method to further improve reliability and energy efficiency. Results show that Hybrid-M achieves comparable accuracy while reducing energy consumption by 67% to 83% compared to MACM-based VSS using only AS mode, across various many-class FSL and ANNS workloads.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Segmented Angular Pre-Processing for Accurate and Efficient In-Memory Vector Similarity SearchChi-Tse Huang, Jen-Chieh Wang, Hsiang-Yun Cheng, An-Yeu Andy WuDAC 2025
- ICE: An Intelligent Cognition Engine with 3D NAND-based In-Memory Computing for Vector Similarity Search AccelerationHan-Wen Hu, Wei-Chen Wang, Yuan-Hao Chang, Yung-Chun Lee 等MICRO 2022 · 被引用 27 次
- Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceQilin Zheng, Ziru Li, Jonathan Ku, Yitu Wang 等DAC 2024 · 被引用 2 次
- MIRACLE: Multimodal Information Retrieval via a Combined In-Memory Processing and Content Addressable Memory ApproachXuehui Liu, Xueyan Wang, Tianyang Yu, Chen Cheng 等DAC 2025 · 被引用 1 次
- Compact and Efficient CAM Architecture through Combinatorial Encoding and Self-Terminating Searching for In-Memory-Searching AcceleratorWeikai Xu, Jin Luo, Qianqian Huang, Ru HuangDAC 2024 · 被引用 1 次
