BOSS: Bandwidth-Optimized Search Accelerator for Storage-Class Memory
Jun Heo, Seung Yul Lee, Sunhong Min, Yeonhong Park, Sungjun Jung, Tae Jun Ham, Jae W. Lee
摘要
Search is one of the most popular and important web services. The inverted index is the standard data structure adopted by most full-text search engines. Recently, custom hardware accelerators for inverted index search have emerged to demonstrate much higher throughput than the conventional CPU or GPU. However, less attention has been paid to addressing the memory capacity pressure with inverted index. The conventional DDRx DRAM memory system significantly increases the system cost to make a terabyte-scale main memory. Instead, a shared memory pool composed of storage-class memory (SCM) devices is a promising alternative for scaling memory capacity at a much lower cost. However, this SCM-based pooled memory poses new challenges caused by the limited bandwidth of both SCM devices and the shared interconnect to the host CPU. Thus, we propose BOSS, the first near-data processing (NDP) architecture for inverted index search on SCM-based pooled memory, which maintains high throughput of query processing in this bandwidth- constrained environment. BOSS mitigates the impact of low bandwidth of SCM devices by employing early-termination search algorithms, reducing the footprint of intermediate data, and introducing a programmable decompression module that can select the best compression scheme for a given inverted index. Furthermore, BOSS includes a top-k selection module in hardware to substantially reduce the host-accelerator bandwidth consumption. Compared to Apache Lucene, a production-grade search engine library, running on 8 CPU cores, BOSS achieves a geomean speedup of 8.1× on various complex query types, while reducing the average energy consumption by 189×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CXL-ANNS: Software-Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor SearchJunhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Seungjun Lee 等USENIX ATC 2023 · 被引用 75 次
- BAAP: Coupling Compute-in-SRAM with DRAM Banks for Near-Memory ProcessingCecilio C. Tamarit, Socrates S. Wong, Akshati Vaishnav, José F. MartínezISCA 2026
它引用的顶会 Paper3
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz 等FAST 2020 · 被引用 470 次
- IIU: Specialized Architecture for Inverted Index SearchJun Heo, Jaeyeon Won, Yejin Lee, Shivam Bharuka 等ASPLOS 2020 · 被引用 7 次
相关 Paper
- NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data ProcessingYitu Wang, Shiyu Li, Qilin Zheng, Linghao Song 等ISCA 2024 · 被引用 26 次
- UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing ArchitecturesTongxin Xie, Zhenhua Zhu, Bing Li, Yukai He 等HPCA 2025 · 被引用 9 次
- Physical vs. Logical Indexing with IDEA: Inverted Deduplication-Aware IndexAsaf Levi, Philip Shilane, Sarai Sheinvald, Gala YadgarFAST 2024 · 被引用 7 次
- Scalable Distributed Inverted List Indexes in Disaggregated MemoryManuel Widmoser, Daniel Kocher, Nikolaus AugstenSIGMOD 2024 · 被引用 5 次
- Columnar Formatted Inverted Index for Highly-Paralleled, Vectorized Query ProcessingWeichen Zhao, Minghao Zhao, Huiqi Hu, Weining QianICDE 2025
