ASMCap: An Approximate String Matching Accelerator for Genome Sequence Analysis Based on Capacitive Content Addressable Memory
Hongtao Zhong, Zhonghao Chen, Wenqin Huangfu, Chen Wang, Yixin Xu, Tianyi Wang, Yao Yu, Yongpan Liu, Vijaykrishnan Narayanan, Huazhong Yang, Xueqing Li
Abstract
Genome sequence analysis is a powerful tool in medical and scientific research. Considering the inevitable sequencing errors and genetic variations, approximate string matching (ASM) has been adopted in practice for genome sequencing. However, with exponentially increasing bio-data, ASM hardware acceleration is facing severe challenges in improving the throughput and energy efficiency with the accuracy constraint.
This paper presents ASMCap, an ASM acceleration approach for genome sequence analysis with hardware-algorithm cooptimization. At the circuit level, ASMCap adopts charge-domain computing based on the capacitive multi-level content addressable memories (ML-CAMs), and outperforms the state-of-the-art ML-CAM-based ASM accelerators EDAM with higher accuracy and energy efficiency. ASMCap also has misjudgment correction capability with two proposed hardware-friendly strategies, namely the Hamming-Distance Aid Correction (HDAC) for the substitution-dominant edits and the Threshold-Aware Sequence Rotation (TASR) for the consecutive indels. Evaluation results show that ASMCap can achieve an average of 1.2x (from 74.7% to 87.6%) and up to 1.8x (from 46.3% to 81.2%) higher F1 score (the key metric of accuracy), 1.4x speedup, and 10.8x energy efficiency improvement compared with EDAM. Compared with the other ASM accelerators, including ResMA based on the comparison matrix, and SaVI based on the seeding strategy, ASMCap achieves an average improvement of 174x and 61x speedup, and 8.7e3x and 943x higher energy efficiency, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- EDAM: edit distance tolerant approximate matching content addressable memoryRobert Hanhan, Esteban Garzón, Zuher Jahshan, Adam Teman et al.ISCA 2022 · 37 citations
- GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence AnalysisDamla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina et al.MICRO 2020 · 23 citations
- ReSMA: accelerating approximate string matching using ReRAM-based content addressable memoryHuize Li, Hai Jin, Long Zheng, Yu Huang et al.DAC 2022 · 7 citations
Related papers
- CASA: An Energy-Efficient and High-Speed CAM-based SMEM Seeding Accelerator for Genome AlignmentYi Huang, Lingkun Kong, Dibei Chen, Zhiyu Chen et al.MICRO 2023 · 5 citations
- Cognitive Correlative Encoding for Genome Sequence Matching in Hyperdimensional SystemPrathyush Poduval, Zhuowen Zou, Xunzhao Yin, Elaheh Sadredini et al.DAC 2021 · 39 citations
- GenStore: a high-performance in-storage processing system for genome sequence analysisNika Mansouri-Ghiasi, Jisung Park, Harun Mustafa, Jeremie S. Kim et al.ASPLOS 2022 · 74 citations
- QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis AlgorithmsJulian Pavon, Iván Vargas Valdivieso, Carlos Rojas, César Hernández et al.ISCA 2024 · 5 citations
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim et al.ISCA 2022 · 66 citations
