Noema: Hardware-Efficient Template Matching for Neural Population Pattern Detection
Ameer M. S. Abdelhadi, Eugene Sha, Ciaran Bannon, Hendrik Steenland, Andreas Moshovos
Abstract
Repeating patterns of activity across neurons is thought to be key to understanding how the brain represents, reacts, and learns. Advances in imaging and electrophysiology allow us to observe activities of groups of neurons in real-time, with ever increasing detail. Detecting patterns over these activity streams is an effective means to explore the brain, and to detect memories, decisions, and perceptions in real-time while driving effectors such as robotic arms, or augmenting and repairing brain function. Template matching is a popular algorithm for detecting recurring patterns in neural populations and has primarily been implemented on commodity systems. Unfortunately, template matching is memory intensive and computationally expensive. This has prevented its use in portable applications, such as neuroprosthetics, which are constrained by latency, form-factor, and energy. We present Noema a dedicated template matching hardware accelerator that overcomes these limitations. Noema is designed to overcome the key bottlenecks of existing implementations: binning that converts the incoming bit-serial neuron activity streams into a stream of aggregate counts, memory storage and traffic for the templates and the binned stream, and the extensive use of floating-point arithmetic. The key innovation in Noema is a reformulation of template matching that enables computations to proceed progressively as data is received without binning while generating numerically identical results. This drastically reduces latency when most computations can now use simple, area- and energy efficient bit- and integer-arithmetic units. Furthermore, Noema implements template encoding to greatly reduce template memory storage and traffic. Noema is a hierarchical and scalable design where the bulk of its units are low-cost and can be readily replicated and their frequency can be adjusted to meet a variety of energy, area, and computation constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SCALO: An Accelerator-Rich Distributed System for Scalable Brain-Computer InterfacingKarthik Sriram, Raghavendra Pradyumna Pothukuchi, Michal Gerasimiuk, Muhammed Ugur et al.ISCA 2023 · 18 citations
- MINDFUL: Safe, Implantable, Large-Scale Brain-Computer Interfaces from a System-Level Design PerspectiveGuy Eichler, Yatin Gilhotra, Nanyu Zeng, Martha A. Kim et al.MICRO 2025 · 3 citations
- IDEA-GP: Instruction-Driven Architecture with Efficient Online Workload Allocation for Geometric PerceptionSuquan Zhang, Yu Hu, Yunfei Xiang, Dawei Zhao et al.ISCA 2025
Related papers
- Marple: Scalable Spike Sorting for Untethered Brain-Machine InterfacingEugene Sha, Andy Wei Liu, Kareem Ibrahim, Mostafa Mahmoud et al.ASPLOS 2024 · 6 citations
- FlexAmata: A Universal and Efficient Adaption of Applications to Spatial Automata Processing AcceleratorsElaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea Stan et al.ASPLOS 2020 · 23 citations
- Software-hardware codesign for efficient in-memory regular pattern matchingLingkun Kong, Qixuan Yu, Agnishom Chattopadhyay, Alexis Le Glaunec et al.PLDI 2022 · 23 citations
- An Energy-Efficient Kalman Filter Architecture with Tunable Accuracy for Brain-Computer InterfacesGuy Eichler, Joseph Zuckerman, Luca P. CarloniDAC 2025 · 1 citation
- RAP: Reconfigurable Automata ProcessorZiyuan Wen, Alexis Le Glaunec, Konstantinos Mamouras, Kaiyuan YangISCA 2025 · 2 citations
