Noema: Hardware-Efficient Template Matching for Neural Population Pattern Detection
Ameer M. S. Abdelhadi, Eugene Sha, Ciaran Bannon, Hendrik Steenland, Andreas Moshovos
摘要
Repeating patterns of activity across neurons is thought to be key to understanding how the brain represents, reacts, and learns. Advances in imaging and electrophysiology allow us to observe activities of groups of neurons in real-time, with ever increasing detail. Detecting patterns over these activity streams is an effective means to explore the brain, and to detect memories, decisions, and perceptions in real-time while driving effectors such as robotic arms, or augmenting and repairing brain function. Template matching is a popular algorithm for detecting recurring patterns in neural populations and has primarily been implemented on commodity systems. Unfortunately, template matching is memory intensive and computationally expensive. This has prevented its use in portable applications, such as neuroprosthetics, which are constrained by latency, form-factor, and energy. We present Noema a dedicated template matching hardware accelerator that overcomes these limitations. Noema is designed to overcome the key bottlenecks of existing implementations: binning that converts the incoming bit-serial neuron activity streams into a stream of aggregate counts, memory storage and traffic for the templates and the binned stream, and the extensive use of floating-point arithmetic. The key innovation in Noema is a reformulation of template matching that enables computations to proceed progressively as data is received without binning while generating numerically identical results. This drastically reduces latency when most computations can now use simple, area- and energy efficient bit- and integer-arithmetic units. Furthermore, Noema implements template encoding to greatly reduce template memory storage and traffic. Noema is a hierarchical and scalable design where the bulk of its units are low-cost and can be readily replicated and their frequency can be adjusted to meet a variety of energy, area, and computation constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SCALO: An Accelerator-Rich Distributed System for Scalable Brain-Computer InterfacingKarthik Sriram, Raghavendra Pradyumna Pothukuchi, Michal Gerasimiuk, Muhammed Ugur 等ISCA 2023 · 被引用 18 次
- MINDFUL: Safe, Implantable, Large-Scale Brain-Computer Interfaces from a System-Level Design PerspectiveGuy Eichler, Yatin Gilhotra, Nanyu Zeng, Martha A. Kim 等MICRO 2025 · 被引用 3 次
- IDEA-GP: Instruction-Driven Architecture with Efficient Online Workload Allocation for Geometric PerceptionSuquan Zhang, Yu Hu, Yunfei Xiang, Dawei Zhao 等ISCA 2025
相关 Paper
- Marple: Scalable Spike Sorting for Untethered Brain-Machine InterfacingEugene Sha, Andy Wei Liu, Kareem Ibrahim, Mostafa Mahmoud 等ASPLOS 2024 · 被引用 6 次
- FlexAmata: A Universal and Efficient Adaption of Applications to Spatial Automata Processing AcceleratorsElaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea Stan 等ASPLOS 2020 · 被引用 23 次
- Software-hardware codesign for efficient in-memory regular pattern matchingLingkun Kong, Qixuan Yu, Agnishom Chattopadhyay, Alexis Le Glaunec 等PLDI 2022 · 被引用 23 次
- An Energy-Efficient Kalman Filter Architecture with Tunable Accuracy for Brain-Computer InterfacesGuy Eichler, Joseph Zuckerman, Luca P. CarloniDAC 2025 · 被引用 1 次
- RAP: Reconfigurable Automata ProcessorZiyuan Wen, Alexis Le Glaunec, Konstantinos Mamouras, Kaiyuan YangISCA 2025 · 被引用 2 次
