NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
Hanchen Yang, Zishen Wan, Ritik Raj, Joongun Park, Ziwei Li, Ananda Samajdar, Arijit Raychowdhury, Tushar Krishna
Abstract
Neuro-Symbolic AI (NSAI) is an emerging paradigm that integrates neural networks with symbolic reasoning to enhance the transparency, reasoning capabilities, and data efficiency of AI systems. Recent NSAI systems have gained traction due to their exceptional performance in reasoning tasks and human-AI collaborative scenarios. Despite these algorithmic advancements, executing NSAI tasks on existing hardware (e.g., CPUs, GPUs, TPUs) remains challenging, due to their heterogeneous computing kernels, high memory intensity, and unique memory access patterns. Moreover, current NSAI algorithms exhibit significant variation in operation types and scales, making them incompatible with existing ML accelerators. These challenges highlight the need for a versatile and flexible acceleration framework tailored to NSAI workloads. In this paper, we propose NSFlow, an FPGA-based acceleration framework designed to achieve high efficiency, scalability, and versatility across NSAI systems. NSFlow features a design architecture generator that identifies workload data dependencies and creates optimized dataflow architectures, as well as a reconfigurable array with flexible compute units, re-organizable memory, and mixed-precision capabilities. Evaluating across NSAI workloads, NSFlow achieves speedup over Jetson TX2, more than over GPU, speedup over TPU-like systolic array, and more than over Xilinx DPU. NSFlow also demonstrates enhanced scalability, with only runtime increase when symbolic workloads scale by . To the best of our knowledge, NSFlow is the first framework to enable real-time generalizable NSAI algorithms acceleration, demonstrating a promising solution for next-generation cognitive systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6750f16-c44f-4bff-9fcc-59fa82b1c68cCited by top-tier papers1
Ask how each one uses itBuilds on6
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- Stratified Rule-Aware Network for Abstract Visual ReasoningSheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei et al.AAAI 2021 · 126 citations
- MIMONets: Multiple-Input-Multiple-Output Neural Networks Exploiting Computation in SuperpositionNicolas Menet, Michael Hersche, Geethan Karunaratne, Luca Benini et al.NeurIPS 2023 · 32 citations
- FALCON: Fast Visual Concept Learning by Integrating Images, Linguistic descriptions, and Conceptual RelationsLingjie Mei, Jiayuan Mao, Ziqi Wang, Chuang Gan et al.ICLR 2022 · 28 citations
- CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-DesignZishen Wan, Hanchen Yang, Ritik Raj, Che-Kai Liu et al.HPCA 2025 · 6 citations
Related papers
- Lobster: A GPU-Accelerated Framework for Neurosymbolic ProgrammingPaul Biberstein, Ziyang Li, Joseph Devietti, Mayur NaikASPLOS 2026 · 1 citation
- DOLPHIN: A Programmable Framework for Scalable Neurosymbolic LearningAaditya Naik, Jason Liu, Claire Wang, Amish Sethi et al.ICML 2025
- KLay: Accelerating Arithmetic Circuits for Neurosymbolic AIJaron Maene, Vincent Derkinderen, Pedro Zuidberg Dos MartiresICLR 2025
- CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM ArchitecturesYingjie Qi, Jianlei Yang, Yiou Wang, Yikun Wang et al.DAC 2025 · 2 citations
- TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAsNeha Prakriya, Yuze Chi, Suhail Basalama, Linghao Song et al.ASPLOS 2024 · 6 citations
