Accelerator Polymorphism: Transcending Domain-Specific Architectures with Robotics
Hanyang Xu, Seongryong Oh, Yubin Lee, Ashwin Rohit Alagiri Rajan, Rohan Mahapatra, Om Patil, Yuchuan Li, Jongse Park, Hadi Esmaeilzadeh
摘要
Robotics is poised for a ChatGPT-like leap, but progress is constrained by the lack of a compute substrate suited to its cross-domain computational demands. Unlike LLMs and neural networks, which are largely confined to a single domain of tensor operations, robotics spans multiple algorithmic domains, including graph search, control theory, optimization, and more, often within a single application pipeline. This cross-domain diversity, coupled with stringent power, latency, and cost constraints, demands a new design point-one that transcends the siloed boundaries of domain-specific accelerators without reintroducing the overheads of general-purpose processing. To address this challenge, we propose a novel architectural concept: accelerator polymorphism, whereby a single physical substrate dynamically shape-shifts itself from one specialized execution model to another without sacrificing the efficiency of specialization by reverting to general-purpose design. We instantiate this concept in a concrete microarchitecture, Morphatron, that dynamically shape-shifts across five execution morphas-queuecentric SIMD, graph, tree, vector, systolic-each tailored to distinct algorithmic domains in robotics. At the heart of Morphatron is Morpha Core, a novel SIMD processor specialized for queuecentric execution and built on a key algorithmic insight: queues are a unifying data structure across robotic workloads. Morpha Core elevates queues to first-class SIMD operands, integrates memory allocation directly into the pipeline as an architectural primitive, and thereby enables hardware acceleration of dynamically-sized data structures. Morphatron is a multi-core accelerator composed of a two-dimensional array of Morpha Cores, designed to support the diverse access and parallelism patterns revealed by our indepth study of end-to-end robotic workloads. We evaluate the performance and power efficiency of Morphatron using six end-to-end robotic applications from the RoWild benchmarks [1]. Compared to the NVIDIA Jetson Orin Nano GPU, Morphatron achieves a 5.5× speedup and a 6.6× improvement in performance-per-watt, on average. Relative to the ARM A78 CPU, Morphatron delivers a 7.7× speedup and a 7.7× improvement in performance-per-watt.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- A Computational Stack for Cross-Domain AccelerationSean Kinzer, Joon Kyung Kim, Soroush Ghodrati, Brahmendra Reddy Yatham 等HPCA 2021 · 被引用 9 次
- Robomorphic computing: a design methodology for domain-specific accelerators parameterized by robot morphologySabrina M. Neuman, Brian Plancher, Thomas Bourgeat, Thierry Tambe 等ASPLOS 2021 · 被引用 43 次
- MetaMorph: Learning Universal Controllers with TransformersAgrim Gupta, Linxi Fan, Surya Ganguli, Li Fei-FeiICLR 2022 · 被引用 130 次
- PolymorPIC: Embedding Polymorphic Processing-in-Cache in RISC-V based Processor for Full-stack Efficient AI InferenceCheng Zou, Ziling Wei, Jun Yan Lee, Chen Nie 等MICRO 2025 · 被引用 4 次
- Data Motion Acceleration: Chaining Cross-Domain Multi AcceleratorsShu-Ting Wang, Hanyang Xu, Amin Mamandipoor, Rohan Mahapatra 等HPCA 2024 · 被引用 10 次
