Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic Manipulation
Yiyang Huang, Yuhui Hao, Bo Yu, Feng Yan, Yuxin Yang, Feng Min, Yinhe Han, Lin Ma, Shaoshan Liu, Qiang Liu, Yiming Gan
Abstract
Embodied AI robots have the potential to fundamentally improve the way human beings live and manufacture. Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks. In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis. Such an execution pipeline creates high latency and energy consumption. This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications. We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline. Instead of predicting action for one single frame, Corki predicts the trajectory for the near future to reduce the frequency of LLM inference. The algorithm is coupled with a hardware that accelerates transforming trajectory into actual torque signals used to control robots and an execution pipeline that parallels data * equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionYufeng Zhong, Chengjian Feng, Feng Yan, Fanfan Liu et al.ICCV 2025 · 1 citation
- CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI SystemsTong Xie, Yijiahao Qi, Jinqi Wen, Zishen Wan et al.ASPLOS 2026 · 1 citation
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang et al.NeurIPS 2023 · 453 citations
Related papers
- Enhancing LLM Planning for Robotics Manipulation through Hierarchical Procedural Knowledge GraphsJiacong Zhou, Jiaxu Miao, Xianyun Wang, Jun YuNeurIPS 2025 · 1 citation
- ReCA: Integrated Acceleration for Real-Time and Efficient Cooperative Embodied Autonomous AgentsZishen Wan, Yuhang Du, Mohamed Ibrahim, Jiayi Qian et al.ASPLOS 2025 · 7 citations
- Grounding LLMs in Scientific Discovery via Embodied ActionsBo Zhang, Jinfeng Zhou, Yuxuan Chen, Jianing Yin et al.ICML 2026
- BadRobot: Jailbreaking Embodied LLM Agents in the Physical WorldHangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou et al.ICLR 2025
- R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action PlannerZiyi Bai, Hanxuan Li, Bin Fu, Chuyan Xiong et al.CVPR 2025
