Optimus Prime: Accelerating Data Transformation in Servers
Arash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland, Zilu Tian, Mario Paulo Drumond, Babak Falsafi, Christoph Koch
Abstract
Modern online services are shifting away from monolithic applications to loosely-coupled microservices because of their improved scalability, reliability, programmability and development velocity. Microservices communicating over the datacenter network require data transformation (DT) to convert messages back and forth between their internal formats. This work identifies DT as a bottleneck due to reductions in latency of the surrounding system components, namely application runtimes, protocol stacks, and network hardware. We therefore propose Optimus Prime (OP), a programmable DT accelerator that uses a novel abstraction, an in-memory schema, to represent DT operations. The schema is compatible with today's DT frameworks and enables any compliant accelerator to perform the transformations comprising a request in parallel. Our evaluation shows that OP's DT throughput matches the line rate of today's NICs and has 60x higher throughput compared to software, at a tiny fraction of the CPU's silicon area and power. We also evaluate a set of microservices running on Thrift, and show up to 30% reduction in service latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6007d024-c609-40c8-a43a-ae382e11afc5Cited by top-tier papers26
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 218 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao et al.MICRO 2021 · 42 citations
- Serialization/Deserialization-free State Transfer in Serverless WorkflowsFangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen et al.EuroSys 2024 · 31 citations
- Naos: Serialization-free RDMA networking in JavaKonstantin Taranov, Rodrigo Bruno, Gustavo Alonso, Torsten HoeflerUSENIX ATC 2021 · 31 citations
Related papers
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 24 citations
- AccelFlow: Orchestrating an On-Package Ensemble of Fine-Grained Accelerators for MicroservicesJovan Stojkovic, Abraham Farrell, Zhangxiaowen Gong, Christopher J. Hughes et al.HPCA 2026 · 2 citations
- The NEBULA RPC-Optimized ArchitectureMark Sutherland, Siddharth Gupta, Babak Falsafi, Virendra J. Marathe et al.ISCA 2020 · 29 citations
- Dagger: efficient and fast RPCs in cloud microservices with near-memory reconfigurable NICsNikita Lazarev, Shaojie Xiang, Neil Adit, Zhiru Zhang et al.ASPLOS 2021 · 54 citations
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo et al.EuroSys 2024 · 29 citations
