Zero-Change Object Transmission for Distributed Big Data Analytics
Mingyu Wu, Shuaiwei Wang, Haibo Chen, Binyu Zang
Abstract
Distributed big-data analytics heavily rely on high-level languages like Java and Scala for their reliability and versatility. However, those high-level languages also create obstacles for data exchange. To transfer data across managed runtimes like Java Virtual Machines (JVMs), objects should be transformed into byte arrays by the sender (serialization) and transformed back into objects by the receiver (deserialization). The object serialization and deserialization (OSD) phase introduces considerable performance overhead. Prior efforts mainly focus on optimizing some phases in OSD, so object transformation is still inevitable. Furthermore, they require extra programming efforts to integrate with existing applications, and their transformation also leads to duplicated object transmission. This work proposes Zero-Change Object Transmission (ZCOT), where objects are directly copied among JVMs without any transformations. ZCOT can be used in existing applications with minimal effort, and its object-based transmission can be used for deduplication. The evaluation on state-of-the-art data analytics frameworks indicates that ZCOT can greatly boost the performance of data exchange and thus improve the application performance by up to 23.6%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Serialization/Deserialization-free State Transfer in Serverless WorkflowsFangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen et al.EuroSys 2024 · 31 citations
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li et al.HPCA 2025 · 6 citations
- DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data ProcessingChen Ding, Sicen Li, Kai Lu, Ting Yao et al.USENIX ATC 2025 · 2 citations
Builds on5
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout et al.EuroSys 2020 · 163 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
- A Specialized Architecture for Object Serialization with Applications to Big Data AnalyticsJaeyoung Jang, Sungjun Jung, Sunmin Jeong, Jun Heo et al.ISCA 2020 · 32 citations
- Naos: Serialization-free RDMA networking in JavaKonstantin Taranov, Rodrigo Bruno, Gustavo Alonso, Torsten HoeflerUSENIX ATC 2021 · 31 citations
- Semeru: A Memory-Disaggregated Managed RuntimeChenxi Wang, Haoran Ma, Shi Liu, Yuanqi Li et al.OSDI 2020 · 6 citations
Related papers
- JOSer: Just-In-Time Object Serialization for Heavy Java Serialization WorkloadsChaokun Yang, Pengbo Nie, Ziyi Lin, Weipeng Wang et al.ASPLOS 2026
- HeapBuffers: Why Not Just Using a Binary Serialization Format for Your Managed Memory?Daniele Bonetta, Júnior Löff, Matteo Basso, Walter BinderOOPSLA 2025 · 2 citations
- TeraHeap: Reducing Memory Pressure in Managed Big Data FrameworksIacovos G. Kolokasis, Giannos Evdorou, Shoaib Akram, Christos Kozanitis et al.ASPLOS 2023 · 15 citations
- Translation of Array-Based Loops to Distributed Data-Parallel ProgramsLeonidas Fegaras, Md Hasanuzzaman NoorVLDB 2020 · 13 citations
- Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into SparkBowen Yu, Guanyu Feng, Huanqi Cao, Xiaohan Li et al.VLDB 2022 · 3 citations
