Multi-Dimensional ML-Pipeline Optimization in Cost-Effective Disaggregated Datacenter
Pingyi Huo, Anusha Devulapally, Hasan Al Maruf, Nandhini Chandramoorthy, Meena Arunachalam, Gulsum Gudukbay Akbulut, Mahmut T. Kandemir, Vijaykrishnan Narayanan
Abstract
Machine learning (ML) pipelines deployed in datacenters are becoming increasingly complex and resource intensive, requiring careful optimizations to meet performance and latency requirements.Deployment in NUMA architectures with heterogeneous memory types such as CXL introduces a vast configuration space for memory and thread management.However, existing methods often need help to efficiently navigate this space, relying on significant manual tuning and needing more adaptability across diverse hardware platforms and extensive user-side code modification.This paper presents and experimentally evaluates an adaptive auto-tuning framework that optimizes memory configurations to maximize system throughput while adhering to Service Level Agreements (SLAs) on latency.Our optimization is carried out in two phases, addressing both performance and power efficiency.The proposed framework integrates an extended Berkeley Packet Filter (eBPF)-based kernel module for real-time performance monitoring and a user-space optimization core leveraging Bayesian Optimization and Pareto Optimality to explore high-dimensional configuration spaces.Extensive experimental evaluations on various workloads demonstrate that our framework achieves up to a 48% increase in throughput compared to default NUMA, TPP, Caption, and GPU while reducing search costs by as much as 77%.Furthermore, our two-phase optimization approach incorporates power efficiency, achieving up to 14.3% power savings without compromising throughput.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6b3ee476-83af-4211-aff5-2df24ba30184Related papers
- xCCLTuner: Treating xCCL as Black-Box and Automatically TuningChenxu Wang, Wentao Fan, Zhehao Lin, Peirui Cao et al.INFOCOM 2026
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner et al.ASPLOS 2023 · 255 citations
- cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node CommunicationsXi Wang, Bin Ma, Jongryool Kim, Byungil Koh et al.SC 2025 · 5 citations
- SIDLE: Tree-structure Aware Indexes for CXL-based Heterogeneous MemoryHaoru Zhao, Mingkai Dong, Fangnuo Wu, Haibo ChenVLDB 2026 · 1 citation
- The Holon Approach for Simultaneously Tuning Multiple Components in a Self-Driving Database Management System with Machine Learning via Synthesized Proto-ActionsWilliam Zhang, Wan Shen Lim, Matthew Butrovich, Andrew PavloVLDB 2024 · 13 citations
