Continuously Bulk Loading over Range Partitioned Tables for Large Scale Historical Data
Xiaolong He, Peng Cai, Xuan Zhou, Aoying Zhou
Abstract
To support efficiently and continuously loading large scale historical data into a distributed data management system (DDMS), it needs to balance the bulk workload across machines. The fundamental problem is to estimate the time used to merge currently loaded data (defined as incremental data) into previously loaded data (defined as baseline data) for each partition, referred to as partition merge. In this work, we present a learning-based framework, referred to as LeaBalancer, to balance the merge loads across cluster nodes. In the situation where the system is scheduled to have regular bulk loading tasks, LeaBalancer can learn to predict the partition merge time from the merge logs generated by previous bulk loadings. Nevertheless, it is still difficult to balance the bulk workload only using a single plan phase because of inaccurate merge time prediction or other in-progress heavy workloads during the bulk loading. To resolve this problem, we design a multi-round balancing strategy, and at the beginning of each round LeaBalancer carefully chooses partitions for migration according to the remaining merge loads in each node. Experimental results show that LeaBalancer can adaptively perform load balance under various settings.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b5833b35-b73a-4ecd-96fb-9a2a5a1d03b0Related papers
- Lunule: an agile and judicious metadata load balancer for CephFSYiduo Wang, Cheng Li, Xinyang Shao, Youxu Chen et al.SC 2021 · 18 citations
- Promi: Progressive Live Migration in Distributed Database SystemsZhenghao Ding, Xinyi Zhang, Wei Lu, Wenlong Ma et al.ICDE 2025
- Optimal Online Balanced Graph PartitioningMaciej Pacut, Mahmoud Parham, Stefan SchmidINFOCOM 2021 · 8 citations
- Partitioner Selection with EASE to Optimize Distributed Graph ProcessingNikolai Merkel, Ruben Mayer, Tawkir Ahmed Fakir, Hans-Arno JacobsenICDE 2023 · 3 citations
- Robust Load Balancing with Machine Learned AdviceSara Ahmadian, Hossein Esfandiari, Vahab S. Mirrokni, Binghui PengSODA 2022 · 5 citations
