Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms
Herodotos Herodotou, Elena Kakoulli
Abstract
The recent advancements in storage technologies have popularized the use of tiered storage systems in data-intensive compute clusters. The Hadoop Distributed File System (HDFS), for example, now supports storing data in memory, SSDs, and HDDs, while OctopusFS and hatS offer fine-grained storage tiering solutions. However, the task schedulers of big data platforms (such as Hadoop and Spark) will assign tasks to available resources only based on data locality information, and completely ignore the fact that local data is now stored on a variety of storage media with different performance characteristics. This paper presents Trident, a principled task scheduling approach that is designed to make optimal task assignment decisions based on both locality and storage tier information. Trident formulates task scheduling as a minimum cost maximum matching problem in a bipartite graph and uses a standard solver for finding the optimal solution. In addition, Trident utilizes two novel pruning algorithms for bounding the size of the graph, while still guaranteeing optimality. Trident is implemented in both Spark and Hadoop, and evaluated extensively using a realistic workload derived from Facebook traces as well as an industry-validated benchmark, demonstrating significant benefits in terms of application performance and cluster efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04421842-7679-4475-addd-5135e3a9acdbCited by top-tier papers3
- Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data ProcessingChenghao Lyu, Qi Fan, Fei Song, Arnab Sinha et al.VLDB 2022 · 14 citations
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 9 citations
- S/C: Speeding up Data Materialization with Bounded MemoryZhaoheng Li, Xinyu Pi, Yongjoo ParkICDE 2023 · 7 citations
Builds on1
Related papers
- Adaptive Low-level Storage of Very Large Knowledge GraphsJacopo Urbani, Ceriel J. H. JacobsWWW 2020 · 10 citations
- TCO-driven Storage Provisioning for Exascale Data CentersTimothy Kim, Saurabh Kadekodi, Arif Merchant, Prashant Nema et al.EuroSys 2026
- Spark-based Cloud Data Analytics using Multi-Objective OptimizationFei Song, Khaled Zaouk, Chenghao Lyu, Arnab Sinha et al.ICDE 2021 · 15 citations
- Concealing Compression-accelerated I/O for HPC Applications through In Situ Task SchedulingSian Jin, Sheng Di, Frédéric Vivien, Daoce Wang et al.EuroSys 2024 · 13 citations
- Boosting Task Scheduling Data Locality with Low-latency, HW-accelerated Label PropagationLucas Morais, Juan Miguel De Haro Ruiz, Alfredo Goldman, Guido Araujo et al.MICRO 2025 · 1 citation
