Spangle: A Distributed In-Memory Processing System for Large-Scale Arrays
Sangchul Kim, Bogyeong Kim, Bongki Moon
摘要
With increasing volumes of scientific data, a scalable and parallel computing framework is required for scientific analysis in computer simulations and experiments. Scientific data are commonly generated in multi-dimensional arrays, and the array data model is appropriate to store them for analysis, including for data mining and arithmetic computation. In this paper, we introduce an array processing system called Spangle. It is implemented on top of Apache Spark, a popular map-reduce framework for complex computation workloads. To support array data computation, we extended Resilient Distributed Dataset (RDD) based on the array data model named ArrayRDD. ArrayRDD is an inherently parallel data structure that provides fault-tolerance. In addition, by adopting the array data model, Spangle provides an interface for expressing machine learning algorithms, which heavily rely on linear algebra. We tailored two popular algorithms, PageRank and Stochastic Gradient Descent, for large-scale datasets in Spangle.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Translation of Array-Based Loops to Distributed Data-Parallel ProgramsLeonidas Fegaras, Md Hasanuzzaman NoorVLDB 2020 · 被引用 13 次
- Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear AlgebraShangyu Luo, Dimitrije Jankov, Binhang Yuan, Chris JermaineSIGMOD 2021 · 被引用 9 次
- Scalable Robust Graph Embedding with SparkChi Thang Duong, Dung Hoang, Hongzhi Yin, Matthias Weidlich 等VLDB 2022 · 被引用 3 次
- Array-based Data Management for GenomicsOlha Horlova, Abdulrahman Kaitoua, Stefano CeriICDE 2020 · 被引用 3 次
- SpDISTAL: Compiling Distributed Sparse Tensor ComputationsRohan Yadav, Alex Aiken, Fredrik KjolstadSC 2022 · 被引用 7 次
