A-Scan: Efficient Scale-Up Analytics via Throughput-Guided Data Movement
Hamish Nicholson, Aunn Raza, Viktor Sanca, Anastasia Ailamaki
摘要
Modern scale-up analytical systems, from single servers to rack-scale deployments, face increasing complexity in data movement across their storage hierarchies, NUMA nodes, and high-speed interconnects like CXL and PCIe. The performance gap between optimal and suboptimal data-movement strategies can be substantial, as the best approach depends on complex interactions between hardware capabilities (storage bandwidth, interconnect speeds, near-storage compute) and workload characteristics such as query selectivity, access patterns, and data volume. While existing systems employ datamovement strategies fixed at query planning time, the growing diversity of storage technologies, memory topologies, and analytical workloads makes such static approaches brittle, demanding adaptive runtime selection. We propose Throughput-Guided Data Movement that uses end-to-end operator-pipeline throughput as the adaptation signal to select how and where bytes are realized at query runtime. Throughput implicitly captures the interplay between hardware and query properties without an explicit performance model for data movement. We integrate throughput-guided data movement into a leaf operator, Adaptive Scan (A-Scan), that encapsulates data-movement within the operator and selects among them at runtime to maximize observed end-to-end throughput. A-Scan keeps the query plan fixed and presents the same rowset interface to its parent operator while internally controlling how and where data are realized across nodes and memories. We validate A-Scan by implementing it in a push-based, code-generating analytical engine, showing that it approaches the per-query best static choice across tested configurations and interference conditions, without hardware-specific tuning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Mosaic: A Budget-Conscious Storage Engine for Relational Database SystemsLukas Vogel, Alexander van Renen, Satoshi Imamura, Viktor Leis 等VLDB 2020
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast StorageJigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten BinnigVLDB 2026
- ADAMANT: A Query Executor with Plug-In Interfaces for Easy Co-processor IntegrationBala Gurumurthy, David Broneske, Gabriel Campero Durand, Thilo Pionteck 等ICDE 2023 · 被引用 1 次
- LAQy: Efficient and Reusable Query Approximations via Lazy SamplingViktor Sanca, Periklis Chrysogelos, Anastasia AilamakiSIGMOD 2023 · 被引用 5 次
- Tango: A Cross-layer Approach to Managing I/O Interference over Local Ephemeral StorageZhenbo Qiao, Qirui Tian, Zhenlu Qin, Jinzhen Wang 等SC 2024 · 被引用 2 次
