AWARE: Workload-aware, Redundancy-exploiting Linear Algebra
Sebastian Baunsgaard, Matthias Boehm
Abstract
Compression is an effective technique for fitting data in available memory, reducing I/O, and increasing instruction parallelism. While data systems primarily rely on lossless compression, modern machine learning (ML) systems exploit the approximate nature of ML and mostly use lossy compression via low-precision floating- or fixed-point representations. The resulting unknown impact on learning progress, and model accuracy, however, create trust concerns, that require trial and error, and are problematic for declarative ML pipelines. Given the trend towards increasingly complex, composite ML pipelines---with outer loops for hyper-parameter tuning, feature selection, and data cleaning/augmentation---it is hard for a user to infer the impact of lossy compression. Sparsity exploitation is a common lossless scheme used to improve performance without this uncertainty. Evolving this concept to general redundancy-exploiting compression is a natural next step. Existing work on lossless compression and compressed linear algebra (CLA) enable such exploitation to a degree, but face challenges for general applicability. In this paper, we address these limitations with a workload-aware compression framework, comprising a broad spectrum of new compression schemes and kernels. Instead of a data-centric approach that optimizes compression ratios, our workload-aware compression summarizes the workload of an ML pipeline, and optimizes the compression and execution plan to minimize execution time. On various micro benchmarks and end-to-end ML pipelines, we observe improvements for individual operations up to 10,000x and ML algorithms up to νmprint6.6 x compared to uncompressed operations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Galley: Modern Query Optimization for Sparse Tensor ProgramsKyle Deeds, Willow Ahrens, Magdalena Balazinska, Dan SuciuSIGMOD 2025 · 3 citations
- Morphing-based Compression for Data-centric ML PipelinesSebastian Baunsgaard, Matthias BoehmVLDB 2026
- stratum: A System Infrastructure for Massive Agent-Centric ML WorkloadsArnab Phani, Elias Strauss, Sebastian SchelterVLDB 2026
- Reducing the GPU Memory Bottleneck with Lossless Compression for MLAditya K. Kamath, Arvind Krishnamurthy, Marco Canini, Simon PeterEuroSys 2026
Builds on16
- Approximate Analytics System over Compressed Time Series with Tight Deterministic Error GuaranteesChunbin Lin, Etienne Boursier, Yannis PapakonstantinouVLDB 2020 · 505 citations
- Pruning Neural Networks at Initialization: Why Are We Missing the Mark?Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICLR 2021 · 261 citations
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training DataFelipe Petroski Such, Aditya Rawal, Joel Lehman, Kenneth O. Stanley et al.ICML 2020 · 180 citations
- Combiner: Full Attention Transformer with Sparse Computation CostHongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang et al.NeurIPS 2021 · 99 citations
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
Related papers
- A Portable, Fast, DCT-based Compressor for AI AcceleratorsMilan Shah, Xiaodong Yu, Sheng Di, Michela Becchi et al.HPDC 2024 · 4 citations
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 8 citations
- SpZip: Architectural Support for Effective Data Compression In Irregular ApplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2021 · 31 citations
- Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to AskZhaoyuan Su, Ammar Ahmed, Zirui Wang, Ali Anwar et al.VLDB 2024
- Improving Matrix-vector Multiplication via Lossless Grammar-Compressed MatricesPaolo Ferragina, Giovanni Manzini, Travis Gagie, Dominik Köppl et al.VLDB 2022 · 17 citations
