Memory safe computations with XLA compiler
Artem Artemev, Yuze An, Tilman Roeder, Mark van der Wilk
Abstract
Software packages like TensorFlow and PyTorch are designed to support linear algebra operations, and their speed and usability determine their success. However, by prioritising speed, they often neglect memory requirements. As a consequence, the implementations of memory-intensive algorithms that are convenient in terms of software design can often not be run for large problems due to memory overflows. Memoryefficient solutions require complex programming approaches with significant logic outside the computational framework. This impairs the adoption and use of such algorithms. To address this, we developed an XLA compiler extension 1 that adjusts the computational data-flow representation of an algorithm according to a user-specified memory limit. We show that k-nearest neighbour and sparse Gaussian process regression methods can be run at a much larger scale on a single device, where standard implementations would have failed. Our approach leads to better use of hardware resources. We believe that further focus on removing memory constraints at a compiler level will widen the range of machine learning methods that can be developed in the future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa626fca-6d12-431e-9a89-a0276884d956Cited by top-tier papers2
- A differentiable brain simulator bridging brain simulation and brain-inspired computingChaoming Wang, Tianqiu Zhang, Sichao He, Hongyaoxing Gu et al.ICLR 2024 · 8 citations
- (A)iSpy: Parasitic Trojans for Machine Learning InfrastructureHabibur Rahaman, Qipan Xu, Zafaryab Haider, Prabuddha Chakraborty et al.CCS 2026
Builds on4
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee et al.ICLR 2020 · 254 citations
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
- Fast geometric learning with symbolic matricesJean Feydy, Joan Alexis Glaunès, Benjamin Charlier, Michael M. BronsteinNeurIPS 2020 · 53 citations
- Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate GradientsArtem Artemev, David R. Burt, Mark van der WilkICML 2021 · 28 citations
Related papers
- Compilation of Modular and General Sparse WorkspacesGenghan Zhang, Olivia Hsu, Fredrik KjolstadPLDI 2024 · 6 citations
- Compiler Support for Sparse Tensor ConvolutionsPeiming Liu, Alexander J. Root, Anlun Xu, Yinying Li et al.OOPSLA 2024 · 5 citations
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
- AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesZhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long et al.ASPLOS 2022 · 78 citations
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 69 citations
