METAL: Caching Multi-level Indexes in Domain-Specific Architectures
Anagha Molakalmur Anil Kumar, Aditya Prasanna, Jonathan Balkind, Arrvindh Shriraman
Abstract
State-of-the-art domain specific architectures (DSAs) work with sparse data, and need hardware support for index datastructures [31,43,57,61]. Indexes are more space-efficient for sparse-data, and reduce DRAM bandwidth, if data reuse can be managed. However, indexes exhibit dynamic accesses, chase pointers, and need to walk-and-search. This inflates the working set and thrashes the cache. We observe that the cache organization itself is responsible for this behavior.
We develop METAL, a portable caching idiom that enables DSAs to employ index data-structures. METAL decouples reuse of the index metadata from data reuse, and optimizes it independently. We propose two ideas: i) IX-Cache: A cache that leverages range tags to short-circuits index walks, and reduces the working set. IX-cache helps capture the trade-off between wider index nodes that maximize reach vs those that are closer to leaf and minimize walk latency. ii) Reuse Patterns: An interface to explicitly manage the cache. Patterns orchestrate cache insertions and bypass as we dynamically traverse different index regions. METAL improves performance vs. streaming DSAs by 7.8×, address-caches by 4.1×, and state-of-the-art DSA-cache [50] by 2.4×. We reduce DRAM energy by 1.6× vs. prior state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e8a3a2a-19b1-4d37-a2e7-a09d15f6604fBuilds on10
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 158 citations
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- Terrace: A Hierarchical Graph Container for Skewed Dynamic GraphsPrashant Pandey, Brian Wheatman, Helen Xu, Aydin BuluçSIGMOD 2021 · 53 citations
Related papers
- X-cache: a modular architecture for domain-specific cachesAli Sedaghati, Milad Hakimi, Reza Hojabr, Arrvindh ShriramanISCA 2022 · 11 citations
- RANGE-BLOCKS: A Synchronization Facility for Domain-Specific ArchitecturesAnagha Molakalmur Anil Kumar, Aditya Prasanna, Arrvindh ShriramanASPLOS 2025
- NOMAD: Enabling Non-blocking OS-managed DRAM Cache via Tag-Data DecouplingYoungin Kim, Hyeonjin Kim, William J. SongHPCA 2023 · 10 citations
- Profile-Guided Temporal PrefetchingMengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang et al.ISCA 2025 · 4 citations
- SeaCache: Efficient and Adaptive Caching for Sparse AcceleratorsXintong Li, Jinchen Jiang, Mingyu GaoMICRO 2025 · 1 citation
