Sparse Binary Transformers for Multivariate Time Series Modeling
Matt Gorbett, Hossein Shirazi, Indrakshi Ray
Abstract
Compressed Neural Networks have the potential to enable deep learning across new applications and smaller computational environments. However, understanding the range of learning tasks in which such models can succeed is not well studied. In this work, we apply sparse and binary-weighted Transformers to multivariate time series problems, showing that the lightweight models achieve accuracy comparable to that of dense floating-point Transformers of the same structure. Our model achieves favorable results across three time series learning tasks: classification, anomaly detection, and single-step forecasting. Additionally, to reduce the computational complexity of the attention mechanism, we apply two modifications, which show little to no decline in model performance: 1) in the classification task, we apply a fixed mask to the query, key, and value activations, and 2) for forecasting and anomaly detection, which rely on predicting outputs at a single point in time, we propose an attention mask to allow computation only at the current time step. Together, each compression technique and attention modification substantially reduces the number of non-zero operations necessary in the Transformer. We measure the computational savings of our approach over a range of metrics including parameter count, bit size, and floating point operation (FLOPs) count, showing up to a 53× reduction in storage size and up to 10.5× reduction in FLOPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a56fe3b2-8785-444d-8beb-aac4dd196bb6Cited by top-tier papers1
Ask how each one uses itBuilds on25
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
Related papers
- Towards Lightweight Time Series Forecasting: A Patch-Wise Transformer with Weak Data EnrichingMeng Wang, Jintao Yang, Bin Yang, Hui Li et al.ICDE 2025 · 10 citations
- Deeply Tensor Compressed Transformers for End-to-End Object DetectionPeining Zhen, Ziyang Gao, Tianshu Hou, Yuan Cheng et al.AAAI 2022 · 19 citations
- TSLANet: Rethinking Transformers for Time Series Representation LearningEmadeldeen Eldele, Mohamed Ragab, Zhenghua Chen, Min Wu et al.ICML 2024 · 159 citations
- Continual Transformers: Redundancy-Free Attention for Online InferenceLukas Hedegaard, Arian Bakhtiarnia, Alexandros IosifidisICLR 2023 · 3 citations
- Efficient Transformer Inference with Statically Structured Sparse AttentionSteve Dai, Hasan Genc, Rangharajan Venkatesan, Brucek KhailanyDAC 2023 · 11 citations
