Interpretable Lightweight Transformer via Unrolling of Learned Graph Smoothness Priors
Viet Ho Tam Thuc Do, Parham Eftekhar, Seyed Alireza Hosseini, Gene Cheung, Philip A. Chou
Abstract
We build interpretable and lightweight transformer-like neural networks by unrolling iterative optimization algorithms that minimize graph smoothness priors -- the quadratic graph Laplacian regularizer (GLR) and the -norm graph total variation (GTV) -- subject to an interpolation constraint. The crucial insight is that a normalized signal-dependent graph learning module amounts to a variant of the basic self-attention mechanism in conventional transformers. Unlike"black-box"transformers that require learning of large key, query and value matrices to compute scaled dot products as affinities and subsequent output embeddings, resulting in huge parameter sets, our unrolled networks employ shallow CNNs to learn low-dimensional features per node to establish pairwise Mahalanobis distances and construct sparse similarity graphs. At each layer, given a learned graph, the target interpolated signal is simply a low-pass filtered output derived from the minimization of an assumed graph smoothness prior, leading to a dramatic reduction in parameter count. Experiments for two image interpolation applications verify the restoration performance, parameter efficiency and robustness to covariate shift of our graph-based unrolled networks compared to conventional transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm UnrollingJunyi Yao, Parham Eftekhar, Gene Cheung, Xujin Chris Liu et al.ICLR 2026 · 1 citation
- Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric AttentionWeida Wang, Changyong He, Jin Zeng, Di QiuICCV 2025 · 1 citation
- Lightweight and Interpretable Transformer via Unrolling of Mixed Graph Algorithms for Traffic ForecastJi Qi, Mingxiao Liu, VIET THUC, Yuzhe Li et al.ICML 2026
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
Related papers
- CLG-INet: Coupled Local-Global Interactive Network for Image RestorationYuqi Jiang, Chune Zhang, Shuo Jin, Jiao Liu et al.ACM MM 2023 · 4 citations
- Compacter: A Lightweight Transformer for Image RestorationZhijian Wu, Jun Li, Yang Hu, Dingjiang HuangACM MM 2024
- A Flexible Framework for Designing Trainable Priors with Adaptive Smoothing and Game EncodingBruno Lecouat, Jean Ponce, Julien MairalNeurIPS 2020 · 16 citations
- Wavy TransformerSatoshi Noguchi, Yoshinobu KawaharaNeurIPS 2025 · 1 citation
- Dual Graph Regularized Deep Unfolding Network for Guided Depth Map Super-resolutionZhiwei Zhong, Peilin Chen, Qiangqiang Shen, Bo Li et al.CVPR 2026 · 3 citations
