Interpretable Lightweight Transformer via Unrolling of Learned Graph Smoothness Priors
Viet Ho Tam Thuc Do, Parham Eftekhar, Seyed Alireza Hosseini, Gene Cheung, Philip A. Chou
摘要
We build interpretable and lightweight transformer-like neural networks by unrolling iterative optimization algorithms that minimize graph smoothness priors -- the quadratic graph Laplacian regularizer (GLR) and the -norm graph total variation (GTV) -- subject to an interpolation constraint. The crucial insight is that a normalized signal-dependent graph learning module amounts to a variant of the basic self-attention mechanism in conventional transformers. Unlike"black-box"transformers that require learning of large key, query and value matrices to compute scaled dot products as affinities and subsequent output embeddings, resulting in huge parameter sets, our unrolled networks employ shallow CNNs to learn low-dimensional features per node to establish pairwise Mahalanobis distances and construct sparse similarity graphs. At each layer, given a learned graph, the target interpolated signal is simply a low-pass filtered output derived from the minimization of an assumed graph smoothness prior, leading to a dramatic reduction in parameter count. Experiments for two image interpolation applications verify the restoration performance, parameter efficiency and robustness to covariate shift of our graph-based unrolled networks compared to conventional transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm UnrollingJunyi Yao, Parham Eftekhar, Gene Cheung, Xujin Chris Liu 等ICLR 2026 · 被引用 1 次
- Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric AttentionWeida Wang, Changyong He, Jin Zeng, Di QiuICCV 2025 · 被引用 1 次
- Lightweight and Interpretable Transformer via Unrolling of Mixed Graph Algorithms for Traffic ForecastJi Qi, Mingxiao Liu, VIET THUC, Yuzhe Li 等ICML 2026
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 被引用 208 次
相关 Paper
- CLG-INet: Coupled Local-Global Interactive Network for Image RestorationYuqi Jiang, Chune Zhang, Shuo Jin, Jiao Liu 等ACM MM 2023 · 被引用 4 次
- Compacter: A Lightweight Transformer for Image RestorationZhijian Wu, Jun Li, Yang Hu, Dingjiang HuangACM MM 2024
- A Flexible Framework for Designing Trainable Priors with Adaptive Smoothing and Game EncodingBruno Lecouat, Jean Ponce, Julien MairalNeurIPS 2020 · 被引用 16 次
- Wavy TransformerSatoshi Noguchi, Yoshinobu KawaharaNeurIPS 2025 · 被引用 1 次
- Dual Graph Regularized Deep Unfolding Network for Guided Depth Map Super-resolutionZhiwei Zhong, Peilin Chen, Qiangqiang Shen, Bo Li 等CVPR 2026 · 被引用 3 次
