A Constrained Optimization Perspective of Unrolled Transformers
Javier Porras-Valenzuela, Samar Hadou, Alejandro Ribeiro
Abstract
We introduce a constrained optimization framework for training transformers that behave like optimization descent algorithms. Specifically, we enforce layerwise descent constraints on the objective function and replace standard empirical risk minimization (ERM) with a primal-dual training scheme. This approach yields models whose intermediate representations decrease the loss monotonically in expectation across layers. We apply our method to both unrolled transformer architectures and conventional pretrained transformers on tasks of video denoising and text classification. Across these settings, we observe constrained transformers achieve stronger robustness to perturbations and maintain higher out-of-distribution generalization, while preserving in-distribution performance. The code is available at https://github.com/ jotaporras/constrained-unrolledtransformers
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
Related papers
- Loss Shaping Constraints for Long-Term Time Series ForecastingIgnacio Hounie, Javier Porras-Valenzuela, Alejandro RibeiroICML 2024 · 7 citations
- Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot ModelsKaican Li, Weiyan Xie, Yongxiang Huang, Didan Deng et al.NeurIPS 2024 · 5 citations
- Approximation Theory for Lipschitz Continuous TransformersTakashi Furuya, Davide Murari, Carola-Bibiane SchönliebICML 2026 · 4 citations
- LORE: Lagrangian-Optimized Robust Embeddings for Visual EncodersBorna Khodabandeh, Amirabbas Afzali, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi et al.NeurIPS 2025
- Hierarchically Robust Representation LearningQi Qian, Juhua Hu, Hao LiCVPR 2020
