TRAM: Bridging Trust Regions and Sharpness Aware Minimization
Tom Sherborne, Naomi Saphra, Pradeep Dasigi, Hao Peng
Abstract
Sharpness-aware minimization (SAM) reports improving domain generalization by reducing the loss surface curvature in the parameter space. However, generalization during fine-tuning is often more dependent on the transferability of representations in the function space. Trust-region methods (TR) target this goal by regularizing representation curvature to reduce catastrophic forgetting of pre-trained task-agnostic information while adopting task-specific skills. We consider unifying these strategies for low curvature in both parameter space and function space to improve out-of-domain (OOD) generalization. We propose Trust Region Aware Minimization (TRAM), a SAM algorithm fine-tuning for low parameter sharpness and smooth, informative representations preserving pre-trained structure. TRAM uses a trust region bound to inform the SAM adversarial neighborhood, introducing an awareness of function curvature within optimization for flatter minima. We empirically validate TRAM in vision (cross-dataset adaptation) and text (OOD language modeling, zero-shot cross-lingual transfer) tasks where robust domain transfer and representation generality are critical. TRAM outperforms SAM- and TR-based optimization across all tasks, notably surpassing competing methods for hard transfer between anticorrelated domains. TRAM establishes a novel standard in fine-tuning for domain-generalizable models with minimal additional computation over previous sharpness-aware methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant ProblemsBingcong Li, Liang Zhang, Niao HeNeurIPS 2024 · 14 citations
- Avoiding spurious sharpness minimization broadens applicability of SAMSidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. DauphinICML 2025
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Better Fine-Tuning by Reducing Representational CollapseArmen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal et al.ICLR 2021 · 20 citations
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu et al.NeurIPS 2024 · 48 citations
- Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware MinimizationJinping Wang, Qinhan Liu, Zhiwu Xie, Zhiqiang GaoICML 2026
- Sharpness-Aware Gradient Matching for Domain GeneralizationPengfei Wang, Zhaoxiang Zhang, Zhen Lei, Lei ZhangCVPR 2023
- Tilted Sharpness-Aware MinimizationTian Li, Tianyi Zhou, Jeff A. BilmesICML 2025
