Optimal Transport Model Distributional Robustness
Van-Anh Nguyen, Trung Le, Anh Tuan Bui, Thanh-Toan Do, Dinh Q. Phung
Abstract
Distributional robustness is a promising framework for training deep learning models that are less vulnerable to adversarial examples and data distribution shifts. Previous works have mainly focused on exploiting distributional robustness in the data space. In this work, we explore an optimal transport-based distributional robustness framework in model spaces. Specifically, we examine a model distribution within a Wasserstein ball centered on a given model distribution that maximizes the loss. We have developed theories that enable us to learn the optimal robust center model distribution. Interestingly, our developed theories allow us to flexibly incorporate the concept of sharpness awareness into training, whether it's a single model, ensemble models, or Bayesian Neural Networks, by considering specific forms of the center model distribution. These forms include a Dirac delta distribution over a single model, a uniform distribution over several models, and a general Bayesian Neural Network. Furthermore, we demonstrate that Sharpness-Aware Minimization (SAM) is a specific case of our framework when using a Dirac delta distribution over a single model, while our framework can be seen as a probabilistic extension of SAM. To validate the effectiveness of our framework in the aforementioned settings, we conducted extensive experiments, and the results reveal remarkable improvements compared to the baselines. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 962510ff-d34c-447f-ba68-9662a078259fCited by top-tier papers5
- Flat Seeking Bayesian Neural NetworksVan-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do et al.NeurIPS 2023 · 14 citations
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient InfluenceQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui et al.ICLR 2026 · 10 citations
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 2 citations
- Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference OptimizationHaocheng Luo, Zehang Deng, Thanh-Toan Do, Mehrtash Harandi et al.ICLR 2026 · 1 citation
- Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation ModelsNgoc-Quan Pham, Tuan Truong, Quyen Tran, Tan Minh Nguyen et al.ICML 2025
Builds on23
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
Related papers
- A Unified Wasserstein Distributional Robustness Framework for Adversarial TrainingAnh Tuan Bui, Trung Le, Quan Hung Tran, He Zhao et al.ICLR 2022 · 54 citations
- Universal generalization guarantees for Wasserstein distributionally robust modelsTam Le, Jérôme MalickICLR 2025
- Bootstrap Your Uncertainty: Adaptive Robust Classification Driven by Optimal-TransportJiawei Huang, Minming Li, Hu DingNeurIPS 2025
- SAM as an Optimal Relaxation of BayesThomas Möllenhoff, Mohammad Emtiyaz KhanICLR 2023 · 3 citations
- Adversarial Distributional Training for Robust Deep LearningYinpeng Dong, Zhijie Deng, Tianyu Pang, Jun Zhu et al.NeurIPS 2020 · 154 citations
