LDT: Layer-Decomposition Training Makes Networks More Generalizable
Zaizuo Tang, Zongqi Yang, Yu-Bin Yang
摘要
Domain generalization methods can effectively enhance network performance on test samples with unknown distributions by isolating gradients between unstable and stable parameters. However, existing methods employ relatively coarse-grained partitioning of stable versus unstable parameters, leading to misclassified unstable parameters that degrade network feature processing capabilities. We first provide a theoretical analysis of gradient perturbations caused by unstable parameters. Based on this foundation, we propose Layer-Decomposition Training (LDT), which conducts fine-grained layer-wise partitioning guided by parameter instability levels, substantially improving parameter update stability. Furthermore, to address gradient amplitude disparities within stable layers and unstable layers respectively, we introduce a Dynamic Parameter Update (DPU) strategy that adaptively determines layer-specific update coefficients according to gradient variations, optimizing feature learning efficiency. Extensive experiments across diverse tasks (super-resolution, classification, semantic segmentation) and architectures (Transformer, Mamba, CNN) demonstrate LDT's superior generalization capability. Our code is available at https://github.com/ZaizuoTang/LDT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Spatially-Adaptive Feature Modulation for Efficient Image Super-ResolutionLong Sun, Jiangxin Dong, Jinhui Tang, Jinshan PanICCV 2023 · 被引用 211 次
相关 Paper
- Parameter Exchange for Robust Dynamic Domain GeneralizationLuojun Lin, Zhifeng Shen, Zhishu Sun, Yuanlong Yu 等ACM MM 2023 · 被引用 7 次
- Gradient Smoothing: Coupling Layer-wise Updates for Improved OptimizationHaoming Meng, Anton Sugolov, Vardan PapyanICML 2026
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao 等ICLR 2026 · 被引用 5 次
- Domain Generalization with Vital Phase AugmentationIngyun Lee, Wooju Lee, Hyun MyungAAAI 2024 · 被引用 12 次
- March on Data Imperfections: Domain Division and Domain Generalization for Semantic SegmentationHai Xu, Hongtao Xie, Zheng-Jun Zha, Sun'ao Liu 等ACM MM 2020 · 被引用 4 次
