START: A Generalized State Space Model with Saliency-Driven Token-Aware Transformation
Jintao Guo, Lei Qi, Yinghuan Shi, Yang Gao
摘要
Domain Generalization (DG) aims to enable models to generalize to unseen target domains by learning from multiple source domains. Existing DG methods primarily rely on convolutional neural networks (CNNs), which inherently learn texture biases due to their limited receptive fields, making them prone to overfitting source domains. While some works have introduced transformer-based methods (ViTs) for DG to leverage the global receptive field, these methods incur high computational costs due to the quadratic complexity of self-attention. Recently, advanced state space models (SSMs), represented by Mamba, have shown promising results in supervised learning tasks by achieving linear complexity in sequence length during training and fast RNN-like computation during inference. Inspired by this, we investigate the generalization ability of the Mamba model under domain shifts and find that input-dependent matrices within SSMs could accumulate and amplify domain-specific features, thus hindering model generalization. To address this issue, we propose a novel SSM-based architecture with saliency-based token-aware transformation (namely START), which achieves state-of-the-art (SOTA) performances and offers a competitive alternative to CNNs and ViTs. Our START can selectively perturb and suppress domain-specific features in salient tokens within the input-dependent matrices of SSMs, thus effectively reducing the discrepancy between different domains. Extensive experiments on five benchmarks demonstrate that START outperforms existing SOTA DG methods with efficient linear complexity. Our code is available at https://github.com/lingeringlight/START.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LDT: Layer-Decomposition Training Makes Networks More GeneralizableZaizuo Tang, Zongqi Yang, Yu-Bin YangICLR 2026
- Test-time Domain Generalization for Image Super-resolutionZaizuo Tang, Yu-Bin YangICLR 2026
它引用的顶会 Paper48
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
相关 Paper
- DGMamba: Domain Generalization via Generalized State Space ModelShaocong Long, Qianyu Zhou, Xiangtai Li, Xuequan Lu 等ACM MM 2024 · 被引用 15 次
- MaskViM: Domain Generalized Semantic Segmentation with State Space ModelsJiahao Li, Yang Lu, Yuan Xie, Yanyun QuAAAI 2025 · 被引用 1 次
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic dataTianyi Chen, Pengxiao Lin, Zhiwei Wang, Zhi-Qin John XuNeurIPS 2025 · 被引用 4 次
- DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang 等CVPR 2026
- PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space ModelHao Yang, Qianyu Zhou, Haijia Sun, Xiangtai Li 等AAAI 2025 · 被引用 3 次
