IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models
Tao Feng, Lizhen Qu, Zhuang Li, Haolan Zhan, Yuncheng Hua, Gholamreza Haffari
Abstract
Machine learning models have made incredible progress, but they still struggle when applied to examples from unseen domains. This study focuses on a specific problem of domain generalization, where a model is trained on one source domain and tested on multiple target domains that are unseen during training. We propose IMO: Invariant features Masks for Out-of-Distribution text classification, to achieve OOD generalization by learning domain-invariant features. During training, IMO employs a greedy algorithm to learn sparse representations for each layer in a top-down manner. It performs better than the opposite direction and learning of sparse representations for all layers simultaneously. Our comprehensive experiments show that IMO substantially outperforms strong baselines such as prompt-based methods and large language models, in terms of various evaluation metrics and settings. 1 * Corresponding Author. 1 Codes are available at https://github.com/ WilliamsToTo/IMO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext feb85cf5-75cf-4cfb-ad0d-2bd44ba8a34aCited by top-tier papers4
- CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language ModelsKairong Han, Wenshuo Zhao, Ziyu Zhao, Ye Jun Jian et al.EMNLP 2025 · 3 citations
- IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular DataTao Feng, Lizhen Qu, Niket Tandon, Gholamreza HaffariACL 2025
- On the Reliability of Large Language Models for Causal DiscoveryTao Feng, Lizhen Qu, Niket Tandon, Zhuang Li et al.ACL 2025
- Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction SteeringLi Zheng, Xin Zhang, Shuyi He, Fei Li et al.ACL 2026
Builds on20
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman et al.ICML 2020 · 266 citations
Related papers
- Learning Beyond Domains: Misleading Prompts and Pseudo-Label Contrast for Text Domain GeneralizationQizhi Li, Xuyang Wang, Yingke Chen, Ming Yan et al.AAAI 2026 · 1 citation
- Domain Generalization in CLIP via Learning with Diverse Text PromptsChangsong Wen, Zelin Peng, Yu Huang, Xiaokang Yang et al.CVPR 2025
- Domain Generalization by Learning and Removing Domain-specific FeaturesYu Ding, Lei Wang, Bin Liang, Shuming Liang et al.NeurIPS 2022 · 75 citations
- Provable Domain Generalization via Invariant-Feature Subspace RecoveryHaoxiang Wang, Haozhe Si, Bo Li, Han ZhaoICML 2022 · 38 citations
- Disentangled Prompt Representation for Domain GeneralizationDe Cheng, Zhipeng Xu, Xinyang Jiang, Nannan Wang et al.CVPR 2024
