Multimodal Phased Transformer for Sentiment Analysis
Junyan Cheng, Iordanis Fostiropoulos, Barry W. Boehm, Mohammad Soleymani
Abstract
Multimodal Transformers achieve superior performance in multimodal learning tasks. However, the quadratic complexity of the selfattention mechanism in Transformers limits their deployment in low-resource devices and makes their inference and training computationally expensive. We propose multimodal Sparse Phased Transformer (SPT) to alleviate the problem of self-attention complexity and memory footprint. SPT uses a sampling function to generate a sparse attention matrix and compress a long sequence to a shorter sequence of hidden states. SPT concurrently captures interactions between the hidden states of different modalities at every layer. To further improve the efficiency of our method, we use Layer-wise parameter sharing and Factorized Co-Attention that share parameters between Cross Attention Blocks, with minimal impact on task performance. We evaluate our model with three sentiment analysis datasets and achieve comparable or superior performance compared with the existing methods, with a 90% reduction in the number of parameters. We conclude that (SPT) along with parameter sharing can capture multimodal interactions with reduced model size and improved sample efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a103d4f-6dc8-4bae-81e6-a4785d82dc07Cited by top-tier papers6
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu et al.EMNLP 2022 · 206 citations
- A Multi-Focus-Driven Multi-Branch Network for Robust Multimodal Sentiment AnalysisChuanqi Tao, Jiaming Li, Tianzi Zang, Peng GaoAAAI 2025 · 10 citations
- Bridging Neural and Symbolic Representations with Transitional Dictionary LearningJunyan Cheng, Peter ChinICLR 2024
- Faith: An Efficient Framework for Transformer Verification on GPUsBoyuan Feng, Tianqi Tang, Yuke Wang, Zhaodong Chen et al.USENIX ATC 2022
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive LearningJiangfeng Sun, Sihao He, Zhonghong Ou, Meina SongAAAI 2026
Builds on12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
Related papers
- Multimodal Transformers are Hierarchical Modal-wise Heterogeneous GraphsYijie Jin, Junjie Peng, Xuanchao Lin, Haochen Yuan et al.ACL 2025
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 11 citations
- Parameter Efficient Multimodal Transformers for Video Representation LearningSangho Lee, Youngjae Yu, Gunhee Kim, Thomas M. Breuel et al.ICLR 2021 · 90 citations
- Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion RecognitionZirun Guo, Tao Jin, Zhou ZhaoACL 2024 · 33 citations
- Diffuser: Efficient Transformers with Multi-Hop Attention Diffusion for Long SequencesAosong Feng, Irene Li, Yuang Jiang, Rex YingAAAI 2023 · 20 citations
