MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm Detection
Yifeng Xie, Zhihong Zhu, Xin Chen, Zhanpeng Chen, Zhiqi Huang
Abstract
In the field of multi-modal learning, model parameters are typically large, necessitating the use of parameter-efficient fine-tuning (PEFT) techniques. These methods have been pivotal in enhancing training efficiency for downstream tasks in almost all situations. However, directly applying PEFT methods struggles to fully address the intricate demands of multi-modal tasks, such as multi-modal sarcasm detection (MSD), which demands the extraction and comparison of cues from different modalities. MSD, particularly when reliant on textual and visual modalities, faces challenges in identifying sarcasm's incongruity. This issue often arises from the lack of intermodality interaction during tuning, resulting in a disconnect between textual and visual information. In this paper, we introduce a novel approach called Bi-directional Adapter (BA), designated as MoBA. This approach is designed to minimize training parameters while enhancing the model's ability to interpret sarcasm across modalities. By facilitating an exchange between textual and visual information through a low-rank representation, our method adeptly captures the nuances of sarcastic expressions with a reduced number of training parameters. Our empirical studies, carried out on two publicly accessible and emerging datasets, demonstrate that our model substantially improves sarcasm detection accuracy. These findings indicate that our approach provides a more reliable and efficient solution to address the complexities of MSD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3067e826-a094-4772-bdbc-59f2a1e3996eCited by top-tier papers2
- MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm DetectionHaochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou et al.CVPR 2026 · 4 citations
- SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image ComprehensionYue Jiang, Haiwei Xue, Minghao Han, Mingcheng Li et al.AAAI 2026 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Parameter-, Memory-, Time-Efficient Multi-Task Dense Vision AdaptationHaiming Yao, Wei Luo, Qiyu Chen, Jianxing Liao et al.AAAI 2026
- PEFT-BoA: Parameter-Efficient Fine-Tuning with Bag-of-Adapters for Multi-Modal Object Re-identificationHongchao Li, Guangxing Liu, Xixi Wang, Baihe Liang et al.AAAI 2026
- HALoRA: Low-Rank Adaptation with Hierarchical Budget Allocation for Efficient Vision-Language AlignmentLetian Zhang, Guanghao Meng, Xudong Ren, Jinpeng WangAAAI 2026
- DisLoRA: Task-specific Low-Rank Adaptation via Orthogonal Basis from Singular Value DecompositionShe Yifei, Xinhao Wei, Yulong WangEMNLP 2025
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen et al.ICLR 2026 · 5 citations
