Debiased Multimodal Understanding for Human Language Sequences
Zhi Xu, Dingkang Yang, Mingcheng Li, Yuzheng Wang, Zhaoyu Chen, Jiawei Chen, Jinjie Wei, Lihua Zhang
摘要
Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language ModelsJiawei Chen, Dingkang Yang, Yue Jiang, Mingcheng Li 等ACM MM 2024 · 被引用 5 次
- Sampling to Distill: Knowledge Transfer from Open-World DataYuzheng Wang, Zhaoyu Chen, Jie Zhang, Dingkang Yang 等ACM MM 2024 · 被引用 5 次
- Multi-Modal Image Fusion via Intervention-Stable Feature LearningXue Wang, Zheng Guan, Wenhua Qian, Chengchao Wang 等CVPR 2026 · 被引用 3 次
- Dual-Path Counterfactual Integration for Multimodal Aspect-Based Sentiment ClassificationRui Liu, Jiahao Cao, Jiaqian Ren, Xu Bai 等EMNLP 2025 · 被引用 1 次
- Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete ModalitiesMingcheng Li, Dingkang Yang, Xiao Zhao, Shuaibing Wang 等CVPR 2024
它引用的顶会 Paper21
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 被引用 737 次
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du 等ACM MM 2022 · 被引用 260 次
- How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionDingkang Yang, Kun Yang, Yuzheng Wang, Jing Liu 等NeurIPS 2023 · 被引用 160 次
相关 Paper
- Causal Intervention for Subject-Deconfounded Facial Action Unit RecognitionYingjie Chen, Diqi Chen, Tao Wang, Yizhou Wang 等AAAI 2022 · 被引用 36 次
- Context De-Confounded Emotion RecognitionDingkang Yang, Zhaoyu Chen, Yuzheng Wang, Shunli Wang 等CVPR 2023
- Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference PerspectiveSijie Mai, Shiqin HanACL 2026
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive LearningJiangfeng Sun, Sihao He, Zhonghong Ou, Meina SongAAAI 2026
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu 等CVPR 2026 · 被引用 7 次
