Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
Divyam Madaan, Taro Makino, Sumit Chopra, Kyunghyun Cho
摘要
Supervised multi-modal learning involves mapping multiple modalities to a target label. Previous studies in this field have concentrated on capturing in isolation either the inter-modality dependencies (the relationships between different modalities and the label) or the intra-modality dependencies (the relationships within a single modality and the label). We argue that these conventional approaches that rely solely on either inter- or intra-modality dependencies may not be optimal in general. We view the multi-modal learning problem from the lens of generative models where we consider the target as a source of multiple modalities and the interaction between them. Towards that end, we propose inter-&intra-modality modeling (I2M2) framework, which captures and integrates both the inter- and intra-modality dependencies, leading to more accurate predictions. We evaluate our approach using real-world healthcare and vision-and-language datasets with state-of-the-art models, demonstrating superior performance over traditional methods focusing only on one type of modality dependency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time SeriesPayal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia 等NeurIPS 2025 · 被引用 19 次
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 被引用 17 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 被引用 3 次
- Information-Theoretic Decomposition for Multimodal Interaction LearningZequn Yang, Yake Wei, Haotian Ni, Zhihao Xu 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
相关 Paper
- MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual InvariantChenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang 等CVPR 2024 · 被引用 20 次
- M3TR: Multi-modal Multi-label Recognition with TransformerJiawei Zhao, Yifan Zhao, Jia LiACM MM 2021 · 被引用 45 次
- Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisBing Cao, Han Zhang, Nannan Wang, Xinbo Gao 等AAAI 2020 · 被引用 94 次
- Towards the Causal Complete Cause of Multi-Modal Representation LearningJingyao Wang, Siyu Zhao, Wenwen Qiang, Jiangmeng Li 等ICML 2025
- Transformer-based Label Set Generation for Multi-modal Multi-label Emotion DetectionXincheng Ju, Dong Zhang, Junhui Li, Guodong ZhouACM MM 2020 · 被引用 64 次
