Efficient Quantification of Multimodal Interaction at Sample Level
Zequn Yang, Hongfa Wang, Di Hu
Abstract
Interactions between modalities-redundancy, uniqueness, and synergy-collectively determine the composition of multimodal information. Understanding these interactions is crucial for analyzing information dynamics in multimodal systems, yet their accurate sample-level quantification presents significant theoretical and computational challenges. To address this, we introduce the Lightweight Sample-wise Multimodal Interaction (LSMI) estimator, rigorously grounded in pointwise information theory. We first develop a redundancy estimation framework, employing an appropriate pointwise information measure to quantify this most decomposable and measurable interaction. Building upon this, we propose a general interaction estimation method that employs efficient entropy estimation, specifically tailored for sample-wise estimation in continuous distributions. Extensive experiments on synthetic and real-world datasets validate LSMI's precision and efficiency. Crucially, our sample-wise approach reveals fine-grained sample-and category-level dynamics within multimodal data, enabling practical applications such as redundancy-informed sample partitioning, targeted knowledge distillation, and interaction-aware model ensembling. The code is available at https://github. com/GeWu-Lab/LSMI_Estimator .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8ad1b61-0000-4f54-b5d3-399eea7f70f4Cited by top-tier papers4
- Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-ExpertsXing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang et al.ICLR 2026 · 1 citation
- Information-Theoretic Decomposition for Multimodal Interaction LearningZequn Yang, Yake Wei, Haotian Ni, Zhihao Xu et al.CVPR 2026 · 1 citation
- To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal PerformanceWanlong Fang, Tianle Zhang, Alvin ChanAAAI 2026
- Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language ModelsYuriel Ryan, Ip Man, Adriel Kuek, Paul Pu Liang et al.ICML 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz et al.ICLR 2020 · 152 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
Related papers
- Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkPaul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling et al.NeurIPS 2023 · 120 citations
- Partial Information Decomposition via Normalizing Flows in Latent Gaussian DistributionsWenyuan Zhao, Adithya Balachandran, Chao Tian, Paul Pu LiangNeurIPS 2025 · 5 citations
- Gaussian Partial Information Decomposition: Bias Correction and Application to High-dimensional DataPraveen Venkatesh, Corbett Bennett, Sam Gale, Tamina K. Ramirez et al.NeurIPS 2023 · 28 citations
- SΩI: Score-based O-INFORMATION EstimationMustapha Bounoua, Giulio Franzese, Pietro MichiardiICML 2024 · 3 citations
- Enhancing Multimodal Cooperation via Sample-Level Modality ValuationYake Wei, Ruoxuan Feng, Zihe Wang, Di HuCVPR 2024 · 24 citations
