Lune

ICML2026Top-tier venue

Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

Yuriel Ryan, Ip Man, Adriel Kuek, Paul Pu Liang, Roy Lee

2026Year

Abstract

Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting the shared information between modalities to compensate for the impaired one. To this end, we analyze multimodal interactions -redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information provided by the modalities -to determine their impacts on model reliability. Specifically, amplifying redundant interactions would increase this exploitable shared information to resolve these issues; yet, modern instruction datasets often eliminate redundancies to prioritize visual grounding. We bridge this gap through a self-captioning workflow featuring a MULTIMODAL INTERACTION GATE: a mechanism to convert unique interactions into redundant interactions. Our findings suggest that increasing redundancy can reduce visual induced errors by 38.3% and improve consistency by 16.8%. Code available at https://github.com/yurie lryan/Multimodal-Interaction-Tun ing.git

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f12761d8-eb9f-4a93-b52a-9773c0048b15

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines