Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
Hongcheng Liu, Yuhao Wang, Zhe Chen, Pingjie Wang, Zhiyuan Zhu, Yixuan Hou, Yanfeng Wang, Yu Wang
Abstract
Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring synergistic omni-modal reasoning. Beyond understanding global multimodal context, effective reasoning also hinges on fine-grained crossmodal alignment, especially identifying shared referents across modalities, yet this aspect has been largely overlooked. To bridge this gap, we formalize the challenge as a cross-modal coreference problem, where a model must localize a referent in a source modality and reidentify it in a target modality. Building on this paradigm, we introduce CROSSOMNI, a dataset comprising nine tasks equipped with human-designed reasoning rationales to evaluate and enhance this capability. Experiments on 13 Omni-LLMs reveal systematic weaknesses in cross-modal coreference, which we attribute to the absence of coreference-aware thinking patterns. To address this, we enhance crossmodal alignment via two strategies: a trainingfree In-Context Learning method and a trainingbased SFT+GRPO framework designed to induce such thinking patterns. Both approaches yield substantial performance gains and generalize effectively to collaborative reasoning tasks. Overall, our findings highlight crossmodal coreference as a crucial missing piece for advancing robust omni-modal reasoning. Auto-Generation Human Verification Double Check Consistency Test-set Distribution Q: What is the occupation of the person who said ''when did he get out of prison''? Option: A. nurse B. officer C. sergeant D. detective
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang et al.ICLR 2026 · 162 citations
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang et al.ICLR 2026 · 143 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
Related papers
- ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance DecodingYiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu et al.ICLR 2026 · 5 citations
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng et al.CVPR 2026 · 3 citations
- R2-MultiOmnia: Leading Multilingual Multimodal Reasoning via Self-TrainingLeonardo Ranaldi, Federico Ranaldi, Giulia PucciACL 2025 · 9 citations
- DiMA: Distinguishing Resident and Tourist Preferences via Multi-Modal LLM Alignment for Out-of-Town Cross-Domain RecommendationFan Zhang, Jinpeng Chen, Tao Wang, Huan Li et al.AAAI 2026
- Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin et al.ACL 2026
