MOBI: Monolithic Graph-Language Modeling Beyond Modality Interference
Zhiyao Zhou, Yugang Ji, Ziwen Xu, Zhuonan Zheng, Sheng Zhou, Weigao Wen, Jiawei Chen, Ming Gu, Chun Chen, Can Wang
Abstract
Graph-Language Models (GLMs) aim to endow LLMs with structure-grounded reasoning ability, yet existing solutions often struggle with modality interference : structural information can disrupt pretrained linguistic reasoning, while language cues can overwhelm structural signals. Mainstream modular GLMs with an external graph encoder attempt to mitigate the interference by separating graph encoding from language decoding. This separation fails to strike an effective balance between modality fusion and interference, exhibiting limited cross-modal interaction while leaving interference between modalities largely unresolved. To tackle the above challenges, we propose MOBI (Monolithic Graph-Language Modeling Beyond Modality Interference), a monolithic graph-language model that unifies graph encoding and language decoding within a single backbone for end-to-end graph-text fusion. Specifically, MOBI overcomes modality interference via (i) dual-pathway transformer that preserves pretrained linguistic knowledge while acquiring structural understanding, (ii) progressive interaction scheduling that suppresses cross-modal noise by dynamically regulating the information flow, and (iii) correlation-guided attribute perturbation that discourages textual shortcuts on text-attributed graphs. Extensive experiments on 11 datasets across various settings demonstrate that MOBI consistently outperforms modular baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63ed2f0d-5999-490c-99ed-3a500cf160f5Builds on40
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu et al.CVPR 2026 · 2 citations
- Graph Reasoning Transformers for Knowledge-Aware Question AnsweringRuilin Zhao, Feng Zhao, Liang Hu, Guandong XuAAAI 2024 · 10 citations
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal ReasoningYucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar et al.ICLR 2026 · 3 citations
- Graph Language ModelsMoritz Plenz, Anette FrankACL 2024
- DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality ExpertsMingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu et al.ACM MM 2025
