MOBI: Monolithic Graph-Language Modeling Beyond Modality Interference
Zhiyao Zhou, Yugang Ji, Ziwen Xu, Zhuonan Zheng, Sheng Zhou, Weigao Wen, Jiawei Chen, Ming Gu, Chun Chen, Can Wang
摘要
Graph-Language Models (GLMs) aim to endow LLMs with structure-grounded reasoning ability, yet existing solutions often struggle with modality interference : structural information can disrupt pretrained linguistic reasoning, while language cues can overwhelm structural signals. Mainstream modular GLMs with an external graph encoder attempt to mitigate the interference by separating graph encoding from language decoding. This separation fails to strike an effective balance between modality fusion and interference, exhibiting limited cross-modal interaction while leaving interference between modalities largely unresolved. To tackle the above challenges, we propose MOBI (Monolithic Graph-Language Modeling Beyond Modality Interference), a monolithic graph-language model that unifies graph encoding and language decoding within a single backbone for end-to-end graph-text fusion. Specifically, MOBI overcomes modality interference via (i) dual-pathway transformer that preserves pretrained linguistic knowledge while acquiring structural understanding, (ii) progressive interaction scheduling that suppresses cross-modal noise by dynamically regulating the information flow, and (iii) correlation-guided attribute perturbation that discourages textual shortcuts on text-attributed graphs. Extensive experiments on 11 datasets across various settings demonstrate that MOBI consistently outperforms modular baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper40
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu 等CVPR 2026 · 被引用 2 次
- Graph Reasoning Transformers for Knowledge-Aware Question AnsweringRuilin Zhao, Feng Zhao, Liang Hu, Guandong XuAAAI 2024 · 被引用 10 次
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal ReasoningYucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar 等ICLR 2026 · 被引用 3 次
- Graph Language ModelsMoritz Plenz, Anette FrankACL 2024
- DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality ExpertsMingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu 等ACM MM 2025
