Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision?
Zihao Li, Lecheng Zheng, Bowen Jin, Dongqi Fu, Baoyu Jing, Yikun Ban, Jingrui He, Jiawei Han
Abstract
While great success has been achieved in building vision models with Contrastive Language-Image Pre-training (CLIP) over internet-scale image-text pairs, building transferable Graph Neural Networks (GNNs) with CLIP pipeline is challenging because of the scarcity of labeled data and text supervision, different levels of downstream tasks, and the conceptual gaps between domains. In this work, to address these issues, we propose a multi-modal prompt learning paradigm to effectively adapt pre-trained GNN to downstream tasks and data, given only a few semantically labeled samples, each with extremely weak text supervision. Our new paradigm embeds the graphs directly in the same space as the Large Language Models (LLMs) by learning both graph prompts and text prompts simultaneously. We demonstrate the superior performance of our paradigm in few-shot, multi-tasklevel, and cross-domain settings. Moreover, we build the first CLIP-style zero-shot classification prototype that can generalize GNNs to unseen classes with extremely weak text supervision. The code is available at https: //github.com/Violet24K/Morpher .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 740a28da-4f09-4b23-944c-063362740afaCited by top-tier papers6
- HeroFilter: Adaptive Spectral Graph Filter for Varying Heterophilic RelationsShuaicheng Zhang, Haohui Wang, Junhong Lin, Xiaojie Guo et al.NeurIPS 2025 · 5 citations
- OwlEye: Zero-Shot Learner for Cross-Domain Graph Data Anomaly DetectionLecheng Zheng, Dongqi Fu, Zihao Li, Jingrui HeICLR 2026 · 2 citations
- Geometric Constraints for Small Language Models to Understand and Expand Scientific TaxonomiesLiri Fang, Dongqi Fu, Jiawei Han, Jingrui He et al.ICLR 2026
- Learnable Spatial-Temporal Positional Encoding for Link PredictionKatherine Tieu, Dongqi Fu, Zihao Li, Ross Maciejewski et al.ICML 2025
- SelfElicit: Your Language Model Secretly Knows Where is the Relevant EvidenceZhining Liu, Rana Ali Amjad, Ravinarayana Adkathimar, Tianxin Wei et al.ACL 2025
Builds on64
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
Related papers
- GraphCLIP: Enhancing Transferability in Graph Foundation Models for Text-Attributed GraphsYun Zhu, Haizhou Shi, Xiaotang Wang, Yongchao Liu et al.WWW 2025 · 54 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 21 citations
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan et al.CVPR 2023
- Edge Prompt Tuning for Graph Neural NetworksXingbo Fu, Yinhan He, Jundong LiICLR 2025 · 140 citations
