Grasp: Refining Semantic Graphs into Purified Knowledge for Cross-Modal Communication
Liang Chen, Xiaoding Wang, Limei Lin, Dajin Wang, Zhiquan Liu, Jie Wu
Abstract
The explosive growth of multimodal web data demands communication that transmits meaning rather than raw bits. Existing semantic-communication systems often fail under noise, missing modalities, and distribution shifts because they optimize surface features instead of modality-invariant knowledge. We present Grasp, a knowledge-centric framework for cross-modal communication. Grasp segments streams into semantic blocks and builds a graph over them; a lightweight Graph Neural Networks (GNN) produces schedulable, importance-weighted representations. At its core is knowledge purification: we minimize a conditional mutual information upper bound to perform a three-way disentanglement-strongly related, weakly related, and task-irrelevant components-so that only essential semantics are transmitted while non-essential factors are suppressed. To maintain synchrony, we introduce one-totwo temporal contrastive learning to achieve triple alignment of video, audio, and text despite sampling asynchrony. For efficient transmission, Grasp uses a cross-modal shared vector-quantization codebook-a discrete knowledge codebook-updated by multimodal attention. At the receiver, a soft-recovery mechanism leverages this shared knowledge to robustly reconstruct semantics under low signal-to-noise ratio (SNR) or missing modalities, yielding graceful degradation. Across web tasks-including cross-modal retrieval and missing-modality inference-Grasp improves knowledge consistency, semantic fidelity, and downstream performance over strong baselines while maintaining low latency. These results show that communication structured around purified knowledge is key to building robust, semantic-aware systems for the modern web. CCS Concepts • Theory of computation → Semantics and reasoning; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- SeqCare: Sequential Training with External Medical Knowledge Graph for Diagnosis Prediction in Healthcare DataYongxin Xu, Xu Chu, Kai Yang, Zhiyuan Wang et al.WWW 2023 · 41 citations
- Mutual Information Estimation via Normalizing FlowsIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya et al.NeurIPS 2024 · 30 citations
- Unveiling Discrete Clues: Superior Healthcare Predictions for Rare DiseasesChuang Zhao, Hui Tang, Jiheng Zhang, Xiaomeng LiWWW 2025 · 8 citations
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationLiao Qu, Huichao Zhang, Yiheng Liu, Xu Wang et al.CVPR 2025
Related papers
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan et al.CVPR 2021
- Scaling Multimodal Pre-Training via Cross-Modality Gradient HarmonizationJunru Wu, Yi Liang, Feng Han, Hassan Akbari et al.NeurIPS 2022 · 20 citations
- From Bits to Tokens: Knowledge-Driven Generative Communication of Multimodal DataXingyu Chen, Zihao Feng, Wuqiong Zhao, Jianrong Ding et al.NSDI 2026 · 1 citation
- CLOP: Video-and-Language Pre-Training with Knowledge RegularizationsGuohao Li, Hu Yang, Feng He, Zhifan Feng et al.ACM MM 2022 · 1 citation
- Multimodal Knowledge Graph Error Detection with Disentanglement VAE and Multi-Grained Triplet ConfidenceXuhui Sui, Ying Zhang, Yu Zhao, Baohang Zhou et al.WWW 2025 · 2 citations
