Grasp: Refining Semantic Graphs into Purified Knowledge for Cross-Modal Communication
Liang Chen, Xiaoding Wang, Limei Lin, Dajin Wang, Zhiquan Liu, Jie Wu
摘要
The explosive growth of multimodal web data demands communication that transmits meaning rather than raw bits. Existing semantic-communication systems often fail under noise, missing modalities, and distribution shifts because they optimize surface features instead of modality-invariant knowledge. We present Grasp, a knowledge-centric framework for cross-modal communication. Grasp segments streams into semantic blocks and builds a graph over them; a lightweight Graph Neural Networks (GNN) produces schedulable, importance-weighted representations. At its core is knowledge purification: we minimize a conditional mutual information upper bound to perform a three-way disentanglement-strongly related, weakly related, and task-irrelevant components-so that only essential semantics are transmitted while non-essential factors are suppressed. To maintain synchrony, we introduce one-totwo temporal contrastive learning to achieve triple alignment of video, audio, and text despite sampling asynchrony. For efficient transmission, Grasp uses a cross-modal shared vector-quantization codebook-a discrete knowledge codebook-updated by multimodal attention. At the receiver, a soft-recovery mechanism leverages this shared knowledge to robustly reconstruct semantics under low signal-to-noise ratio (SNR) or missing modalities, yielding graceful degradation. Across web tasks-including cross-modal retrieval and missing-modality inference-Grasp improves knowledge consistency, semantic fidelity, and downstream performance over strong baselines while maintaining low latency. These results show that communication structured around purified knowledge is key to building robust, semantic-aware systems for the modern web. CCS Concepts • Theory of computation → Semantics and reasoning; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- SeqCare: Sequential Training with External Medical Knowledge Graph for Diagnosis Prediction in Healthcare DataYongxin Xu, Xu Chu, Kai Yang, Zhiyuan Wang 等WWW 2023 · 被引用 41 次
- Mutual Information Estimation via Normalizing FlowsIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya 等NeurIPS 2024 · 被引用 30 次
- Unveiling Discrete Clues: Superior Healthcare Predictions for Rare DiseasesChuang Zhao, Hui Tang, Jiheng Zhang, Xiaomeng LiWWW 2025 · 被引用 8 次
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationLiao Qu, Huichao Zhang, Yiheng Liu, Xu Wang 等CVPR 2025
相关 Paper
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan 等CVPR 2021
- Scaling Multimodal Pre-Training via Cross-Modality Gradient HarmonizationJunru Wu, Yi Liang, Feng Han, Hassan Akbari 等NeurIPS 2022 · 被引用 20 次
- From Bits to Tokens: Knowledge-Driven Generative Communication of Multimodal DataXingyu Chen, Zihao Feng, Wuqiong Zhao, Jianrong Ding 等NSDI 2026 · 被引用 1 次
- CLOP: Video-and-Language Pre-Training with Knowledge RegularizationsGuohao Li, Hu Yang, Feng He, Zhifan Feng 等ACM MM 2022 · 被引用 1 次
- Multimodal Knowledge Graph Error Detection with Disentanglement VAE and Multi-Grained Triplet ConfidenceXuhui Sui, Ying Zhang, Yu Zhao, Baohang Zhou 等WWW 2025 · 被引用 2 次
