From Bits to Tokens: Knowledge-Driven Generative Communication of Multimodal Data
Xingyu Chen, Zihao Feng, Wuqiong Zhao, Jianrong Ding, Ke Sun, Xinyu Zhang
Abstract
Classical communication systems strive for bit-by-bit reconstruction, yet this objective often misaligns with downstream application tasks, such as perception and decision-making with sensor data. This mismatch is amplified in wireless settings, where packet losses and channel dynamics make strict bit-fidelity both costly and fragile. In this paper, we introduce Knowledge-Driven Communication (KDC), a framework that transmits semantic knowledge rather than raw bits by leveraging pretrained knowledge bases. KDC features a task-aware transmitter, which uses multimodal foundation models to abstract source data into tokenized embeddings and prioritize semantically critical content, enabling zero-shot adaptation without task-specific retraining. On the receiver side, KDC employs pretrained knowledge bases and incrementally updated context to reconstruct task-relevant information, enabling graceful degradation even under data loss. We implement a full KDC prototype and evaluate it over diverse data modalities and wireless networks. KDC operates as an application-layer codec wrapper and receiver-side restoration module, fully compatible with existing wireless communication protocols and source/channel coding mechanisms. Experiments show that KDC consistently outperforms state-of-theart codecs and learned baselines, achieving high task accuracy with a fraction of the transmitted data, while maintaining robustness under challenging wireless conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Feature Weighting and Boosting for Few-Shot SegmentationKhoi Nguyen, Sinisa TodorovicICCV 2019 · 402 citations
Related papers
- Grasp: Refining Semantic Graphs into Purified Knowledge for Cross-Modal CommunicationLiang Chen, Xiaoding Wang, Limei Lin, Dajin Wang et al.WWW 2026
- Channel-Adaptive Denoising Diffusion Models for Reliable Semantic CommunicationsWei Du, Bo YangINFOCOM 2025 · 6 citations
- AdaSem: Adaptive Goal-Oriented Semantic Communications for End-to-End Camera RelocalizationQi Liao, Tze-Yang TungINFOCOM 2024 · 8 citations
- DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementYichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An et al.ICLR 2025
- Knowledge-Adaptation PriorsMohammad Emtiyaz Khan, Siddharth SwaroopNeurIPS 2021 · 32 citations
