Learning based Multi-modality Image and Video Compression
Guo Lu, Tianxiong Zhong, Jing Geng, Qiang Hu, Dong Xu
摘要
Multi-modality (i.e., multi-sensor) data is widely used in various vision tasks for more accurate or robust perception. However, the increased data modalities bring new challenges for data storage and transmission. The existing data compression approaches usually adopt individual codecs for each modality without considering the correlation between different modalities. This work proposes a multi-modality compression framework for infrared and visible image pairs by exploiting the cross-modality redundancy. Specifically, given the image in the reference modality (e.g., the infrared image), we use the channel-wise alignment module to produce the aligned features based on the affine transform. Then the aligned feature is used as the context information for compressing the image in the current modality (e.g., the visible image), and the corresponding affine coefficients are losslessly compressed at negligible cost. Furthermore, we introduce the Transformer-based spatial alignment module to exploit the correlation between the intermediate features in the decoding procedures for different modalities. Our framework is very flexible and easily extended for multi-modality video compression. Experimental results show our proposed framework outperforms the traditional and learning-based single modality compression methods on the FLIR and KAIST datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic CompressionYuan Tian, Guo Lu, Guangtao Zhai, Zhiyong GaoICCV 2023 · 被引用 29 次
- VRVVC: Variable-Rate NeRF-Based Volumetric Video CompressionQiang Hu, Houqiang Zhong, Zihan Zheng, Xiaoyun Zhang 等AAAI 2025 · 被引用 11 次
- You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed VideosXiang Fang, Daizong Liu, Pan Zhou, Guoshun NanCVPR 2023
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson 等ICCV 2019 · 被引用 258 次
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 被引用 233 次
- Robust Multi-Modality Multi-Object TrackingWenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang 等ICCV 2019 · 被引用 221 次
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 被引用 207 次
相关 Paper
- TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationZeyu Wang, Fabien Colonnier, Jinghong Zheng, Jyotibdha Acharya 等ACM MM 2023 · 被引用 28 次
- BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-IdentificationHaoxuan Xu, Guanglin NiuCVPR 2026 · 被引用 3 次
- Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image FusionYanglin Deng, Tianyang Xu, Chunyang Cheng, Hui Li 等CVPR 2026
- End-to-End RGB-D Image Compression via Exploiting Channel-Modality RedundancyHuiming Zheng, Wei GaoAAAI 2024 · 被引用 15 次
- Infrared-Privileged UAV Detection via Cross-Modal Vector-QuantizationZhibo Lou, Ruijie Zhang, Zeyu Luo, Qianxi Cao 等AAAI 2026
