Visual Redundancy Removal of Composite Images via Multimodal Learning
Wuyuan Xie, Shukang Wang, Rong Zhang, Miaohui Wang
Abstract
Composite images are generated by combining two or more different photographs, and their content is typically heterogeneous. However, existing unimodal visual redundancy prediction methods are difficult to accurately model the complex characteristics of this image type. In this paper, we investigate the visual redundancy modeling of composite images from an end-to-end multimodal perspective, including four cross-media modalities (i.e., text, brightness, color, and segmentation). Specifically, we design a two-stage cross-modal alignment module based on self-attention mechanism and contrastive learning, and develop a fusion module based on a cross-modal augmentation paradigm. Further, we establish the first cross-media visual redundancy dataset for composite images, which contains 413 groups of cross-modal data and generates 13629 realistic compression distortions using the latest versatile video coding (VVC) standard. Experimental results on nine benchmark datasets demonstrate the effectiveness of our method, outperforming seven representative methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- DDJND: Dual Domain Just Noticeable Difference in Multi-Source Content Images with Structural DiscrepancyMiaohui Wang, Zhenming Li, Wuyuan XieAAAI 2025 · 2 citations
- The Last Byte: Learning Just Enough for Machine-Oriented Image CompressionWuyuan Xie, Zhenming Li, Ye Liu, Jian Jin et al.AAAI 2026
- Visual Redundancy Removal for Composite Images: A Benchmark Dataset and a Multi-Visual-Effects Driven Incremental MethodMiaohui Wang, Rong Zhang, Lirong Huang, Yanshan LiAAAI 2024
Related papers
- Heterogeneous Feature Fusion and Cross-modal Alignment for Composed Image RetrievalGangjian Zhang, Shikui Wei, Huaxin Pang, Yao ZhaoACM MM 2021 · 34 citations
- Cross Modal Compression: Towards Human-comprehensible Semantic CompressionJiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma et al.ACM MM 2021 · 24 citations
- Learning based Multi-modality Image and Video CompressionGuo Lu, Tianxiong Zhong, Jing Geng, Qiang Hu et al.CVPR 2022 · 26 citations
- End-to-End RGB-D Image Compression via Exploiting Channel-Modality RedundancyHuiming Zheng, Wei GaoAAAI 2024 · 15 citations
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
