Visual Redundancy Removal of Composite Images via Multimodal Learning
Wuyuan Xie, Shukang Wang, Rong Zhang, Miaohui Wang
摘要
Composite images are generated by combining two or more different photographs, and their content is typically heterogeneous. However, existing unimodal visual redundancy prediction methods are difficult to accurately model the complex characteristics of this image type. In this paper, we investigate the visual redundancy modeling of composite images from an end-to-end multimodal perspective, including four cross-media modalities (i.e., text, brightness, color, and segmentation). Specifically, we design a two-stage cross-modal alignment module based on self-attention mechanism and contrastive learning, and develop a fusion module based on a cross-modal augmentation paradigm. Further, we establish the first cross-media visual redundancy dataset for composite images, which contains 413 groups of cross-modal data and generates 13629 realistic compression distortions using the latest versatile video coding (VVC) standard. Experimental results on nine benchmark datasets demonstrate the effectiveness of our method, outperforming seven representative methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- DDJND: Dual Domain Just Noticeable Difference in Multi-Source Content Images with Structural DiscrepancyMiaohui Wang, Zhenming Li, Wuyuan XieAAAI 2025 · 被引用 2 次
- The Last Byte: Learning Just Enough for Machine-Oriented Image CompressionWuyuan Xie, Zhenming Li, Ye Liu, Jian Jin 等AAAI 2026
- Visual Redundancy Removal for Composite Images: A Benchmark Dataset and a Multi-Visual-Effects Driven Incremental MethodMiaohui Wang, Rong Zhang, Lirong Huang, Yanshan LiAAAI 2024
相关 Paper
- Heterogeneous Feature Fusion and Cross-modal Alignment for Composed Image RetrievalGangjian Zhang, Shikui Wei, Huaxin Pang, Yao ZhaoACM MM 2021 · 被引用 34 次
- Cross Modal Compression: Towards Human-comprehensible Semantic CompressionJiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma 等ACM MM 2021 · 被引用 24 次
- Learning based Multi-modality Image and Video CompressionGuo Lu, Tianxiong Zhong, Jing Geng, Qiang Hu 等CVPR 2022 · 被引用 26 次
- End-to-End RGB-D Image Compression via Exploiting Channel-Modality RedundancyHuiming Zheng, Wei GaoAAAI 2024 · 被引用 15 次
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui 等CVPR 2023
