Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative Learning
Yi Bin, Junrong Liao, Yujuan Ding, Haoxuan Li, Yang Yang, See-Kiong Ng, Heng Tao Shen
摘要
Cross-modal coherence modeling is essential for intelligent systems to help them organize and structure information, thereby understanding and creating content of the physical world coherently like human-beings. Previous work on cross-modal coherence modeling attempted to leverage the order information from another modality to assist the coherence recovering of the target modality. Despite of the effectiveness, labeled associated coherency information is not always available and might be costly to acquire, making the cross-modal guidance hard to leverage. To tackle this challenge, this paper explores a new way to take advantage of cross-modal guidance without gold labels on coherency, and proposes the Weak Cross-Modal Guided Ordering (WeGO) model. More specifically, it leverages high-confidence predicted pairwise order in one modality as reference information to guide the coherence modeling in another. An iterative learning paradigm is further designed to jointly optimize the coherence modeling in two modalities with selected guidance from each other. The iterative cross-modal boosting also functions in inference to further enhance coherence prediction in each modality. Experimental results on two public datasets have demonstrated that the proposed method outperforms existing methods for cross-modal coherence modeling tasks. Major technical modules have been evaluated effective through ablation studies. Codes are available at: https://github.com/scvready123/IterWeGO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- CyCLIP: Cyclic Contrastive Language-Image PretrainingShashank Goel, Hritik Bansal, Sumit Bhatia, Ryan A. Rossi 等NeurIPS 2022 · 被引用 192 次
- PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model PretrainingYuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang 等NeurIPS 2022 · 被引用 168 次
- Joint Multi-modal Aspect-Sentiment Analysis with Auxiliary Cross-modal Relation DetectionXincheng Ju, Dong Zhang, Rong Xiao, Junhui Li 等EMNLP 2021 · 被引用 130 次
相关 Paper
- Non-Autoregressive Cross-Modal Coherence ModellingYi Bin, Wenhao Shi, Jipeng Zhang, Yujuan Ding 等ACM MM 2022 · 被引用 12 次
- Cross-Modal Coherence for Text-to-Image RetrievalMalihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia 等AAAI 2022 · 被引用 11 次
- Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion RecognitionWenjue He, Xiaofeng Zhu, Zheng ZhangAAAI 2026 · 被引用 1 次
- An Anchor-based Relative Position Embedding Method for Cross-Modal TasksYa Wang, Xingwu Sun, Fengzong Lian, Zhanhui Kang 等EMNLP 2022 · 被引用 1 次
- Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency LearningYafei Zhang, Lingqi Kong, Huafeng Li, Jie WenICCV 2025 · 被引用 10 次
