Non-Autoregressive Cross-Modal Coherence Modelling
Yi Bin, Wenhao Shi, Jipeng Zhang, Yujuan Ding, Yang Yang, Heng Tao Shen
Abstract
Modelling the coherence of information is important for human to perceive and prehend the physical world. Existing works on coherence modelling mainly focus on single modality, which overlook the effect of information integration and semantic consistency across modalities. To fill the research gap, this paper targets at the cross-modal coherence modelling, specifically, the cross-modal ordering task. The task requires to not only explore the coherence information in single modality, but also leverage cross-modal information to model the semantic consistency between modalities. To this end, we propose a Non-Autoregressive Cross-modal Ordering Net (NACON) adopting a basic encoder-decoder architecture. Specifically, NACON is equipped with an order-invariant context encoder to model the unordered input set and a non-autoregressive decoder to generate ordered sequences in parallel. We devise a cross-modal positional attention module in NACON to take advantage of the cross-modal order guidance. To alleviate the repetition problem of non-autoregressive models, we introduce an elegant exclusive loss to constrain the ordering exclusiveness between positions and elements. We conduct extensive experiments on two assembled datasets to support our task, SIND and TACoS-Ordering. Experimental results show that the proposed NACON can effectively leverage cross-modal guidance and recover the correct order of the elements.The code is available at https://github.com/YiBin-CHN/CMCM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative EliminationHaoxuan Li, Yi Bin, Junrong Liao, Yang Yang et al.ACM MM 2023 · 42 citations
- GalleryGPT: Analyzing Paintings with Large Multimodal ModelsYi Bin, Wenhao Shi, Yujuan Ding, Zhiqiang Hu et al.ACM MM 2024 · 35 citations
- Leveraging Weak Cross-Modal Guidance for Coherence Modelling via Iterative LearningYi Bin, Junrong Liao, Yujuan Ding, Haoxuan Li et al.ACM MM 2024 · 3 citations
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu et al.SIGIR 2026
Related papers
- An Anchor-based Relative Position Embedding Method for Cross-Modal TasksYa Wang, Xingwu Sun, Fengzong Lian, Zhanhui Kang et al.EMNLP 2022 · 1 citation
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 124 citations
- Order-Agnostic Cross Entropy for Non-Autoregressive Machine TranslationCunxiao Du, Zhaopeng Tu, Jing JiangICML 2021 · 93 citations
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics PriorityHengqi Liu, Wanting Zhou, Longteng Kong, Fangxiang Feng et al.CVPR 2026
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
