UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog System
Zhiyuan Ma, Jianjun Li, Guohui Li, Yongjing Cheng
Abstract
As a more natural and intelligent interaction manner, multimodal task-oriented dialog system recently has received great attention and many remarkable progresses have been achieved. Nevertheless, almost all existing studies follow the pipeline to first learn intra-modal features separately and then conduct simple feature concatenation or attention-based feature fusion to generate responses, which hampers them from learning inter-modal interactions and conducting cross-modal feature alignment for generating more intention-aware responses. To address these issues, we propose UniTranSeR, a Unified Transformer Semantic Representation framework with feature alignment and intention reasoning for multimodal dialog systems. Specifically, we first embed the multimodal features into a unified Transformer semantic space to prompt inter-modal interactions, and then devise a feature alignment and intention reasoning (FAIR) layer to perform cross-modal entity alignment and fine-grained key-value reasoning, so as to effectively identify user’s intention for generating more accurate responses. Experimental results verify the effectiveness of UniTranSeR, showing that it significantly outperforms state-of-the-art approaches on the representative MMD dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dba14235-cac7-40aa-92cb-860376e06f62Cited by top-tier papers7
- Generative Multi-Modal Knowledge Retrieval with Large Language ModelsXinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma et al.AAAI 2024 · 35 citations
- Neural Residual Diffusion Models for Deep Scalable Vision GenerationZhiyuan Ma, Liangliang Zhao, Biqing Qi, Bowen ZhouNeurIPS 2024 · 15 citations
- LMD: Faster Image Reconstruction with Latent Masking DiffusionZhiyuan Ma, Zhihuan Yu, Jianjun Li, Bowen ZhouAAAI 2024 · 15 citations
- Dual Semantic Knowledge Composed Multimodal Dialog SystemsXiaolin Chen, Xuemeng Song, Yinwei Wei, Liqiang Nie et al.SIGIR 2023 · 8 citations
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 8 citations
Builds on11
- Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented DialogLibo Qin, Xiao Xu, Wanxiang Che, Yue Zhang et al.ACL 2020 · 90 citations
- Continual Learning in Task-Oriented Dialogue SystemsAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon et al.EMNLP 2021 · 68 citations
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 48 citations
- Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic ElementsWeidong He, Zhi Li, Dongcai Lu, Enhong Chen et al.ACM MM 2020 · 31 citations
- Multimodal Dialog System: Relational Graph-based Context-aware Question UnderstandingHaoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei et al.ACM MM 2021 · 31 citations
Related papers
- MaTCR: Modality-Aligned Thought Chain Reasoning for Multimodal Task-Oriented Dialogue GenerationYiting Liu, Liang Li, Beichen Zhang, Shan Huang et al.ACM MM 2023 · 3 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
- Intention Reasoning Network for Multi-Domain End-to-end Task-Oriented DialogueZhiyuan Ma, Jianjun Li, Zezheng Zhang, Guohui Li et al.EMNLP 2021 · 5 citations
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li et al.ICDE 2024 · 14 citations
