Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach
Xingyu Li, Chen Gong, Guohong Fu
Abstract
Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content. In the era of rapidly growing multimodal content and social media, MCR is particularly crucial for interpreting user interactions and bridging text-visual references to improve communication and personalization. However, MCR research for real-world dialogues remains unexplored due to the lack of sufficient data resources. To address this gap, we introduce TikTalkCoref, the first Chinese multimodal coreference dataset for social media in real-world scenarios, derived from the popular Douyin short-video platform. This dataset pairs short videos with corresponding textual dialogues from user comments and includes manually annotated coreference clusters for both person mentions in the text and the coreferential person head regions in the corresponding video frames. We also present an effective benchmark approach for MCR, focusing on the celebrity domain, and conduct extensive experiments on our dataset, providing reliable benchmark results for this newly constructed dataset. We release the TikTalk-Coref dataset to facilitate future research on MCR for real-world social media dialogues at https://github.com/lxystaruni/TikTalkCoref .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b4d278f-7ce0-4c22-bcae-acb9b4b60fe0Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal ConversationsSatwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak DamavandiEMNLP 2021 · 54 citations
- CCMB: A Large-scale Chinese Cross-modal BenchmarkChunyu Xie, Heng Cai, Jincheng Li, Fanjing Kong et al.ACM MM 2023 · 10 citations
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu et al.ACM MM 2023 · 10 citations
Related papers
- A Chinese Multimodal Social Video Dataset for Controversy DetectionTianjiao Xu, Aoxuan Chen, Yuxi Zhao, Jinfei Gao et al.ACM MM 2024 · 4 citations
- Short Video Ordering via Position Decoding and Successor PredictionShiping Ge, Qiang Chen, Zhiwei Jiang, Yafeng Yin et al.SIGIR 2024 · 1 citation
- Tencent-MVSE: A Large-Scale Benchmark Dataset for Multi-Modal Video Similarity EvaluationZhaoyang Zeng, Yongsheng Luo, Zhenhua Liu, Fengyun Rao et al.CVPR 2022 · 5 citations
- Who are you referring to? Coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenICCV 2023 · 8 citations
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai et al.ACM MM 2021 · 134 citations
