Situation and Behavior Understanding by Trope Detection on Films
Chen-Hsi Chang, Hung-Ting Su, Juiheng Hsu, Yu-Siang Wang, Yu-Cheng Chang, Zhe Yu Liu, Ya-Liang Chang, Wen-Feng Cheng, Ke-Jyun Wang, Winston H. Hsu
摘要
The human ability of deep cognitive skills is crucial for the development of various real-world applications that process diverse and abundant user generated input. While recent progress of deep learning and natural language processing have enabled learning system to reach human performance on some benchmarks requiring shallow semantics, such human ability still remains challenging for even modern contextual embedding models, as pointed out by many recent studies [9, 10, 22, 24, 32] . Existing machine comprehension datasets assume sentence-level input, lack of casual or motivational inferences, or can be answered with question-answer bias. Here, we present a challenging novel task, trope detection on films, in an effort to create a situation and behavior understanding for machines. Tropes are frequently used storytelling devices for creative works. Comparing to existing movie tag prediction tasks, tropes are more sophisticated as they can vary widely, from a moral concept to a series of circumstances, and embedded with motivations and cause-and-effects. We introduce a new dataset, Tropes in Movie Synopses (TiMoS), with 5623 movie synopses and 95 different tropes collecting from a Wikipedia-style database, TVTropes. We present a multi-stream comprehension network (MulCom) leveraging multi-level attention of words, sentences, and role relations. Experimental result demonstrates that modern models including BERT contextual embedding, movie tag prediction systems, and relational networks, perform at most 37% of human performance (23.97/64.87) in terms of F1 score. Our MulCom outperforms all modern baselines, by 1.5 to 5.0 F1 score and 1.5 to 3.0 mean of average precision (mAP) score. We also provide a detailed analysis and human evaluation to pave ways for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TaleStream: Supporting Story Ideation with Trope KnowledgeJean-Peïc Chou, Alexa Fay Siu, Nedim Lipka, Ryan A. Rossi 等UIST 2023 · 被引用 15 次
- Using Web Data to Reveal 22-Year History of Sneaker DesignsSungkyu Park, Hyeonho Song, Sungwon Han, Berhane Weldegebriel 等WWW 2022 · 被引用 2 次
它引用的顶会 Paper3
- To Test Machine Comprehension, Start by Defining ComprehensionJesse Dunietz, Gregory Burnham, Akash Bharadwaj, Owen Rambow 等ACL 2020 · 被引用 7 次
- Probing Linguistic SystematicityEmily Goodwin, Koustuv Sinha, Timothy J. O'DonnellACL 2020 · 被引用 4 次
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan 等CVPR 2020
相关 Paper
- MetaGPT: A Large Vision-Language Model for Meme Metaphor UnderstandingBo Xu, Chenyuan Wang, Xinyu Chen, Hongfei Lin 等AAAI 2026
- ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsRujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon 等EMNLP 2021 · 被引用 30 次
- Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-MindMo Yu, Qiujing Wang, Shunchi Zhang, Yisi Sang 等ICML 2024 · 被引用 22 次
- MemeQA: Holistic Evaluation for Meme UnderstandingKhoi P. N. Nguyen, Terrence Li, Derek Lou Zhou, Gabriel Xiong 等ACL 2025 · 被引用 3 次
- Explaining (Sarcastic) Utterances to Enhance Affect Understanding in Multimodal DialoguesShivani Kumar, Ishani Mondal, Md. Shad Akhtar, Tanmoy ChakrabortyAAAI 2023 · 被引用 22 次
