Situation and Behavior Understanding by Trope Detection on Films
Chen-Hsi Chang, Hung-Ting Su, Juiheng Hsu, Yu-Siang Wang, Yu-Cheng Chang, Zhe Yu Liu, Ya-Liang Chang, Wen-Feng Cheng, Ke-Jyun Wang, Winston H. Hsu
Abstract
The human ability of deep cognitive skills is crucial for the development of various real-world applications that process diverse and abundant user generated input. While recent progress of deep learning and natural language processing have enabled learning system to reach human performance on some benchmarks requiring shallow semantics, such human ability still remains challenging for even modern contextual embedding models, as pointed out by many recent studies [9, 10, 22, 24, 32] . Existing machine comprehension datasets assume sentence-level input, lack of casual or motivational inferences, or can be answered with question-answer bias. Here, we present a challenging novel task, trope detection on films, in an effort to create a situation and behavior understanding for machines. Tropes are frequently used storytelling devices for creative works. Comparing to existing movie tag prediction tasks, tropes are more sophisticated as they can vary widely, from a moral concept to a series of circumstances, and embedded with motivations and cause-and-effects. We introduce a new dataset, Tropes in Movie Synopses (TiMoS), with 5623 movie synopses and 95 different tropes collecting from a Wikipedia-style database, TVTropes. We present a multi-stream comprehension network (MulCom) leveraging multi-level attention of words, sentences, and role relations. Experimental result demonstrates that modern models including BERT contextual embedding, movie tag prediction systems, and relational networks, perform at most 37% of human performance (23.97/64.87) in terms of F1 score. Our MulCom outperforms all modern baselines, by 1.5 to 5.0 F1 score and 1.5 to 3.0 mean of average precision (mAP) score. We also provide a detailed analysis and human evaluation to pave ways for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75611e62-e08b-4ef5-b605-e7af72be750dCited by top-tier papers2
- TaleStream: Supporting Story Ideation with Trope KnowledgeJean-Peïc Chou, Alexa Fay Siu, Nedim Lipka, Ryan A. Rossi et al.UIST 2023 · 15 citations
- Using Web Data to Reveal 22-Year History of Sneaker DesignsSungkyu Park, Hyeonho Song, Sungwon Han, Berhane Weldegebriel et al.WWW 2022 · 2 citations
Builds on3
- To Test Machine Comprehension, Start by Defining ComprehensionJesse Dunietz, Gregory Burnham, Akash Bharadwaj, Owen Rambow et al.ACL 2020 · 7 citations
- Probing Linguistic SystematicityEmily Goodwin, Koustuv Sinha, Timothy J. O'DonnellACL 2020 · 4 citations
- Violin: A Large-Scale Dataset for Video-and-Language InferenceJingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan et al.CVPR 2020
Related papers
- MetaGPT: A Large Vision-Language Model for Meme Metaphor UnderstandingBo Xu, Chenyuan Wang, Xinyu Chen, Hongfei Lin et al.AAAI 2026
- ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsRujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon et al.EMNLP 2021 · 30 citations
- Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-MindMo Yu, Qiujing Wang, Shunchi Zhang, Yisi Sang et al.ICML 2024 · 22 citations
- MemeQA: Holistic Evaluation for Meme UnderstandingKhoi P. N. Nguyen, Terrence Li, Derek Lou Zhou, Gabriel Xiong et al.ACL 2025 · 3 citations
- Explaining (Sarcastic) Utterances to Enhance Affect Understanding in Multimodal DialoguesShivani Kumar, Ishani Mondal, Md. Shad Akhtar, Tanmoy ChakrabortyAAAI 2023 · 22 citations
