Modular Memorability: Tiered Representations for Video Memorability Prediction
Théo Dumont, Juan Segundo Hevia, Camilo Luciano Fosco
摘要
The question of how to best estimate the memorability of visual content is currently a source of debate in the memorability community. In this paper, we propose to explore how different key properties of images and videos affect their consolidation into memory. We analyze the impact of several features and develop a model that emulates the most important parts of a proposed "pathway to memory": a simple but effective way of representing the different hurdles that new visual content needs to surpass to stay in memory. This framework leads to the construction of our M3-S model, a novel memorability network that processes input videos in a modular fashion. Each module of the network emulates one of the four key steps of the pathway to memory: raw encoding, scene understanding, event understanding and memory consolidation. We find that the different representations learned by our modules are non-trivial and substantially different from each other. Additionally, we observe that certain representations tend to perform better at the task of memorability prediction than others, and we introduce an in-depth ablation study to support our results. Our proposed approach surpasses the state of the art on the two largest video memorability datasets and opens the door to new applications in the field. Our code is available at https://github.com/tekal-ai/modular- memorability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- GANalyze: Toward Visual Definitions of Cognitive Image PropertiesLore Goetschalckx, Alex Andonian, Aude Oliva, Phillip IsolaICCV 2019 · 被引用 345 次
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 被引用 285 次
- VideoMem: Constructing, Analyzing, Predicting Short-Term and Long-Term Video MemorabilityRomain Cohendet, Claire-Hélène Demarty, Ngoc Q. K. Duong, Martin EngilbergeICCV 2019 · 被引用 56 次
相关 Paper
- Predicting Event Memorability from Contextual Visual SemanticsQianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju 等NeurIPS 2021 · 被引用 6 次
- Video Entailment via Reaching a Structure-Aware Cross-modal ConsensusXuan Yao, Junyu Gao, Mengyuan Chen, Changsheng XuACM MM 2023 · 被引用 4 次
- Visual Consensus Modeling for Video-Text RetrievalShuqiang Cao, Bairui Wang, Wei Zhang, Lin MaAAAI 2022 · 被引用 24 次
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji 等ACM MM 2021 · 被引用 41 次
- Knowledge-based Temporal Fusion Network for Interpretable Online Video Popularity PredictionShisong Tang, Qing Li, Xiaoteng Ma, Ci Gao 等WWW 2022 · 被引用 29 次
