Modular Memorability: Tiered Representations for Video Memorability Prediction
Théo Dumont, Juan Segundo Hevia, Camilo Luciano Fosco
Abstract
The question of how to best estimate the memorability of visual content is currently a source of debate in the memorability community. In this paper, we propose to explore how different key properties of images and videos affect their consolidation into memory. We analyze the impact of several features and develop a model that emulates the most important parts of a proposed "pathway to memory": a simple but effective way of representing the different hurdles that new visual content needs to surpass to stay in memory. This framework leads to the construction of our M3-S model, a novel memorability network that processes input videos in a modular fashion. Each module of the network emulates one of the four key steps of the pathway to memory: raw encoding, scene understanding, event understanding and memory consolidation. We find that the different representations learned by our modules are non-trivial and substantially different from each other. Additionally, we observe that certain representations tend to perform better at the task of memorability prediction than others, and we introduce an in-depth ablation study to support our results. Our proposed approach surpasses the state of the art on the two largest video memorability datasets and opens the door to new applications in the field. Our code is available at https://github.com/tekal-ai/modular- memorability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- GANalyze: Toward Visual Definitions of Cognitive Image PropertiesLore Goetschalckx, Alex Andonian, Aude Oliva, Phillip IsolaICCV 2019 · 345 citations
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 285 citations
- VideoMem: Constructing, Analyzing, Predicting Short-Term and Long-Term Video MemorabilityRomain Cohendet, Claire-Hélène Demarty, Ngoc Q. K. Duong, Martin EngilbergeICCV 2019 · 56 citations
Related papers
- Predicting Event Memorability from Contextual Visual SemanticsQianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju et al.NeurIPS 2021 · 6 citations
- Video Entailment via Reaching a Structure-Aware Cross-modal ConsensusXuan Yao, Junyu Gao, Mengyuan Chen, Changsheng XuACM MM 2023 · 4 citations
- Visual Consensus Modeling for Video-Text RetrievalShuqiang Cao, Bairui Wang, Wei Zhang, Lin MaAAAI 2022 · 24 citations
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji et al.ACM MM 2021 · 41 citations
- Knowledge-based Temporal Fusion Network for Interpretable Online Video Popularity PredictionShisong Tang, Qing Li, Xiaoteng Ma, Ci Gao et al.WWW 2022 · 29 citations
