VideoMem: Constructing, Analyzing, Predicting Short-Term and Long-Term Video Memorability
Romain Cohendet, Claire-Hélène Demarty, Ngoc Q. K. Duong, Martin Engilberge
摘要
Humans share a strong tendency to memorize/forget some of the visual information they encounter. This paper focuses on providing computational models for the prediction of the intrinsic memorability of visual content. To address this new challenge, we introduce a large scale dataset (VideoMem) composed of 10,000 videos annotated with memorability scores. In contrast to previous work on image memorability -where memorability was measured a few minutes after memorization -memory performance is measured twice: a few minutes after memorization and again 24-72 hours later. Hence, the dataset comes with short-term and long-term memorability annotations. After an in-depth analysis of the dataset, we investigate several deep neural network based models for the prediction of video memorability. Our best model using a ranking loss achieves a Spearman's rank correlation of 0.494 for short-term memorability prediction, while our proposed model with attention mechanism provides insights of what makes a content memorable. The VideoMem dataset with pre-extracted features is publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video UnderstandingMinghui Wu, Chenxu Zhao, Anyang Su, Donglin Di 等ACM MM 2024 · 被引用 8 次
- Predicting Event Memorability from Contextual Visual SemanticsQianli Xu, Fen Fang, Ana Garcia del Molino, Vigneshwaran Subbaraju 等NeurIPS 2021 · 被引用 6 次
- How to Take a Memorable Picture? Empowering Users with Actionable FeedbackFrancesco Laiti, Davide Talon, Jacopo Staiano, Elisa RicciCVPR 2026
- Teaching Human Behavior Improves Content Understanding Abilities Of VLMsSomesh Kumar Singh, Harini S. I, Yaman Kumar Singla, Changyou Chen 等ICLR 2025
- Modular Memorability: Tiered Representations for Video Memorability PredictionThéo Dumont, Juan Segundo Hevia, Camilo Luciano FoscoCVPR 2023
相关 Paper
- Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingJingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang 等NeurIPS 2025 · 被引用 25 次
- Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming VideoHongxiang Jiang, Zengrui Ge, Guo Chen, Qixiong Wang 等ICLR 2026
- FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality AssessmentSijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao 等ACM MM 2025 · 被引用 8 次
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan 等CVPR 2022 · 被引用 158 次
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
