WhyAct: Identifying Action Reasons in Lifestyle Vlogs
Oana Ignat, Santiago Castro, Hanwen Miao, Weiji Li, Rada Mihalcea
Abstract
We aim to automatically identify human action reasons in online videos. We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them. We introduce and make publicly available the WhyAct dataset, consisting of 1,077 visual actions manually annotated with their reasons. We describe a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f3bc181-636c-4dd9-a04e-d766911560eaBuilds on6
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- Video2Commonsense: Generating Commonsense Descriptions to Enrich Video CaptioningZhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral et al.EMNLP 2020 · 61 citations
- GLUCOSE: GeneraLized and COntextualized Story ExplanationsNasrin Mostafazadeh, Aditya Kalyanpur, Lori Moon, David W. Buchanan et al.EMNLP 2020 · 7 citations
Related papers
- YTCommentQA: Video Question Answerability in Instructional VideosSaelyne Yang, Sunghyun Park, Yunseok Jang, Moontae LeeAAAI 2024 · 6 citations
- Oops! Predicting Unintentional Action in VideoDave Epstein, Boyuan Chen, Carl VondrickCVPR 2020
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect TimesOlga Loginova, Sofía Ortega LoguinovaACL 2025
- Multi-modal Action Chain Abductive ReasoningMengze Li, Tianbao Wang, Jiahe Xu, Kairong Han et al.ACL 2023 · 11 citations
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual ScenariosLu Zhu, Tiantian Geng, Yangye Chen, Teng Wang et al.AAAI 2026 · 1 citation
