Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention
Katsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro Okada
摘要
Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions, and a method for the newly proposed task that effectively utilizes multi-modal data of video and motion sensors, or inertial measurement units (IMUs). While conventional video captioning tasks have difficulty in dealing with detailed descriptions of human activities due to the limited view of a fixed camera, egocentric vision has greater potential to be used for generating the finer-grained descriptions of human activities on the basis of a much closer view. In addition, we utilize wearable-sensor data as auxiliary information to mitigate the inherent problems in egocentric vision: motion blur, self-occlusion, and out-of-camera-range activities. We propose a method for effectively utilizing the sensor data in combination with the video data on the basis of an attention mechanism that dynamically determines the modality that requires more attention, taking the contextual information into account. We compared the proposed sensor-fusion method with strong baselines on the MMAC Captions dataset and found that using sensor data as supplementary information to the egocentric-video data was beneficial, and that our proposed method outperformed the strong baselines, demonstrating the effectiveness of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Retrieval-Augmented Egocentric Video CaptioningJilan Xu, Yifei Huang, Junlin Hou, Guo Chen 等CVPR 2024 · 被引用 16 次
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View FusionXiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu 等CVPR 2026 · 被引用 1 次
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event PredictionQile Su, Shoutai Zhu, Shuai Zhang, Baoyu Liang 等ACM MM 2025
- Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationZhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 等ACM MM 2025
它引用的顶会 Paper13
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 被引用 186 次
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu 等ACL 2020 · 被引用 168 次
- Ego-Pose Estimation and Forecasting As Real-Time PD ControlYe Yuan, Kris KitaniICCV 2019 · 被引用 147 次
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt 等ICCV 2019 · 被引用 108 次
相关 Paper
- Towards Fine-Grained Human Motion Video CaptioningGuorui Song, Guocun Wang, Zhe Huang, Jing Lin 等ACM MM 2025
- Multi-Perspective Video CaptioningYi Bin, Xindi Shang, Bo Peng, Yujuan Ding 等ACM MM 2021 · 被引用 14 次
- EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUsZhen Fan, Peng Dai, Zhuo Su, Xu Gao 等AAAI 2025 · 被引用 13 次
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 被引用 50 次
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla 等CVPR 2024
