Temporal Sentiment Localization: Listen and Look in Untrimmed Videos
Zhicheng Zhang, Jufeng Yang
摘要
Video sentiment analysis aims to uncover the underlying attitudes of viewers, which has a wide range of applications in real world. Existing works simply classify a video into a single sentimental category, ignoring the fact that sentiment in untrimmed videos may appear in multiple segments with varying lengths and unknown locations. To address this, we propose a challenging task, i.e., Temporal Sentiment Localization (TSL), to find which parts of the video convey sentiment. To systematically investigate fully- and weakly-supervised settings for TSL, we first build a benchmark dataset named TSL-300, which is consisting of 300 videos with a total length of 1,291 minutes. Each video is labeled in two ways, one of which is frame-by-frame annotation for the fully-supervised setting, and the other is single-frame annotation, i.e., only a single frame with strong sentiment is labeled per segment for the weakly-supervised setting. Due to the high cost of labeling a densely annotated dataset, we propose TSL-Net in this work, employing single-frame supervision to localize sentiment in videos. In detail, we generate the pseudo labels for unlabeled frames using a greedy search strategy, and fuse the affective features of both visual and audio modalities to predict the temporal sentiment distribution. Here, a reverse mapping strategy is designed for feature fusion, and a contrastive loss is utilized to maintain the consistency between the original feature and the reverse prediction. Extensive experiments show the superiority of our method against the state-of-the-art approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi 等CVPR 2024 · 被引用 137 次
- ExtDM: Distribution Extrapolation Diffusion Model for Video PredictionZhicheng Zhang, Junyao Hu, Wentao Cheng, Danda Pani Paudel 等CVPR 2024 · 被引用 24 次
- Ordinal Label Distribution LearningChangsong Wen, Xin Zhang, Xingxu Yao, Jufeng YangICCV 2023 · 被引用 22 次
- LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented DiffusionPancheng Zhao, Peng Xu, Pengda Qin, Deng-Ping Fan 等CVPR 2024 · 被引用 14 次
- Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLMJunxiao Ma, Jingjing Wang, Jiamin Luo, Peiying Yu 等WWW 2025 · 被引用 11 次
它引用的顶会 Paper15
- CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of ModalityWenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu 等ACL 2020 · 被引用 376 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park 等ICCV 2019 · 被引用 285 次
- 3C-Net: Category Count and Center Loss for Weakly-Supervised Action LocalizationSanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, Ling ShaoICCV 2019 · 被引用 174 次
- Weakly-supervised Temporal Action Localization by Uncertainty ModelingPilhyeon Lee, Jinglu Wang, Yan Lu, Hyeran ByunAAAI 2021 · 被引用 141 次
相关 Paper
- Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment LocalizationCailing Han, Zhangbin Li, Jinxing Zhou, Wei Qian 等CVPR 2026 · 被引用 1 次
- Representation Learning through Multimodal Attention and Time-Sync Comments for Affective Video Content AnalysisJicai Pan, Shangfei Wang, Lin FangACM MM 2022 · 被引用 18 次
- Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in VideosJuncheng Li, Junlin Xie, Linchao Zhu, Long Qian 等ACM MM 2022 · 被引用 8 次
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao 等AAAI 2026 · 被引用 4 次
- Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksZiyi Liu, Le Wang, Qilin Zhang, Zhanning Gao 等ICCV 2019 · 被引用 122 次
