Understanding Human Preferences: Towards More Personalized Video to Text Generation
Yihan Wu, Ruihua Song, Xu Chen, Hao Jiang, Zhao Cao, Jin Yu
摘要
While previous video to text models have achieved remarkable successes, they mostly focus on how to understand the video contents in a general sense, but fail to capture the human personalized preferences, which is highly demanded for an engaging multimodal chatbots. Different from user modeling in collaborative filtering, there is no other user behaviors in inference as a real-time video stream is coming. In this paper, we formally define the task of personalized video commenting task and design an end-to-end personalized framework for solving this task. In specific, we argue that the personalization for video comment generation can be reflected in two aspects, that is, (1) for the same video, different users may comment on different clips, and (2) for the same clip, different people may also express various opinions with diverse commentary styles. Motivated by these considerations, we design our framework based on two components. The first one is a clip selector, which is responsible for predicting the clips that the user may comment in the video. The second one is a text generator, which aims to produce the comment based on the above predicted clips and the user's preference. In our framework, these two components are optimized in an end-to-end manner to mutually enhance each other, where we design confidence-aware scheduled sampling and iterative inference strategies to solve the problem that the ground truth clips are absent in the inference phase. As the absence of personalized video to text dataset, we collect and release a new dataset for studying this problem. We conduct extensive experiments to demonstrate the effectiveness of our model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- RAP: Retrieval-Augmented Personalization for Multimodal Large Language ModelsHaoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li 等CVPR 2025
相关 Paper
- VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments GenerationWeiying Wang, Jieting Chen, Qin JinACM MM 2020 · 被引用 26 次
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu 等ACM MM 2023 · 被引用 10 次
- VCMaster: Generating Diverse and Fluent Live Video Comments Based on Multimodal ContextsManman Zhang, Ge Luo, Yuchen Ma, Sheng Li 等ACM MM 2023 · 被引用 3 次
- A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationBenhui Zhang, Junyu Gao, Yuan YuanACM MM 2024 · 被引用 14 次
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic 等EMNLP 2022 · 被引用 2 次
