Understanding Human Preferences: Towards More Personalized Video to Text Generation
Yihan Wu, Ruihua Song, Xu Chen, Hao Jiang, Zhao Cao, Jin Yu
Abstract
While previous video to text models have achieved remarkable successes, they mostly focus on how to understand the video contents in a general sense, but fail to capture the human personalized preferences, which is highly demanded for an engaging multimodal chatbots. Different from user modeling in collaborative filtering, there is no other user behaviors in inference as a real-time video stream is coming. In this paper, we formally define the task of personalized video commenting task and design an end-to-end personalized framework for solving this task. In specific, we argue that the personalization for video comment generation can be reflected in two aspects, that is, (1) for the same video, different users may comment on different clips, and (2) for the same clip, different people may also express various opinions with diverse commentary styles. Motivated by these considerations, we design our framework based on two components. The first one is a clip selector, which is responsible for predicting the clips that the user may comment in the video. The second one is a text generator, which aims to produce the comment based on the above predicted clips and the user's preference. In our framework, these two components are optimized in an end-to-end manner to mutually enhance each other, where we design confidence-aware scheduled sampling and iterative inference strategies to solve the problem that the ground truth clips are absent in the inference phase. As the absence of personalized video to text dataset, we collect and release a new dataset for studying this problem. We conduct extensive experiments to demonstrate the effectiveness of our model.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 068a2438-b920-4efa-a478-b2df9939c8abCited by top-tier papers2
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- RAP: Retrieval-Augmented Personalization for Multimodal Large Language ModelsHaoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li et al.CVPR 2025
Related papers
- VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments GenerationWeiying Wang, Jieting Chen, Qin JinACM MM 2020 · 26 citations
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu et al.ACM MM 2023 · 10 citations
- VCMaster: Generating Diverse and Fluent Live Video Comments Based on Multimodal ContextsManman Zhang, Ge Luo, Yuchen Ma, Sheng Li et al.ACM MM 2023 · 3 citations
- A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationBenhui Zhang, Junyu Gao, Yuan YuanACM MM 2024 · 14 citations
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic et al.EMNLP 2022 · 2 citations
