Teaching Human Behavior Improves Content Understanding Abilities Of VLMs
Somesh Kumar Singh, Harini S. I, Yaman Kumar Singla, Changyou Chen, Rajiv Ratn Shah, Veeky Baths, Balaji Krishnamurthy
摘要
Communication is defined as "Who says what to whom with what effect." A message from a communicator generates downstream receiver effects, also known as behavior. Receiver behavior, being a downstream effect of the message, carries rich signals about it. Even after carrying signals about the message, the behavior signal is often ignored while training vision language models. We show that training VLMs on receiver behavior can actually help improve their content-understanding abilities. We demonstrate that training VLMs to predict receiver behaviors, such as likes, comments, and replay graphs, which are available at scale, enhances the VLM's performance across a broad range of downstream content understanding tasks. We show this performance increase over 6 types of behavior, 46 different tasks covering image, video, text and audio over 26 benchmark datasets across both 0-shot and fine-tuning settings, outperforming many supervised baselines on diverse tasks ranging from emotion recognition to captioning by upto 150%. We note that since receiver behavior, such as likes, comments, and replay graphs, is collected by default on the internet and does not need any human annotations to be useful, the performance improvement we get after training on this data is essentially free-lunch. We also release BLIFT, our Behaviour-LLaVA IFT dataset comprising 730k images and videos with their receiver behavior collected from multiple platforms on which we train our models to achieve this. The dataset and code are available at behavior-in-the-wild.github.io/behavior-llava. 00:52 That ending put a smile on my face . Won't make me change insurance companies, but it's a great commercial. Sad that State Farm no longer covers us in Bay Area! "Neighbah, Papah, Baaah, Labah, Backstabah"… Lol!! Arnie just can't get that hard R out.! 😂 😂 Temporal Understanding Character Understanding Context Understanding User Opinion Understanding The sheep part was hilarious Cognitive Understanding Channel: YouTube Speaker (Who says): State Farm Insurance Receivers (To whom): YouTube Subscribers (What) Video Content Arnold and Danny are always a beautiful duet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
相关 Paper
- Large Content And Behavior Models To Understand, Simulate, And Optimize Content And BehaviorAshmit Khandelwal, Aditya Agrawal, Aanisha Bhattacharyya, Yaman Kumar 等ICLR 2024 · 被引用 11 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual QuestionsWenbo Hu, Yifan Xu, Yi Li, Weiyue Li 等AAAI 2024 · 被引用 209 次
- LLaVAction: evaluating and training multi-modal large language models for action understandingHaozhe Qi, Shaokai Ye, Alexander Mathis, Mackenzie W. MathisICLR 2026 · 被引用 3 次
- Multi-Task Gaze Communication UnderstandingCheng Peng, Oya ÇeliktutanACM MM 2025
