Teaching Human Behavior Improves Content Understanding Abilities Of VLMs
Somesh Kumar Singh, Harini S. I, Yaman Kumar Singla, Changyou Chen, Rajiv Ratn Shah, Veeky Baths, Balaji Krishnamurthy
Abstract
Communication is defined as "Who says what to whom with what effect." A message from a communicator generates downstream receiver effects, also known as behavior. Receiver behavior, being a downstream effect of the message, carries rich signals about it. Even after carrying signals about the message, the behavior signal is often ignored while training vision language models. We show that training VLMs on receiver behavior can actually help improve their content-understanding abilities. We demonstrate that training VLMs to predict receiver behaviors, such as likes, comments, and replay graphs, which are available at scale, enhances the VLM's performance across a broad range of downstream content understanding tasks. We show this performance increase over 6 types of behavior, 46 different tasks covering image, video, text and audio over 26 benchmark datasets across both 0-shot and fine-tuning settings, outperforming many supervised baselines on diverse tasks ranging from emotion recognition to captioning by upto 150%. We note that since receiver behavior, such as likes, comments, and replay graphs, is collected by default on the internet and does not need any human annotations to be useful, the performance improvement we get after training on this data is essentially free-lunch. We also release BLIFT, our Behaviour-LLaVA IFT dataset comprising 730k images and videos with their receiver behavior collected from multiple platforms on which we train our models to achieve this. The dataset and code are available at behavior-in-the-wild.github.io/behavior-llava. 00:52 That ending put a smile on my face . Won't make me change insurance companies, but it's a great commercial. Sad that State Farm no longer covers us in Bay Area! "Neighbah, Papah, Baaah, Labah, Backstabah"… Lol!! Arnie just can't get that hard R out.! 😂 😂 Temporal Understanding Character Understanding Context Understanding User Opinion Understanding The sheep part was hilarious Cognitive Understanding Channel: YouTube Speaker (Who says): State Farm Insurance Receivers (To whom): YouTube Subscribers (What) Video Content Arnold and Danny are always a beautiful duet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
Related papers
- Large Content And Behavior Models To Understand, Simulate, And Optimize Content And BehaviorAshmit Khandelwal, Aditya Agrawal, Aanisha Bhattacharyya, Yaman Kumar et al.ICLR 2024 · 11 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual QuestionsWenbo Hu, Yifan Xu, Yi Li, Weiyue Li et al.AAAI 2024 · 209 citations
- LLaVAction: evaluating and training multi-modal large language models for action understandingHaozhe Qi, Shaokai Ye, Alexander Mathis, Mackenzie W. MathisICLR 2026 · 3 citations
- Multi-Task Gaze Communication UnderstandingCheng Peng, Oya ÇeliktutanACM MM 2025
