Learning Human-Human Interactions in Images from Weak Textual Supervision
Morris Alper, Hadar Averbuch-Elor
Abstract
Interactions between humans are diverse and context-dependent, but previous works have treated them as categorical, disregarding the heavy tail of possible interactions. We propose a new paradigm of learning human-human interactions as free text from a single still image, allowing for flexibility in modeling the unlimited space of situations and relationships between people. To overcome the absence of data labelled specifically for this task, we use knowledge distillation applied to synthetic caption data produced by a large language model without explicit supervision. We show that the pseudo-labels produced by this procedure can be used to train a captioning model to effectively understand human-human interactions in images, as measured by a variety of metrics that measure textual and semantic faithfulness and factual groundedness of our predictions. We further show that our approach outperforms SOTA image captioning and situation recognition models on this task. We will releas 1 our code and pseudo-labels along with Waldo and Wenda, a manually-curated test set for still image human-human interaction understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- LiT: Zero-Shot Transfer with Locked-image text TuningXiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner et al.CVPR 2022 · 349 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
Related papers
- Who's Waldo? Linking People Across Text and ImagesClaire Yuqing Cui, Apoorv Khandelwal, Yoav Artzi, Noah Snavely et al.ICCV 2021 · 21 citations
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 5 citations
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image ModelsVatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang et al.EMNLP 2025
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 115 citations
- Zero-Shot Referring Expression Comprehension via Structural Similarity Between Images and CaptionsZeyu Han, Fangrui Zhu, Qianru Lao, Huaizu JiangCVPR 2024
