Learning Human-Human Interactions in Images from Weak Textual Supervision
Morris Alper, Hadar Averbuch-Elor
摘要
Interactions between humans are diverse and context-dependent, but previous works have treated them as categorical, disregarding the heavy tail of possible interactions. We propose a new paradigm of learning human-human interactions as free text from a single still image, allowing for flexibility in modeling the unlimited space of situations and relationships between people. To overcome the absence of data labelled specifically for this task, we use knowledge distillation applied to synthetic caption data produced by a large language model without explicit supervision. We show that the pseudo-labels produced by this procedure can be used to train a captioning model to effectively understand human-human interactions in images, as measured by a variety of metrics that measure textual and semantic faithfulness and factual groundedness of our predictions. We further show that our approach outperforms SOTA image captioning and situation recognition models on this task. We will releas 1 our code and pseudo-labels along with Waldo and Wenda, a manually-curated test set for still image human-human interaction understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- LiT: Zero-Shot Transfer with Locked-image text TuningXiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner 等CVPR 2022 · 被引用 349 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
相关 Paper
- Who's Waldo? Linking People Across Text and ImagesClaire Yuqing Cui, Apoorv Khandelwal, Yoav Artzi, Noah Snavely 等ICCV 2021 · 被引用 21 次
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 被引用 5 次
- AcT2I: Evaluating and Improving Action Depiction in Text-to-Image ModelsVatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang 等EMNLP 2025
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 被引用 115 次
- Zero-Shot Referring Expression Comprehension via Structural Similarity Between Images and CaptionsZeyu Han, Fangrui Zhu, Qianru Lao, Huaizu JiangCVPR 2024
