Intentonomy: A Dataset and Study Towards Human Intent Understanding
Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge J. Belongie, Ser-Nam Lim
Abstract
An image is worth a thousand words, conveying information that goes beyond the mere visual content therein. In this paper, we study the intent behind social media images with an aim to analyze how visual information can facilitate recognition of human intent. Towards this goal, we introduce an intent dataset, Intentonomy, comprising 14K images covering a wide range of everyday scenes. These images are manually annotated with 28 intent categories derived from a social psychology taxonomy. We then systematically study whether, and to what extent, commonly used visual information, i.e., object and context, contribute to human motive understanding. Based on our findings, we conduct further study to quantify the effect of attending to object and context classes as well as textual information in the form of hashtags when training an intent classifier. Our results quantitatively and qualitatively shed light on how visual and textual information can produce observable effects when predicting intent. 1 Annotation details Amazon Mechanical Turk (MTurk) was recruited to collect labels of perceived intent by employing a similarity comparison task that we call "unsatisfactory substitutes". We rely on the notion of "mental imagery" [46] -a quasi-perceptual experience that maps example images to a visual representation in one's mind, along with games
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af962f87-16b8-4866-96c1-6c38ca24d01bCited by top-tier papers10
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 97 citations
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu et al.NeurIPS 2024 · 31 citations
- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in ConversationsHanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou et al.ICLR 2024 · 29 citations
- Exploring Visual Engagement Signals for Representation LearningMenglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie et al.ICCV 2021 · 15 citations
Builds on2
Related papers
- Oops! Predicting Unintentional Action in VideoDave Epstein, Boyuan Chen, Carl VondrickCVPR 2020
- Impressions: Visual Semiotics and Aesthetic Impact UnderstandingJulia Kruk, Caleb Ziems, Diyi YangEMNLP 2023 · 5 citations
- Visual Intention Grounding for Egocentric AssistantsPengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li et al.ICCV 2025 · 2 citations
- Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense DiscoveryFeihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu et al.ACM MM 2024 · 11 citations
- Affection: Learning Affective Explanations for Real-World Visual DataPanos Achlioptas, Maks Ovsjanikov, Leonidas J. Guibas, Sergey TulyakovCVPR 2023
