Point-of-Interest Type Prediction using Text and Images
Danae Sánchez Villegas, Nikolaos Aletras
Abstract
Point-of-interest (POI) type prediction is the task of inferring the type of a place from where a social media post was shared. Inferring a POI's type is useful for studies in computational social science including sociolinguistics, geosemiotics, and cultural geography, and has applications in geosocial networking technologies such as recommendation and visualization systems. Prior efforts in POI type prediction focus solely on text, without taking visual information into account. However in reality, the variety of modalities, as well as their semiotic relationships with one another, shape communication and interactions in social media. This paper presents a study on POI type prediction using multimodal information from text and images available at posting time. For that purpose, we enrich a currently available data set for POI type prediction with the images that accompany the text messages. Our proposed method extracts relevant information from each modality to effectively capture interactions between text and image achieving a macro F1 of 47.21 across eight categories significantly outperforming the state-of-the-art method for POI type prediction based on textonly methods. Finally, we provide a detailed analysis to shed light on cross-modal interactions and the limitations of our best performing model. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a436d116-d70f-4fdb-ba2d-babbe4bdbd66Cited by top-tier papers2
- Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal ClassificationChunpu Xu, Jing LiEMNLP 2022 · 3 citations
- Automatic Identification and Classification of Bragging in Social MediaMali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos AletrasACL 2022
Builds on4
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image WordingsYue Wang, Jing Li, Michael R. Lyu, Irwin KingEMNLP 2020 · 11 citations
- Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!Jack Hessel, Lillian LeeEMNLP 2020 · 3 citations
- Analyzing Political Parody in Social MediaAntonis Maronikolakis, Danae Sanchez Villegas, Daniel Preotiuc-Pietro, Nikolaos AletrasACL 2020 · 3 citations
Related papers
- "I Have No Text in My Post": Using Visual Hints to Model User Emotions in Social MediaJunho Song, Kyungsik Han, Sang-Wook KimWWW 2022 · 5 citations
- MMPOI: A Multi-Modal Content-Aware Framework for POI RecommendationsYang Xu, Gao Cong, Lei Zhu, Lizhen CuiWWW 2024 · 24 citations
- Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood EmbeddingZhecheng Wang, Haoyuan Li, Ram RajagopalAAAI 2020 · 113 citations
- Multimodal Representation with Embedded Visual Guiding Objects for Named Entity Recognition in Social Media PostsZhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen et al.ACM MM 2020 · 139 citations
- POI-Enhancer: An LLM-based Semantic Enhancement Framework for POI Representation LearningJiawei Cheng, Jingyuan Wang, Yichuan Zhang, Jiahao Ji et al.AAAI 2025 · 30 citations
