Adaptive Transformer-Based Conditioned Variational Autoencoder for Incomplete Social Event Classification
Zhangming Li, Shengsheng Qian, Jie Cao, Quan Fang, Changsheng Xu
Abstract
With the rapid development of the Internet and the expanding scale of social media, incomplete social event classification has increasingly become a challenging task. The key for incomplete social event classification is to accurately leverage the image-level and text-level information. However, most of the existing approaches may suffer from the following limitations: (1) Most Generative Models use the available features to generate the incomplete modality features for social events classification while ignoring the rich semantic label information. (2) The majority of existing multi-modal methods just simply concatenate the coarse-grained image features and text features of the event to get the multi-modal features to classify social events, which ignores the irrelevant multi-modal features and limits their modeling capabilities. To tackle these challenges, in this paper, we propose an Adaptive Transformer-Based Conditioned Variational Autoencoder Network (AT-CVAE) for incomplete social event classification. In the AT-CVAE, we propose a novel Transformer-based Conditioned Variational Autoencoder to jointly model the textual information, visual information and label information into a unified deep model, which can generate more discriminative latent features and enhance the performance of incomplete social event classification. Furthermore, the Mixture-of-Experts Mechanism is utilized to dynamically acquire the weights of each multi-modal information, which can better filter out the irrelevant multi-modal information and capture the vitally important information. Extensive experiments are conducted on two public event datasets, demonstrating the superior performance of our AT-CVAE method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e6e9cb1a-8b68-4fc2-bdf3-3b6b9c952deaRelated papers
- Open-World Social Event ClassificationShengsheng Qian, Hong Chen, Dizhan Xue, Quan Fang et al.WWW 2023 · 24 citations
- A Span-based Multimodal Variational Autoencoder for Semi-supervised Multimodal Named Entity RecognitionBaohang Zhou, Ying Zhang, Kehui Song, Wenya Guo et al.EMNLP 2022 · 15 citations
- Graph Convolutional Incomplete Multi-modal HashingXiaobo Shen, Yinfan Chen, Shirui Pan, Weiwei Liu et al.ACM MM 2023 · 16 citations
- Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information ExtractionBaohang Zhou, Ying Zhang, Yu Zhao, Xuhui Sui et al.WWW 2025 · 5 citations
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu et al.ACM MM 2020 · 63 citations
