Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning
Huaizheng Zhang, Yong Luo, Qiming Ai, Yonggang Wen, Han Hu
Abstract
Given the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. However, manually finding relevant ads to match the provided content is labor-intensive, and hence some automatic advertising techniques are developed. Since ads are usually hard to understand only according to its visual appearance due to the contained visual metaphor, some other modalities, such as the contained texts, should be exploited for understanding. To further improve user experience, it is necessary to understand both the ads' topic and sentiment. This motivates us to develop a novel deep multimodal multitask framework that integrates multiple modalities to achieve effective topic and sentiment prediction simultaneously for ads understanding. In particular, in our framework termed DeepAd, we first extract multimodal information from ads and learn high-level and comparable representations. The visual metaphor of the ad is decoded in an unsupervised manner. The obtained representations are then fed into the proposed hierarchical multimodal attention modules to learn task-specific representations for final prediction. A multitask loss function is also designed to jointly train both the topic and sentiment prediction models in an end-to-end manner, where bottom-layer parameters are shared to alleviate over-fitting. We conduct extensive experiments on a large-scale advertisement dataset and achieve state-of-the-art performance for both prediction tasks. The obtained results could be utilized as a benchmark for ads understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4191da9-0068-40ac-a224-c5351ad74b8aCited by top-tier papers2
- MM-AU: Towards Multimodal Understanding of Advertisement VideosDigbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli et al.ACM MM 2023 · 8 citations
- Affect2MM: Affective Analysis of Multimedia Content Using Emotion CausalityTrisha Mittal, Puneet Mathur, Aniket Bera, Dinesh ManochaCVPR 2021
Related papers
- Sentiment and Emotion help Sarcasm? A Multi-task Learning Framework for Multi-Modal Sarcasm, Sentiment and Emotion AnalysisDushyant Singh Chauhan, Dhanush S. R, Asif Ekbal, Pushpak BhattacharyyaACL 2020 · 131 citations
- A Video Is Worth 4096 Tokens: Verbalize Story Videos To Understand Them In Zero ShotAanisha Bhattacharyya, Yaman Singla, Balaji Krishnamurthy, Rajiv Ratn Shah et al.EMNLP 2023 · 9 citations
- Generating Multimodal Metaphorical Features for Meme UnderstandingBo Xu, Junzhe Zheng, Jiayuan He, Yuxuan Sun et al.ACM MM 2024 · 6 citations
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 116 citations
