Seeing Beyond Words: Multimodal Aspect-Level Complaint Detection in Ecommerce Videos
Rishikesh Devanathan, Apoorva Singh, A. S. Poornash, Sriparna Saha
Abstract
Complaints are pivotal expressions within e-commerce communication, yet the intricate nuances of human interaction present formidable challenges for AI agents to grasp comprehensively. While recent attention has been drawn to analyzing complaints within a multimodal context, relying solely on text and images is insufficient for organizations. The true value lies in the ability to pinpoint complaints within the intricate structures of discourse, scrutinizing them at a granular aspect level. Our research delves into the discourse structure of e-commerce video-based product reviews, pioneering a novel task we term Aspect-Level Complaint Detection from Discourse (ACDD). Embedded in a multimodal framework, this task entails identifying aspect categories and assigning complaint/non-complaint labels at a nuanced aspect level. To facilitate this endeavour, we have curated a unique multimodal product review dataset, meticulously annotated at the utterance level with aspect categories and associated complaint labels. To support this undertaking, we introduce a Multimodal Aspect-Aware Complaint Analysis (MAACA) model that incorporates a novel pre-training strategy and a global feature fusion technique across the three modalities. Additionally, the proposed framework leverages a moment retrieval step to identify the relevant portion of the clip, crucial for accurately detecting the fine-grained aspect categories and conducting aspect-level complaint detection. Extensive experiments conducted on the proposed dataset showcase that our framework outperforms unimodal and bimodal baselines, offering valuable insights into the application of video-audio-text representation learning frameworks for downstream tasks. The dataset and code are available at: https://github.com/rdev12/MAACA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Federated Meta-Learning for Emotion and Sentiment Aware Multi-modal Complaint IdentificationApoorva Singh, Siddarth Chandrasekar, Sriparna Saha, Tanmay SenEMNLP 2023 · 5 citations
- Aspect-Based Multimodal Mining: Unveiling Sentiments, Complaints, and Beyond in User-Generated ContentMamta, Gopendra Vikram Singh, Deepak Raju Kori, Asif EkbalACM MM 2024 · 1 citation
- AbCoRD: Exploiting multimodal generative approach for Aspect-based Complaint and Rationale DetectionRaghav Jain, Apoorva Singh, Vivek Kumar Gangwar, Sriparna SahaACM MM 2023 · 2 citations
- JDDC 2.1: A Multimodal Chinese Dialogue Dataset with Joint Tasks of Query Rewriting, Response Generation, Discourse Parsing, and SummarizationNan Zhao, Haoran Li, Youzheng Wu, Xiaodong HeEMNLP 2022 · 6 citations
- M5Product: Self-harmonized Contrastive Learning for E-commercial Multi-modal PretrainingXiao Dong, Xunlin Zhan, Yangxin Wu, Yunchao Wei et al.CVPR 2022 · 24 citations
