End-Users Know Best: Identifying Undesired Behavior of Alexa Skills Through User Review Analysis
Mohammed Aldeen, Jeffrey Young, Song Liao, Tsu-Yao Chang, Long Cheng, Haipeng Cai, Xiapu Luo, Hongxin Hu
Abstract
The Amazon Alexa marketplace has grown rapidly in recent years due to third-party developers creating large amounts of content and publishing directly to a skills store. Despite the growth of the Amazon Alexa skills store, there have been several reported security and usability concerns, which may not be identified during the vetting phase. However, user reviews can offer valuable insights into the security & privacy, quality, and usability of the skills. To better understand the effects of these problematic skills on end-users, we introduce ReviewTracker, a tool capable of discerning and classifying semantically negative user reviews to identify likely malicious, policy violating, or malfunctioning behavior on Alexa skills. ReviewTracker employs a pre-trained FastText classifier to identify different undesired skill behaviors. We collected over 700,000 user reviews spanning 6 years with more than 200,000 negative sentiment reviews. ReviewTracker was able to identify 17,820 reviews reporting violations related to Alexa policy requirements across 2,813 skills, and 131,855 reviews highlighting different types of user frustrations associated with 9,294 skills. In addition, we developed a dynamic skill testing framework using ChatGPT to conduct two distinct types of tests on Alexa skills: one using a software-based simulation for interaction to explore the actual behaviors of skills and another through actual voice commands to understand the potential factors causing discrepancies between intended skill functionalities and user experiences. Based on the number of the undesired skill behavior reviews, we tested the top identified problematic skills and detected more than 228 skills violating at least one policy requirement. Our results demonstrate that user reviews could serve as a valuable means to identify undesired skill behaviors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88c17731-bb85-4e71-baba-d854571c3edeCited by top-tier papers4
- The Homework Wars: Exploring Emotions, Behaviours, and Conflicts in Parent-Child Homework InteractionsNan Gao, Yibin Liu, Xin Tang, Yanyan Liu et al.UbiComp 2025 · 2 citations
- AdsDP: A Video Dataset for Recognizing and Examining Dark Patterns in iOS In-App AdvertisementsYuxuan Shang, Guanxiao Wang, Mengxia Ren, Haomin Zhang et al.UbiComp 2025 · 1 citation
- SKILLPoV: Towards Accessible and Effective Privacy Notice for Amazon Alexa SkillsJingwen Yan, Song Liao, Mohammed Aldeen, Luyi Xing et al.NDSS 2025
- No Way to Sign Out? Unpacking Non-Compliance with Google Play's App Account Deletion RequirementsJingwen Yan, Song Liao, Jin Ma, Mohammed Aldeen et al.USENIX Security 2025
Builds on13
- Skill Squatting Attacks on Amazon AlexaDeepak Kumar, Riccardo Paccagnella, Paul Murley, Eric Hennenfent et al.USENIX Security 2018 · 177 citations
- Short Text, Large Effect: Measuring the Impact of User Reviews on Android App Security & PrivacyDuc Cuong Nguyen, Erik Derr, Michael Backes, Sven BugielS&P 2019 · 66 citations
- Dangerous Skills Got Certified: Measuring the Trustworthiness of Skill Certification in Voice Personal Assistant PlatformsLong Cheng, Christin Wilson, Song Liao, Jeffrey Young et al.CCS 2020 · 58 citations
- Hark: A Deep Learning System for Navigating Privacy Feedback at ScaleHamza Harkous, Sai Teja Peddinti, Rishabh Khandelwal, Animesh Srivastava et al.S&P 2022 · 37 citations
- CHAMP: Characterizing Undesired App Behaviors from User Comments based on Market PoliciesYangyu Hu, Haoyu Wang, Tiantong Ji, Xusheng Xiao et al.ICSE 2021 · 19 citations
Related papers
- SkillScanner: Detecting Policy-Violating Voice Applications Through Static Analysis at the Development PhaseSong Liao, Long Cheng, Haipeng Cai, Linke Guo et al.CCS 2023 · 7 citations
- Hey Alexa, is this Skill Safe?: Taking a Closer Look at the Alexa Skill EcosystemChristopher Lentzsch, Sheel Jayesh Shah, Benjamin Andow, Martin Degeling et al.NDSS 2021
- SkillDetective: Automated Policy-Violation Detection of Voice Assistant Applications in the WildJeffrey Young, Song Liao, Long Cheng, Hongxin Hu et al.USENIX Security 2022
- SkillExplorer: Understanding the Behavior of Skills in Large ScaleZhixiu Guo, Zijin Lin, Pan Li, Kai ChenUSENIX Security 2020
- Analyzing Ad Prevalence, Characteristics, and Compliance in Alexa SkillsAafaq Sabir, Abhinaya S. B., Dilawer Ahmed, Anupam DasS&P 2025
