Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
Elisa Tsai, Neal Mangaokar, Boyuan Zheng, Haizhong Zheng, Atul Prakash
Abstract
Terms and conditions for online shopping websites often contain terms that can have significant financial consequences for customers. Despite their impact, there is currently no comprehensive understanding of the types and potential risks associated with unfavorable financial terms. Furthermore, there are no publicly available detection systems or datasets to systematically identify or mitigate these terms. In this paper, we take the first steps toward solving this problem with three key contributions. First, we introduce TermMiner, an automated data collection and topic modeling pipeline to understand the landscape of unfavorable financial terms. Second, we create ShopTC-100K, a dataset of terms and conditions from shopping websites in the Tranco top 100K list, comprising 1.8 million terms from 8,251 websites. Consequently, we develop a taxonomy of 22 types from 4 categories of unfavorable financial terms-spanning purchase, post-purchase, account termination, and legal aspects. Third, we build TermLens, an automated detector that uses Large Language Models (LLMs) to identify unfavorable financial terms. Fine-tuned on an annotated dataset, TermLens achieves an F1 score of 94.6% and a false positive rate of 2.3% using GPT-4o. When applied to shopping websites from the Tranco top 100K, we find that 42.06% of these sites contain at least one unfavorable financial term, with such terms being more prevalent on less popular websites. Case studies further highlight the financial risks and customer dissatisfaction associated with unfavorable financial terms, as well as the limitations of existing ecosystem defenses. CCS Concepts • Information systems → Web mining; • Security and privacy → Social engineering attacks; • Social and professional topics → Consumer products policy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7881e386-8f58-452c-8fb2-96cbbda9ab06Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski et al.NDSS 2019 · 826 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Dark Patterns after the GDPR: Scraping Consent Pop-ups and Demonstrating their InfluenceMidas Nouwens, Ilaria Liccardi, Michael Veale, David R. Karger et al.CHI 2020 · 491 citations
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub et al.USENIX Security 2018 · 400 citations
Related papers
- TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial DomainsYidan Sun, Mengying Zhu, Feiyue Chen, Yangyang Wu et al.AAAI 2026 · 1 citation
- Beyond Phish: Toward Detecting Fraudulent e-Commerce Websites at ScaleMarzieh Bitaab, Haehyun Cho, Adam Oest, Zhuoer Lyu et al.S&P 2023
- AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer ContractsDaniel Braun, Florian MatthesACL 2024 · 2 citations
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot PluginsYigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann et al.S&P 2026 · 12 citations
- Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection CapabilitiesMasahiro Kaneko, Timothy BaldwinEMNLP 2025 · 2 citations
