Snippext: Semi-supervised Opinion Mining with Augmented Data
Zhengjie Miao, Yuliang Li, Xiaolan Wang, Wang-Chiew Tan
摘要
Online services are interested in solutions to opinion mining, which is the problem of extracting aspects, opinions, and sentiments from text. One method to mine opinions is to leverage the recent success of pre-trained language models which can be fine-tuned to obtain highquality extractions from reviews. However, fine-tuning language models still requires a non-trivial amount of training data. In this paper, we study the problem of how to significantly reduce the amount of labeled training data required in fine-tuning language models for opinion mining. We describe Snippext, an opinion mining system developed over a language model that is fine-tuned through semi-supervised learning with augmented data. A novelty of Snippext is its clever use of a two-prong approach to achieve state-of-the-art (SOTA) performance with little labeled training data through: (1) data augmentation to automatically generate more labeled training data from existing ones, and (2) a semi-supervised learning technique to leverage the massive amount of unlabeled data in addition to the (limited amount of) labeled data. We show with extensive experiments that Snippext performs comparably and can even exceed previous SOTA results on several opinion mining tasks with only half the training data required. Furthermore, it achieves new SOTA results when all training data are leveraged. By comparison to a baseline pipeline, we found that Snippext extracts significantly more fine-grained opinions which enable new opportunities of downstream applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian 等ACL 2022 · 被引用 91 次
- Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondZhengjie Miao, Yuliang Li, Xiaolan WangSIGMOD 2021 · 被引用 63 次
- Local Additivity Based Data Augmentation for Semi-supervised NERJiaao Chen, Zhenghui Wang, Ran Tian, Zichao Yang 等EMNLP 2020 · 被引用 45 次
- Accurate and Explainable Recommendation via Review RationalizationSicheng Pan, Dongsheng Li, Hansu Gu, Tun Lu 等WWW 2022 · 被引用 22 次
相关 Paper
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- STraTA: Self-Training with Task Augmentation for Better Few-shot LearningTu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon 等EMNLP 2021 · 被引用 25 次
- Conditional Augmentation for Aspect Term Extraction via Masked Sequence-to-Sequence GenerationKun Li, Chengbo Chen, Xiaojun Quan, Qing Ling 等ACL 2020 · 被引用 101 次
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsFei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai 等EMNLP 2021 · 被引用 18 次
- Revisiting Self-training for Few-shot Learning of Language ModelYiming Chen, Yan Zhang, Chen Zhang, Grandee Lee 等EMNLP 2021 · 被引用 35 次
