Low Resource Quantitative Information Extraction via Structure Searching and Prefix-Based Text Generation
Tongliang Li, Zixiang Wang, Zhoujun Li
Abstract
Quantitative information plays an important part in the financial and data analysis areas. Prior work relied on pattern-matching methods and complex hand-crafted rules to extract quantitative information due to the lack of labeled data. Such methods can be unstable and difficult to scale to the open domain. In this paper, we study quantitative information extraction in the low-resource setting. We propose a search-based approach by searching from the syntactic structures to acquire basic training data. The search process is simple yet effective. Then, a prefix-based text-to-text generation method is employed to extract the quantitative information. The prefix design can fully leverage pre-trained language models for text generation to serve the information extraction purpose. Experimental results show that our approaches achieves high performance with a limited amount of labeled data. The extraction result could further boost the performance of other tasks such as quantitative reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
- Prefix-Tuning: Optimizing Continuous Prompts for GenerationXiang Lisa Li, Percy LiangACL 2021
- Unified Structure Generation for Universal Information ExtractionYaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao et al.ACL 2022
Related papers
- STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language ModelsMingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham et al.AAAI 2024 · 22 citations
- Pre-training Language Models for Comparative ReasoningMengxia Yu, Zhihan Zhang, Wenhao Yu, Meng JiangEMNLP 2023 · 1 citation
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 4 citations
- CQE: A Comprehensive Quantity ExtractorSatya Almasian, Vivian Kazakova, Philip Göldner, Michael GertzEMNLP 2023 · 3 citations
- Unsupervised Graph-Text Mutual Conversion with a Unified Pretrained Language ModelYi Xu, Shuqian Sheng, Jiexing Qi, Luoyi Fu et al.ACL 2023
