MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question Answering
Xiusi Chen, Jyun-Yu Jiang, Wei-Cheng Chang, Cho-Jui Hsieh, Hsiang-Fu Yu, Wei Wang
摘要
Recent advances in few-shot question answering (QA) mostly rely on the power of pretrained large language models (LLMs) and fine-tuning in specific settings. Although the pre-training stage has already equipped LLMs with powerful reasoning capabilities, LLMs still need to be fine-tuned to adapt to specific domains to achieve the best results. In this paper, we propose to select the most informative data for fine-tuning, thereby improving the efficiency of the fine-tuning process with comparative or even better accuracy on the open-domain QA task. We present MINPROMPT, a minimal data augmentation framework for opendomain QA based on an approximate graph algorithm and unsupervised question generation. We transform the raw text into a graph structure to build connections between different factual sentences, then apply graph algorithms to identify the minimal set of sentences needed to cover the most information in the raw text. We then generate QA pairs based on the identified sentence subset and train the model on the selected sentences to obtain the final model. Empirical results on several benchmark datasets and theoretical analysis show that MINPROMPT is able to achieve comparable or better results than baselines with a high degree of efficiency, bringing consistent improvements in F-1 scores.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Accurate and Efficient Multivariate Time Series Forecasting via Offline ClusteringYiming Niu, Jinliang Deng, Lulu Zhang, Zimu Zhou 等ICDE 2025 · 被引用 4 次
- M+: Extending MemoryLLM with Scalable Long-Term MemoryYu Wang, Dmitry Krotov, Yuanzhe Hu, Yifan Gao 等ICML 2025
它引用的顶会 Paper9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 被引用 167 次
- Graph Neural Prompting with Large Language ModelsYijun Tian, Huan Song, Zichen Wang, Haozhu Wang 等AAAI 2024 · 被引用 90 次
- Improving Question Generation with Sentence-Level Semantic Matching and Answer Position InferringXiyao Ma, Qile Zhu, Yanlin Zhou, Xiaolin LiAAAI 2020 · 被引用 68 次
- FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text modelsRakesh Chada, Pradeep NatarajanEMNLP 2021 · 被引用 36 次
相关 Paper
- Knowledge Graph Prompting for Multi-Document Question AnsweringYu Wang, Nedim Lipka, Ryan A. Rossi, Alexa F. Siu 等AAAI 2024 · 被引用 290 次
- From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query GenerationZhenyu Tong, Chuan Qin, Chuyu Fang, Kaichun Yao 等KDD 2025 · 被引用 4 次
- Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question GenerationYuanyuan Liang, Jianing Wang, Hanlun Zhu, Lei Wang 等EMNLP 2023 · 被引用 24 次
- Leveraging QA Datasets to Improve Generative Data AugmentationDheeraj Mekala, Tu Vu, Timo Schick, Jingbo ShangEMNLP 2022 · 被引用 8 次
- Knowledge Graph Finetuning Enhances Knowledge Manipulation in Large Language ModelsHanzhu Chen, Xu Shen, Jie Wang, Zehao Wang 等ICLR 2025
