Feeding What You Need by Understanding What You Learned
Xiaoqiang Wang, Bang Liu, Fangli Xu, Bo Long, Siliang Tang, Lingfei Wu
摘要
Machine Reading Comprehension (MRC) reveals the ability to understand a given text passage and answer questions based on it. Existing research works in MRC rely heavily on large-size models and corpus to improve the performance evaluated by metrics such as Exact Match (EM ) and F 1 . However, such a paradigm lacks sufficient interpretation to model capability and can not efficiently train a model with a large corpus. In this paper, we argue that a deep understanding of model capabilities and data properties can help us feed a model with appropriate training data based on its learning status. Specifically, we design an MRC capability assessment framework that assesses model capabilities in an explainable and multi-dimensional manner. Based on it, we further uncover and disentangle the connections between various data properties and model performance. Finally, to verify the effectiveness of the proposed MRC capability assessment framework, we incorporate it into a curriculum learning pipeline and devise a Capability Boundary Breakthrough Curriculum (CBBC) strategy, which performs a model capability-based training to maximize the data value and improve training efficiency. Extensive experiments demonstrate that our approach significantly improves performance, achieving up to an 11.22% / 8.71% improvement of EM / F 1 on MRC tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Adjective Scale Probe: Can Language Models Encode Formal Semantics Information?Wei Liu, Ming Xiang, Nai DingAAAI 2023 · 被引用 7 次
- SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive ViewYongjie Xiao, Hongru Liang, Peixin Qin, Yao Zhang 等ACL 2025 · 被引用 1 次
- FAC²E: Better Understanding Large Language Model Capabilities by Dissociating Language and CognitionXiaoqiang Wang, Lingfei Wu, Tengfei Ma, Bang LiuEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
- Uncertainty-guided Continual Learning with Bayesian Neural NetworksSayna Ebrahimi, Mohamed Elhoseiny, Trevor Darrell, Marcus RohrbachICLR 2020 · 被引用 211 次
- Curriculum Learning for Natural Language UnderstandingBenfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang 等ACL 2020 · 被引用 156 次
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 被引用 92 次
相关 Paper
- Multi-Task Learning with Generative Adversarial Training for Multi-Passage Machine Reading ComprehensionQiyu Ren, Xiang Cheng, Sen SuAAAI 2020 · 被引用 15 次
- Span Selection Pre-training for Question AnsweringMichael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto 等ACL 2020 · 被引用 9 次
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen 等ACL 2020 · 被引用 39 次
- MMM: Multi-Stage Multi-Task Learning for Multi-Choice Reading ComprehensionDi Jin, Shuyang Gao, Jiun-Yu Kao, Tagyoung Chung 等AAAI 2020 · 被引用 72 次
- Enhancing Multiple-choice Machine Reading Comprehension by Punishing Illogical InterpretationsYiming Ju, Yuanzhe Zhang, Zhixing Tian, Kang Liu 等EMNLP 2021 · 被引用 8 次
