MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization
Canwen Xu, Jiaxin Pei, Hongtao Wu, Yiyu Liu, Chenliang Li
摘要
Recently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing. However, there is currently no cross-task dataset in NLP, which hinders the development of multi-task learning. We propose MATINF, the first jointly labeled large-scale dataset for classification, question answering and summarization. MAT-INF contains 1.07 million question-answer pairs with human-labeled categories and usergenerated question descriptions. Based on such rich information, MATINF is applicable for three major NLP tasks, including classification, question answering, and summarization. We benchmark existing methods and a novel multi-task baseline over MATINF to inspire further research. Our comprehensive comparison and experiments over MATINF and other datasets demonstrate the merits held by MAT-INF. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-trainingWangchunshu Zhou, Yan Zeng, Shizhe Diao, Xinsong ZhangICML 2022 · 被引用 17 次
- Coupling Context Modeling with Zero Pronoun Recovering for Document-Level Natural Language GenerationXin Tan, Longyin Zhang, Guodong ZhouEMNLP 2021 · 被引用 6 次
- NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in NorwegianPeng Liu, Lemei Zhang, Terje Nissen Farup, Even W. Lauvrak 等EMNLP 2024 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du 等ACL 2025
- Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation DataXun Zhu, Fanbin Mo, Zheng Zhang, Jiaxi Wang 等ACM MM 2025
- MLSUM: The Multilingual Summarization CorpusThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski 等EMNLP 2020 · 被引用 4 次
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni 等EMNLP 2020 · 被引用 104 次
- Summarizing Community-based Question-Answer PairsTing-Yao Hsu, Yoshi Suhara, Xiaolan WangEMNLP 2022 · 被引用 5 次
