MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization
Canwen Xu, Jiaxin Pei, Hongtao Wu, Yiyu Liu, Chenliang Li
Abstract
Recently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing. However, there is currently no cross-task dataset in NLP, which hinders the development of multi-task learning. We propose MATINF, the first jointly labeled large-scale dataset for classification, question answering and summarization. MAT-INF contains 1.07 million question-answer pairs with human-labeled categories and usergenerated question descriptions. Based on such rich information, MATINF is applicable for three major NLP tasks, including classification, question answering, and summarization. We benchmark existing methods and a novel multi-task baseline over MATINF to inspire further research. Our comprehensive comparison and experiments over MATINF and other datasets demonstrate the merits held by MAT-INF. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-trainingWangchunshu Zhou, Yan Zeng, Shizhe Diao, Xinsong ZhangICML 2022 · 17 citations
- Coupling Context Modeling with Zero Pronoun Recovering for Document-Level Natural Language GenerationXin Tan, Longyin Zhang, Guodong ZhouEMNLP 2021 · 6 citations
- NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in NorwegianPeng Liu, Lemei Zhang, Terje Nissen Farup, Even W. Lauvrak et al.EMNLP 2024 · 1 citation
Builds on1
Related papers
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du et al.ACL 2025
- Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation DataXun Zhu, Fanbin Mo, Zheng Zhang, Jiaxi Wang et al.ACM MM 2025
- MLSUM: The Multilingual Summarization CorpusThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.EMNLP 2020 · 4 citations
- Exploring and Predicting Transferability across NLP TasksTu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni et al.EMNLP 2020 · 104 citations
- Summarizing Community-based Question-Answer PairsTing-Yao Hsu, Yoshi Suhara, Xiaolan WangEMNLP 2022 · 5 citations
