DiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science Articles
Tanishq Gupta, Mohd Zaki, Devanshi Khatsuriya, Kausik Hira, N. M. Anoop Krishnan, Mausam
Abstract
A crucial component in the curation of KB for a scientific domain (e.g., materials science, foods & nutrition, fuels) is information extraction from tables in the domain's published research articles. To facilitate research in this direction, we define a novel NLP task of extracting compositions of materials (e.g., glasses) from tables in materials science papers. The task involves solving several challenges in concert, such as tables that mention compositions have highly varying structures; text in captions and full paper needs to be incorporated along with data in tables; and regular languages for numbers, chemical compounds and composition expressions must be integrated into the model. We release a training dataset comprising 4,408 distantly supervised tables, along with 1,475 manually annotated dev and test tables. We also present DISCOMAT, a strong baseline that combines multiple graph neural networks with several task-specific regular expressions, features, and constraints. We show that DIS-COMAT outperforms recent table processing architectures by significant margins. We release our code and data for further research on this challenging IE task from scientific tables.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5598b701-5877-4f1b-a19d-c88325903cd9Cited by top-tier papers2
- SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific LiteratureDavid Wadden, Kejian Shi, Jacob Morrison, Alan Li et al.EMNLP 2025 · 2 citations
- ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language ModelsBenjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue et al.EMNLP 2024 · 1 citation
Builds on5
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno et al.ACL 2020 · 19 citations
- Topic Transferable Table Question AnsweringSaneem A. Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen et al.EMNLP 2021
- INFOTABS: Inference on Tables as Semi-structured DataVivek Gupta, Maitrey Mehta, Pegah Nokhiz, Vivek SrikumarACL 2020
Related papers
- A Multi-Task Learning Framework for Reading Comprehension of Scientific Tabular DataXu Yang, Meihui Zhang, Ju Fan, Zeyu Luo et al.ICDE 2024 · 1 citation
- MS-Mentions: Consistently Annotating Entity Mentions in Materials Science Procedural TextTim O'Gorman, Zach Jensen, Sheshera Mysore, Kevin Huang et al.EMNLP 2021 · 13 citations
- The SOFC-Exp Corpus and Neural Approaches to Information Extraction in the Materials Science DomainAnnemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl et al.ACL 2020 · 17 citations
- MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema ModelingYu Song, Santiago Miret, Bang LiuACL 2023 · 24 citations
- PubTables-1M: Towards comprehensive table extraction from unstructured documentsBrandon Smock, Rohith Pesala, Robin AbrahamCVPR 2022 · 125 citations
