Explaining Relationships Between Scientific Documents
Kelvin Luu, Xinyi Wu, Rik Koncel-Kedziorski, Kyle Lo, Isabel Cachola, Noah A. Smith
Abstract
We address the task of explaining relationships between two scientific documents using natural language text. This task requires modeling the complex content of long technical documents, deducing a relationship between these documents, and expressing that relationship in text. Successful solutions can help improve researcher efficiency in search and review. In this paper, we operationalize this task by using citing sentences as a proxy. We establish a large dataset for our task. We pretrain a large language model to serve as the foundation for autoregressive approaches to the task. We explore the impact of taking different views on the two documents, including the use of dense representations extracted with scientific information extraction systems. We provide extensive automatic and human evaluations which show the promise of such models, and make clear the challenges for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71f8d471-c09e-439a-96f6-f361b21fb4beCited by top-tier papers16
- Neighborhood Contrastive Learning for Scientific Document Representations with Citation EmbeddingsMalte Ostendorff, Nils Rethmeier, Isabelle Augenstein, Bela Gipp et al.EMNLP 2022 · 46 citations
- Explainable Legal Case Matching via Inverse Optimal Transport-based Rationale ExtractionWeijie Yu, Zhongxiang Sun, Jun Xu, Zhenhua Dong et al.SIGIR 2022 · 45 citations
- DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation GenerationYifan Wang, Yiping Song, Shuai Li, Chaoran Cheng et al.AAAI 2022 · 42 citations
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang et al.EMNLP 2024 · 28 citations
- PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected PapersYoonjoo Lee, Hyeonsu B. Kang, Matt Latzke, Juho Kim et al.CHI 2024 · 20 citations
Builds on6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
- On Extractive and Abstractive Neural Document Summarization with Transformer Language ModelsJonathan Pilault, Raymond Li, Sandeep Subramanian, Chris PalEMNLP 2020 · 186 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
Related papers
- Automatic Generation of Citation Texts in Scholarly Papers: A Pilot StudyXinyu Xing, Xiaosheng Fan, Xiaojun WanACL 2020 · 39 citations
- SciREX: A Challenge Dataset for Document-Level Information ExtractionSarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz BeltagyACL 2020 · 9 citations
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li et al.EMNLP 2021 · 18 citations
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey et al.ACL 2020 · 20 citations
- HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation PredictionQianyue Hao, Jingyang Fan, Fengli Xu, Jian Yuan et al.NeurIPS 2024 · 23 citations
