Structure and Semantics Preserving Document Representations
Natraj Raman, Sameena Shah, Manuela Veloso
摘要
Retrieving relevant documents from a corpus is typically based on the semantic similarity between the document content and query text. The inclusion of structural relationship between documents can benefit the retrieval mechanism by addressing semantic gaps. However, incorporating these relationships requires tractable mechanisms that balance structure with semantics and take advantage of the prevalent pre-train/fine-tune paradigm. We propose here a holistic approach to learning document representations by integrating intra-document content with inter-document relations. Our deep metric learning solution analyzes the complex neighborhood structure in the relationship network to efficiently sample similar/dissimilar document pairs and defines a novel quintuplet loss function that simultaneously encourages document pairs that are semantically relevant to be closer and structurally unrelated to be far apart in the representation space. Furthermore, the separation margins between the documents are varied flexibly to encode the heterogeneity in relationship strengths. The model is fully fine-tunable and natively supports query projection during inference. We demonstrate that it outperforms competing methods on multiple datasets for document retrieval tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 被引用 76 次
- Ladder Loss for Coherent Visual-Semantic EmbeddingMo Zhou, Zhenxing Niu, Le Wang, Zhanning Gao 等AAAI 2020 · 被引用 46 次
- Which *BERT? A Survey Organizing Contextualized EncodersPatrick Xia, Shijie Wu, Benjamin Van DurmeEMNLP 2020 · 被引用 44 次
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey 等ACL 2020 · 被引用 20 次
相关 Paper
- Contextual Document EmbeddingsJohn Xavier Morris, Alexander M. RushICLR 2025
- DocKS-RAG: Optimizing Document-Level Relation Extraction through LLM-Enhanced Hybrid Prompt TuningXiaolong Xu, Yibo Zhou, Haolong Xiang, Xiaoyong Li 等ICML 2025
- A Graph-based Relevance Matching Model for Ad-hoc RetrievalYufeng Zhang, Jinghao Zhang, Zeyu Cui, Shu Wu 等AAAI 2021 · 被引用 26 次
- Exploiting Document Structures and Cluster Consistencies for Event Coreference ResolutionHieu Minh Tran, Duy Phung, Thien Huu NguyenACL 2021
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune RecipeChong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka 等NeurIPS 2025 · 被引用 4 次
