Structure and Semantics Preserving Document Representations
Natraj Raman, Sameena Shah, Manuela Veloso
Abstract
Retrieving relevant documents from a corpus is typically based on the semantic similarity between the document content and query text. The inclusion of structural relationship between documents can benefit the retrieval mechanism by addressing semantic gaps. However, incorporating these relationships requires tractable mechanisms that balance structure with semantics and take advantage of the prevalent pre-train/fine-tune paradigm. We propose here a holistic approach to learning document representations by integrating intra-document content with inter-document relations. Our deep metric learning solution analyzes the complex neighborhood structure in the relationship network to efficiently sample similar/dissimilar document pairs and defines a novel quintuplet loss function that simultaneously encourages document pairs that are semantically relevant to be closer and structurally unrelated to be far apart in the representation space. Furthermore, the separation margins between the documents are varied flexibly to encode the heterogeneity in relationship strengths. The model is fully fine-tunable and natively supports query projection during inference. We demonstrate that it outperforms competing methods on multiple datasets for document retrieval tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d80f1f1-c051-477f-a939-d01ce74fd0dcCited by top-tier papers1
Ask how each one uses itBuilds on7
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
- Ladder Loss for Coherent Visual-Semantic EmbeddingMo Zhou, Zhenxing Niu, Le Wang, Zhanning Gao et al.AAAI 2020 · 46 citations
- Which *BERT? A Survey Organizing Contextualized EncodersPatrick Xia, Shijie Wu, Benjamin Van DurmeEMNLP 2020 · 44 citations
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey et al.ACL 2020 · 20 citations
Related papers
- Contextual Document EmbeddingsJohn Xavier Morris, Alexander M. RushICLR 2025
- DocKS-RAG: Optimizing Document-Level Relation Extraction through LLM-Enhanced Hybrid Prompt TuningXiaolong Xu, Yibo Zhou, Haolong Xiang, Xiaoyong Li et al.ICML 2025
- A Graph-based Relevance Matching Model for Ad-hoc RetrievalYufeng Zhang, Jinghao Zhang, Zeyu Cui, Shu Wu et al.AAAI 2021 · 26 citations
- Exploiting Document Structures and Cluster Consistencies for Event Coreference ResolutionHieu Minh Tran, Duy Phung, Thien Huu NguyenACL 2021
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune RecipeChong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka et al.NeurIPS 2025 · 4 citations
