SEAL: Structure and Element Aware Learning Improves Long Structured Document Retrieval
Xinhao Huang, Zhibo Ren, Yipeng Yu, Ying Zhou, Zulong Chen, Zeyi Wen
Abstract
In long structured document retrieval, existing methods typically fine-tune pre-trained language models (PLMs) using contrastive learning on datasets lacking explicit structural information. This practice suffers from two critical issues: 1) current methods fail to leverage structural features and element-level semantics effectively, and 2) the lack of datasets containing structural metadata. To bridge these gaps, we propose SEAL, a novel contrastive learning framework. It leverages structure-aware learning to preserve semantic hierarchies and masked element alignment for fine-grained semantic discrimination. Furthermore, we release StructDocRetrieval, a long structured document retrieval dataset with rich structural annotations. Extensive experiments on both released and industrial datasets across various modern PLMs, along with online A/B testing, demonstrate consistent performance improvements, boosting NDCG@10 from 79.41% to 82.59% on BGE-M3. The resources are available at this URL. * Equal Contribution † Corresponding Author Query Structured Documents <h1>Python Introduction<h1> <p> Python is a high-level programming language. Getting started: <p> Could you give me an example of a Python program to get started? Python Introduction Python [CLS] is a high-level programming [CLS] language. Getting started: [CLS] is a high-level programming language. Getting started:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3dce9508-f424-4cb4-b151-c53c7d977df9Builds on11
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu et al.EMNLP 2022 · 25 citations
- MGS3: A Multi-Granularity Self-Supervised Code Search FrameworkRui Li, Junfeng Kang, Qi Liu, Liyang He et al.KDD 2025
- SEAL: Semantic-Aware Hierarchical Learning for Generalized Category DiscoveryZhenqi He, Yuanpei Liu, Kai HanNeurIPS 2025 · 10 citations
- DOGR: Leveraging Document-Oriented Contrastive Learning in Generative RetrievalPenghao Lu, Xin Dong, Yuansheng Zhou, Lei Cheng et al.AAAI 2025
- ContraCLM: Contrastive Learning For Causal Language ModelNihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang et al.ACL 2023 · 4 citations
