Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
Arthur Amalvy, Vincent Labatut, Xavier Bost, Hen-Hsen Huang
Abstract
While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such corpora are necessary to fully represent the diversity of data found in the wild in the context of NLP tasks. We tackle this issue by proposing a method to lawfully and publicly share the annotations of copyrighted literary texts. The corpus creator shares the annotations in clear, along with a non-reversible hashed version of the source material. The corpus user must own the source material, and apply the same hash function to their own tokens, in order to match them to the shared annotations. Crucially, our method is robust to reasonable divergences in the version of the copyrighted data owned by the user. As an illustration, we present alignment experiments on different editions of novels. Our results show that our method is able to correctly align 98.7 to 99.79% of tokens depending on the novel, provided the user version is sufficiently close to the corpus creator's version. We publicly release novelshare, a Python implementation of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal et al.NeurIPS 2021 · 315 citations
Related papers
- NLP Reproducibility For All: Understanding Experiences of BeginnersShane Storks, Keunwoo Peter Yu, Ziqiao Ma, Joyce ChaiACL 2023
- Estimating Agreement by Chance for Sequence AnnotationDiya Li, Carolyn P. Rosé, Ao Yuan, Chunxiao ZhouACL 2024
- Allign: Aligning All-Pair Near-Duplicate Passages in Long TextsWeiqi Feng, Dong DengSIGMOD 2021 · 13 citations
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model GenerationTong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min et al.EMNLP 2024 · 4 citations
- GNAT: A General Narrative Alignment ToolTanzir Pial, Steven SkienaEMNLP 2023 · 3 citations
