Hidden Secrets in the arXiv: Discovering, Analyzing, and Preventing Unintentional Information Disclosure in Source Files of Scientific Preprints
Jan Pennekamp, Johannes Lohmöller, David Schütte, Joscha Loos, Martin Henze
Abstract
Preprints are essential for the timely and open dissemination of research. arXiv, the most widely used preprint service, takes the idea of open science one step further by not only publishing the actual preprints but also L A T E X sources and other files used to create them. As known from other contexts, such as GitHub repositories, and anecdotally exemplified for arXiv, making source code publicly available risks disclosing otherwise "hidden" information. Consequently, the public availability of paper sources raises the question of how much sensitive content is (unintentionally) disclosed through them.
In this paper, we systematically answer this question for all 2.7 M arXiv submissions with available source files across three dimensions of source file-induced information disclosure: 1 inclusion of unnecessary files, 2 metadata embedded in files, and 3 irrelevant content in files such as source code comments. Our analysis reveals that nearly every arXiv submission contains some form of "hidden" information. Notable findings range from links to editable web documents for internal coordination over API and private keys to complete Git histories.
While different tools promise to remove such information from source files, we show that they fail to reliably achieve the intended cleaning functionality. To mitigate this situation, we provide ALC-NG to comprehensively remove files, metadata, and comments that are not needed to compile a L A T E X paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7977795-70d5-4c11-8e6e-82c0db1e4b49Builds on7
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 130 citations
- Automated Detection of Password Leakage from Public GitHub RepositoriesRunhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan ZhangICSE 2022 · 36 citations
- "As an AI language model, I cannot": Investigating LLM Denials of User RequestsJoel Wester, Tim Schrills, Henning Pohl, Niels van BerkelCHI 2024 · 35 citations
- LATTE: Improving Latex Recognition for Tables and Formulae with Iterative RefinementNan Jiang, Shanchao Liang, Chengxiao Wang, Jiannan Wang et al.AAAI 2025 · 8 citations
- Leaky Apps: Large-scale Analysis of Secrets Distributed in Android and iOS AppsDavid Schmidt, Sebastian Schrittwieser, Edgar R. WeipplCCS 2025
Related papers
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- arXivEdits: Understanding the Human Revision Process in Scientific WritingChao Jiang, Wei Xu, Samuel StevensEMNLP 2022 · 10 citations
- Not All Those Who Share Are Lost: Analyzing 25 Years of Cybersecurity Artifact Sharing Practices Through Automated DiscoveryDaan Vansteenhuyse, Arthur Bols, Lieven Desmet, Victor Le Pochat et al.USENIX Security 2026
- On the Privacy Risks Caused by Images in Academic PapersSze Yiu Chau, Jianyu Lu, Kin Man Leung, Chi Fung Kwan et al.CCS 2026
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
