Hidden Secrets in the arXiv: Discovering, Analyzing, and Preventing Unintentional Information Disclosure in Source Files of Scientific Preprints
Jan Pennekamp, Johannes Lohmöller, David Schütte, Joscha Loos, Martin Henze
摘要
Preprints are essential for the timely and open dissemination of research. arXiv, the most widely used preprint service, takes the idea of open science one step further by not only publishing the actual preprints but also L A T E X sources and other files used to create them. As known from other contexts, such as GitHub repositories, and anecdotally exemplified for arXiv, making source code publicly available risks disclosing otherwise "hidden" information. Consequently, the public availability of paper sources raises the question of how much sensitive content is (unintentionally) disclosed through them.
In this paper, we systematically answer this question for all 2.7 M arXiv submissions with available source files across three dimensions of source file-induced information disclosure: 1 inclusion of unnecessary files, 2 metadata embedded in files, and 3 irrelevant content in files such as source code comments. Our analysis reveals that nearly every arXiv submission contains some form of "hidden" information. Notable findings range from links to editable web documents for internal coordination over API and private keys to complete Git histories.
While different tools promise to remove such information from source files, we show that they fail to reliably achieve the intended cleaning functionality. To mitigate this situation, we provide ALC-NG to comprehensively remove files, metadata, and comments that are not needed to compile a L A T E X paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 被引用 130 次
- Automated Detection of Password Leakage from Public GitHub RepositoriesRunhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan ZhangICSE 2022 · 被引用 36 次
- "As an AI language model, I cannot": Investigating LLM Denials of User RequestsJoel Wester, Tim Schrills, Henning Pohl, Niels van BerkelCHI 2024 · 被引用 35 次
- LATTE: Improving Latex Recognition for Tables and Formulae with Iterative RefinementNan Jiang, Shanchao Liang, Chengxiao Wang, Jiannan Wang 等AAAI 2025 · 被引用 8 次
- Leaky Apps: Large-scale Analysis of Secrets Distributed in Android and iOS AppsDavid Schmidt, Sebastian Schrittwieser, Edgar R. WeipplCCS 2025
相关 Paper
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 被引用 12 次
- arXivEdits: Understanding the Human Revision Process in Scientific WritingChao Jiang, Wei Xu, Samuel StevensEMNLP 2022 · 被引用 10 次
- Not All Those Who Share Are Lost: Analyzing 25 Years of Cybersecurity Artifact Sharing Practices Through Automated DiscoveryDaan Vansteenhuyse, Arthur Bols, Lieven Desmet, Victor Le Pochat 等USENIX Security 2026
- On the Privacy Risks Caused by Images in Academic PapersSze Yiu Chau, Jianyu Lu, Kin Man Leung, Chi Fung Kwan 等CCS 2026
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney 等ACL 2020 · 被引用 424 次
