Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs
Melika Mobini, Vincent Holst, Floriano Tori, Andres Algaba, Vincent Ginis
Abstract
Large language models are increasingly used to curate bibliographies, raising the question: are their reference lists distinguishable from human ones? We build paired citation graphs, ground truth and GPT-4o-generated (from parametric knowledge), for 10,000 focal papers ( 275k references) from SciSciNet, and added a field-matched random baseline that preserves out-degree and field distributions while breaking latent structure. We compare (i) structure-only node features (degree/closeness/eigenvector centrality, clustering, edge count) with (ii) 3072-D title/abstract embeddings, using an RF on graph-level aggregates and Graph Neural Networks with node features. Structure alone barely separates GPT from ground truth (RF accuracy 0.60) despite cleanly rejecting the random baseline ( 0.89--0.92). By contrast, embeddings sharply increase separability: RF on aggregated embeddings reaches 0.83, and GNNs with embedding node features achieve 93% test accuracy on GPT vs. ground truth. We show the robustness of our findings by replicating the pipeline with Claude Sonnet 4.5 and with multiple embedding models (OpenAI and SPECTER), with RF separability for ground truth vs. Claude and clean rejection of the random baseline. Thus, LLM bibliographies, generated purely from parametric knowledge, closely mimic human citation topology, but leave detectable semantic fingerprints; detection and debiasing should target content signals rather than global graph structure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9720cc51-126d-4548-ad16-971277cffdc8Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Fair Comparison of Graph Neural Networks for Graph ClassificationFederico Errica, Marco Podda, Davide Bacciu, Alessio MicheliICLR 2020 · 508 citations
- LLM-Check: Investigating Detection of Hallucinations in Large Language ModelsGaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha et al.NeurIPS 2024 · 170 citations
- SciRepEval: A Multi-Format Benchmark for Scientific Document RepresentationsAmanpreet Singh, Mike D'Arcy, Arman Cohan, Doug Downey et al.EMNLP 2023 · 45 citations
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and ReadingYutao Wu, Xiao Liu, Yunhao Feng, Jiale Ding et al.WWW 2026 · 1 citation
Related papers
- People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textJenna Russell, Marzena Karpinska, Mohit IyyerACL 2025 · 39 citations
- On Fake News Detection with LLM Enhanced Semantics MiningXiaoxiao Ma, Yuchen Zhang, Kaize Ding, Jian Yang et al.EMNLP 2024 · 23 citations
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp et al.ICML 2024 · 213 citations
- DEMASQ: Unmasking the ChatGPT WordsmithKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiNDSS 2024
- On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic WritingZeyan Liu, Zijun Yao, Fengjun Li, Bo LuoCCS 2024 · 18 citations
