Artificial Text Detection via Examining the Topology of Attention Maps
Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev
Abstract
The impressive capabilities of recent generative models to create texts that are challenging to distinguish from the human-written ones can be misused for generating fake news, product reviews, and even abusive content. Despite the prominent performance of existing methods for artificial text detection, they still lack interpretability and robustness towards unseen models. To this end, we propose three novel types of interpretable topological features for this task based on Topological Data Analysis (TDA) which is currently understudied in the field of NLP. We empirically show that the features derived from the BERT model outperform count-and neural-based baselines up to 10% on three common datasets, and tend to be the most robust towards unseen GPT-style generation models as opposed to existing methods. The probing analysis of the features reveals their sensitivity to the surface and syntactic properties. The results demonstrate that TDA is a promising line with respect to NLP tasks, specifically the ones that incorporate surface and structural information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated TextsEduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii et al.NeurIPS 2023 · 163 citations
- Do Language Models Plagiarize?Jooyoung Lee, Thai Le, Jinghui Chen, Dongwon LeeWWW 2023 · 109 citations
- On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic WritingZeyan Liu, Zijun Yao, Fengjun Li, Bo LuoCCS 2024 · 18 citations
- Hallucination Detection in LLMs with Topological Divergence on Attention GraphsAlexandra Bazarova, Andrei Volodichev, Aleksandr Yugay, Andrey Shulga et al.ACL 2026 · 14 citations
- Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated TextMazal Bethany, Brandon Wherry, Emet Bethany, Nishant Vishwamitra et al.USENIX Security 2024 · 13 citations
Builds on5
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 110 citations
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 32 citations
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 21 citations
- Computing the Testing Error Without a Testing SetCiprian A. Corneanu, Sergio Escalera, Aleix M. MartinezCVPR 2020
Related papers
- Neural Deepfake Detection with Factual Structure of TextWanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang et al.EMNLP 2020 · 43 citations
- DEMASQ: Unmasking the ChatGPT WordsmithKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiNDSS 2024
- Deepfake Text Detection: Limitations and OpportunitiesJiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman et al.S&P 2023
- Non-Existent Relationship: Fact-Aware Multi-Level Machine-Generated Text DetectionYang Wu, Ruijia Wang, Jie WuEMNLP 2025
- On Fake News Detection with LLM Enhanced Semantics MiningXiaoxiao Ma, Yuchen Zhang, Kaize Ding, Jian Yang et al.EMNLP 2024 · 23 citations
