Artificial Text Detection via Examining the Topology of Attention Maps
Laida Kushnareva, Daniil Cherniavskii, Vladislav Mikhailov, Ekaterina Artemova, Serguei Barannikov, Alexander Bernstein, Irina Piontkovskaya, Dmitri Piontkovski, Evgeny Burnaev
摘要
The impressive capabilities of recent generative models to create texts that are challenging to distinguish from the human-written ones can be misused for generating fake news, product reviews, and even abusive content. Despite the prominent performance of existing methods for artificial text detection, they still lack interpretability and robustness towards unseen models. To this end, we propose three novel types of interpretable topological features for this task based on Topological Data Analysis (TDA) which is currently understudied in the field of NLP. We empirically show that the features derived from the BERT model outperform count-and neural-based baselines up to 10% on three common datasets, and tend to be the most robust towards unseen GPT-style generation models as opposed to existing methods. The probing analysis of the features reveals their sensitivity to the surface and syntactic properties. The results demonstrate that TDA is a promising line with respect to NLP tasks, specifically the ones that incorporate surface and structural information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated TextsEduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii 等NeurIPS 2023 · 被引用 163 次
- Do Language Models Plagiarize?Jooyoung Lee, Thai Le, Jinghui Chen, Dongwon LeeWWW 2023 · 被引用 109 次
- On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic WritingZeyan Liu, Zijun Yao, Fengjun Li, Bo LuoCCS 2024 · 被引用 18 次
- Hallucination Detection in LLMs with Topological Divergence on Attention GraphsAlexandra Bazarova, Andrei Volodichev, Aleksandr Yugay, Andrey Shulga 等ACL 2026 · 被引用 14 次
- Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated TextMazal Bethany, Brandon Wherry, Emet Bethany, Nishant Vishwamitra 等USENIX Security 2024 · 被引用 13 次
它引用的顶会 Paper5
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 被引用 110 次
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 被引用 32 次
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 被引用 21 次
- Computing the Testing Error Without a Testing SetCiprian A. Corneanu, Sergio Escalera, Aleix M. MartinezCVPR 2020
相关 Paper
- Neural Deepfake Detection with Factual Structure of TextWanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang 等EMNLP 2020 · 被引用 43 次
- DEMASQ: Unmasking the ChatGPT WordsmithKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiNDSS 2024
- Deepfake Text Detection: Limitations and OpportunitiesJiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman 等S&P 2023
- Non-Existent Relationship: Fact-Aware Multi-Level Machine-Generated Text DetectionYang Wu, Ruijia Wang, Jie WuEMNLP 2025
- On Fake News Detection with LLM Enhanced Semantics MiningXiaoxiao Ma, Yuchen Zhang, Kaize Ding, Jian Yang 等EMNLP 2024 · 被引用 23 次
