Lune

SIGIR2025Top-tier venue

Continual Origin Tracing of LLM-Generated Text

Haoran Li, Quan Wang

2025Year
2Citations
1Top-tier citations

Abstract

The rapid development of large language models (LLMs) raises concerns about their potential misuse. Accurately identifying and tracing the origin of LLM-generated content is crucial for accountability and transparency. Previous methods typically frame origin tracing as multi-class classification with a fixed label set, thus struggle to adapt to new LLMs without frequent retraining. This paper introduces a new task, continual origin tracing of LLM-generated text, which frames origin tracing in a continual learning or, more precisely, class-incremental learning manner, where new LLMs continuously emerge, and a model incrementally learns to identify new LLMs without forgetting old ones. A novel training-free method is further devised for the task, which continually extracts prototypes for emerging LLMs using a frozen pre-trained model, and conducts global and local prototype decorrelation to improve prototype matching, thus favoring more accurate tracing. To facilitate evaluation on the new task, we construct a benchmark comprising text generated by 19 recently released LLMs from 12 vendors that simulates a real-world scenario where these LLMs emerge over time and need to be recognized incrementally across 8 diverse domains. Rigorous evaluations on this benchmark highlight the effectiveness and potential of the proposed method in the new task, offering a promising direction for future research.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get cfc25bdd-aa89-4e48-bbdc-6f93f9472e34

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines