Idiosyncrasies in Large Language Models
Mingjie Sun, Yida Yin, Zhiqiu Xu, J. Zico Kolter, Zhuang Liu
摘要
In this work, we unveil and study idiosyncrasies in Large Language Models (LLMs) -unique patterns in their outputs that can be used to distinguish the models. To do so, we consider a simple classification task: given a particular text output, the objective is to predict the source LLM that generates the text. We evaluate this synthetic task across various groups of LLMs and find that simply fine-tuning text embedding models on LLMgenerated texts yields excellent classification accuracy. Notably, we achieve 97.1% accuracy on held-out validation data in the five-way classification problem involving ChatGPT, Claude, Grok, Gemini, and DeepSeek. Our further investigation reveals that these idiosyncrasies are rooted in word-level distributions. These patterns persist even when the texts are rewritten, translated, or summarized by an external LLM, suggesting that they are also encoded in the semantic content. Additionally, we leverage LLM as judges to generate detailed, open-ended descriptions of each model's idiosyncrasies. Finally, we discuss the broader implications of our findings, including training on synthetic data, inferring model similarity, and robust evaluation of LLMs. Code is available at github.com/locuslab/llm-idiosyncrasies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Auditing Black-Box LLM APIs with a Rank-Based Uniformity TestXiaoyuan Zhu, Yaowen Ye, Tianyi Qiu, Hanlin Zhu 等ICLR 2026 · 被引用 22 次
- Cost-Aware Contrastive Routing for LLMsReza Shirkavand, Shangqian Gao, Peiran Yu, Heng HuangNeurIPS 2025 · 被引用 19 次
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu 等ICLR 2026 · 被引用 15 次
- Log Probability Tracking of LLM APIsTimothee Chauvin, Erwan Le Merrer, Francois Taiani, Gilles TredanICLR 2026 · 被引用 12 次
- Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language ModelsChantal Shaib, Vinith M. Suriyakumar, Byron C. Wallace, Marzyeh GhassemiNeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper15
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
相关 Paper
- Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language ModelsSergio E. Zanotto, Segun AroyehunEMNLP 2025 · 被引用 3 次
- Model-Agnostic Sentiment Distribution Stability Analysis for Robust LLM-Generated Texts DetectionSiyuan Li, Xi Lin, Guangyan Li, Zehao Liu 等AAAI 2026
- Continual Origin Tracing of LLM-Generated TextHaoran Li, Quan WangSIGIR 2025 · 被引用 2 次
- Quantification of Large Language Model DistillationSunbowen Lee, Junting Zhou, Chang Ao, Kaige Li 等ACL 2025
- Profiler: Black-box AI-generated Text Origin Detection via Context-aware Inference Pattern AnalysisHanxi Guo, Siyuan Cheng, Xiaolong Jin, Zhuo Zhang 等EMNLP 2025
