Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
Nay Myat Min, Long H. Pham, Yige Li, Jun Sun
摘要
Large language models (LLMs) can exhibit concept-conditioned semantic divergence: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like responses that evade token-trigger audits. This behavior falls in a blind spot of current safety evaluations, yet carries major societal stakes, as such concept cues can steer content exposure at scale. We formalize this phenomenon and present RAVEN (Response Anomaly Vigilance), a black-box audit that flags cases where a model is simultaneously highly certain and atypical among peers by coupling semantic entropy over paraphrastic samples with cross-model disagreement. In a controlled LoRA fine-tuning study, we implant a concept-conditioned stance using a small biased corpus, demonstrating feasibility without rare token triggers. Auditing five LLM families across twelve sensitive topics (360 prompts per model) and clustering via bidirectional entailment, RAVEN surfaces recurrent, model-specific divergences in 9/12 topics. Concept-level audits complement token-level defenses and provide a practical early-warning signal for release evaluation and post-deployment monitoring against propaganda-like influence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Adversarial Neuron Pruning Purifies Backdoored Deep ModelsDongxian Wu, Yisen WangNeurIPS 2021 · 被引用 441 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Poisoning Web-Scale Training Datasets is PracticalNicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka 等S&P 2024 · 被引用 309 次
相关 Paper
- Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech DetectionSanjeevan Selvaganapathy, Mehwish NasimACL 2026
- Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented ScanningShenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora 等CCS 2026
- What Lurks Within? Concept Auditing for Shared Diffusion Models at ScaleXiaoyong (Brian) Yuan, Xiaolong Ma, Linke Guo, Lan ZhangCCS 2025
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu 等AAAI 2025 · 被引用 6 次
- Same Question, Different Lies: Cross-Context Consistency (C³) for Black-Box Sandbagging DetectionYulong Lin, Pablo Bernabeu-Pérez, Benjamin Arnav, Lennie Wells 等ICML 2026
