Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-Agent
Giulia Pucci, Leonardo Ranaldi
摘要
Large language models (LLMs) have demonstrated capabilities that are highly satisfactory to a wide range of users by adapting to their culture and wisdom. Yet, this can translate into a propensity to produce responses that align with users' viewpoints, even when the latter are wrong. This behaviour is known as sycophancy, the tendency of LLMs to generate misleading responses as long as they align with the user's, inducing bias and reducing reliability. To make interactions consistent, reliable and safe, we introduce X-Agent, an Oversight Reasoning framework that audits human-model dialogues, reasons about them, captures sycophancy and corrects the final outputs. Concretely, X-Agent extends debate-based frameworks by (i) auditing user-model conversations, (ii) applying a defence layer that steers model behaviour and goes beyond user beliefs, and (iii) extracting reasoning traces from evaluations that serve as training signals for mitigating sycophancy, all in a completely unsupervised way. We evaluate X-Agent across diverse scenarios and languages, showing that it consistently detects sycophancy, reduces unwarranted agreement, and improves crossturn consistency, advancing a reasoning-asoverview paradigm for safer user-model and model-model interaction. Our approach introduces a novel paradigm in which reasoning is not merely a means to solve problems, but as a mechanism for overseeing the problem-solving processes of other models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud 等ICLR 2024 · 被引用 762 次
- On scalable oversight with weak LLMs judging strong LLMsZachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen 等NeurIPS 2024 · 被引用 116 次
- Improving Chain-of-Thought Reasoning via Quasi-Symbolic AbstractionsLeonardo Ranaldi, Marco Valentino, André FreitasACL 2025 · 被引用 29 次
- R2-MultiOmnia: Leading Multilingual Multimodal Reasoning via Self-TrainingLeonardo Ranaldi, Federico Ranaldi, Giulia PucciACL 2025 · 被引用 9 次
- Self-Refine Instruction-Tuning for Aligning Reasoning in Language ModelsLeonardo Ranaldi, André FreitasEMNLP 2024 · 被引用 3 次
相关 Paper
- Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-AuditingWenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu 等ACL 2026 · 被引用 3 次
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving TasksJessica Y. Bo, Majeed Kazemitabaar, Mengqing Deng, Michael Inzlicht 等CHI 2026 · 被引用 8 次
- Evaluating and Mitigating Sycophancy in Large Vision-Language ModelsJiayi Gao, Huaiwen ZhangACM MM 2025
- HalluClean: A Unified Framework to Combat Hallucinations in LLMsYaxin Zhao, Yu ZhangAAAI 2026
- RedDebate: Safer Responses Through Multi-Agent Red Teaming DebatesAli Asad, Stephen Obadinma, Radin Shayanfar, Xiaodan ZhuICML 2026 · 被引用 7 次
