Agentic Oversight via Dialectic Reasoning
Leonardo Ranaldi, Federico Ranaldi
Abstract
Debate has emerged as a promising oversight mechanism for Large Language Models (LLMs) amid rising systemic complexity, particularly where models outperform human evaluators. Yet, Debate provides little verifiable evidence for its final judgments, and its scalability across domains remains largely unexplored. To make oversight grounded and scale as capabilities extend, we introduce an Agentic Oversight framework. Specifically, by using Dialectic Argumentation as a reasoning function, we extend this paradigm to multilingual and multimodal spaces. We employ a weak-to-strong oversight approach based on two expert models that evaluate and defend contesting answers, while a third blind judge determines the winner using Dialectic Argumentation. Experts argue only for belief-consistent answers, founding the Debate on disagreements. We experimented with six tasks on our framework in both multilingual and multimodal scenarios, experimenting with different models. After indepth analysis, we show that dialectic argumentation consistently outperforms single-expert baselines. Moreover, we demonstrate that dialectic judgements from a weaker model deliver argument-mediated supervision that, via finetuning, instils unsupervised reasoning signals in expert models. Q: What is dish called? Options: a. Paris. b. Choux pastry c. Pastafrolla d. Ravioli e. Sfogliatella Question Answers & Debate b. Choux pastry e. Sfogliatella Round 1 #Premises P1: Each piece has a distinct fan shell shape with many ne, curved ridges from laminated layers. P2: The crust looks hard-baked and aky, strata are visibly separated. P3: Filling is visible at the open side; icing sugar is over the ridged shell. P4: Some pieces show dark amarena in custard typical of the Sant'Agata variant. #Argument The shell-shaped, ridged, laminated exterior is characteristic of Sfogliatella. The custard-and-cherry nish ts the Neapolitan Sfogliatella Santarosa style. #Premises P1: The items are individual pastries dusted with icing sugar. P2: Each pastry is lled with a pale custard-like cream and dotted with dark cherries. P3: The exterior looks airy and layered, suggesting a light dough. #Argument These appear to be choux-based pastries because they are small, cream-lled, and nished with icing sugar features common to éclairs and cream pus. The structure seems light and layered, which aligns with a steam-leavened dough. The cherries are consistent with patisserie nishes on choux items. Round 2 #Agreements. Accepts icing sugar and custard/cherries as non-diagnostic by themselves. #Disagreements & Challenges C1 Surface morphology shows parallel lamination forming a fan; these are not random cracks the lines are evenly spaced and curve from a narrow tip to a broad mouth. C2 The edge shows compressed, crisp layers, smooth crust of choux. C3 Choux pastries are typically split éclairs or proteroles;[.....] with visible stratied pastry. #Counter-observations O1: The hand-fan geometry is canonical for sfogliatella. O2: The crust appears crackly, with separable sheets a sign of laminated butter dough. O3: The presence of amarena cherries in custard is consistent with Sant'Agata, a recognised variant. #Agreements. Both accept there is custard, cherries, icing sugar. #Disagreements & Challenges C1 (P1,P2): The 'layers' could simply be cracks from baking a light dough. C2 (P4): Cherries and custard are generic patisserie elements. #Counter-observations O1: The pastries seem light and hollow near the lling, akin to choux cavities. O2: The overall size and individual portioning resemble cream pus.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang et al.NeurIPS 2023 · 463 citations
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis et al.ICML 2024 · 244 citations
Related papers
- On scalable oversight with weak LLMs judging strong LLMsZachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen et al.NeurIPS 2024 · 116 citations
- Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-AgentGiulia Pucci, Leonardo RanaldiEMNLP 2025
- Collaborative Disagreement Resolution for Scalable OversightYuyang Jiang, Chacha Chen, Teng Wu, Liwen Sun et al.ICML 2026
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and ReasoningHaocheng Yang, Fengxiang Cheng, Tianjun Yao, Mengyue Yang et al.ICLR 2026
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang et al.EMNLP 2024 · 177 citations
