InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating
Fuyu Wang, Jiangtong Li, Kun Zhu, Changjun Jiang
摘要
With the rapid advancements in large language models (LLMs), debating tasks, such as argument quality assessment and debate process simulation, have made significant progress. However, existing LLM-based debating systems focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity. Furthermore, these systems lack a structured approach to optimize across various dimensionsincluding evaluation metrics, chain-of-thought (CoT) reasoning, and multi-turn debate refinementthereby limiting their effectiveness. To address these interconnected challenges, we propose a dual-component framework: (1) , a novel evaluation system that establishes a multi-dimensional assessment architecture incorporating four subjective criteria (emotional appeal, argument clarity, argument arrangement, and topic relevance) alongside two objective metrics (fact authenticity and logical validity); and (2) , an optimized debating framework employing a phased optimization approach through CoT reasoning enhancement, multi-dimensional Direct Preference Optimization (DPO), and real-time knowledge grounding via web-based Retrieval Augmented Generation (Web-RAG). Empirical evaluations demonstrate that achieves 44 higher correlation with expert judgments compared to existing methods, while shows significant improvements, outperforming baseline models by 57. Source code is available at https://github.com/fywang12/InspireDebate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu 等NeurIPS 2024 · 被引用 182 次
- Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?Haoang Chi, He Li, Wenjing Yang, Feng Liu 等NeurIPS 2024 · 被引用 124 次
相关 Paper
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsAndries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett 等ICML 2024 · 被引用 82 次
- DEBATE, TRAIN, EVOLVE: Self-Evolution of Language Model ReasoningGaurav Srivastava, Zhenyu Bi, Meng Lu, Xuan WangEMNLP 2025 · 被引用 1 次
- M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation EvaluationZhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren 等ACL 2025
- Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language ModelsChen Han, Wenzhen Zheng, Xijin TangEMNLP 2025 · 被引用 2 次
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope 等EMNLP 2025 · 被引用 1 次
