DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios
Tongzhou Yu, Mingjia Li, Hong Qian, Wenkai Wang, Zongbao Zhang, Yaoyu Jiang, Xiangfeng Wang, Aimin Zhou, Jiajun Guo
摘要
Human creativity has emerged as a critical competency in the era of large language models. Assessing creativity in complex, open-ended environments is a grand challenge in data mining, currently hindered by a reliance on standardized simple tasks and the scarcity of fine-grained expert data. As an ecologically valid assessment context, debate reflects multiple dimensions of creativity, encompassing both divergent thinking and convergent thinking. Moreover, debate is a data-rich domain, with a large volume of publicly accessible materials. Current mainstream automated scoring methods are poorly suited to complex settings such as debate, and therefore still rely on costly human evaluation. To this end, this paper proposes DEFINED, a data-efficient computational framework for fine-grained creativity assessment in debate scenarios. DEFINED operationalizes debate creativity through a hierarchical eight-dimensional metric system, implemented via a pre-trained autoregressive language model with a hierarchical scoring head that supports both fine-grained and coarse-grained evaluation. Statements and their associated expert scores were obtained from authentic debate competitions, and a constrained data augmentation strategy was employed to address the elite bias inherent in the original data. DEFINED adopts a mixed-granularity training strategy enabling robust learning from limited fine-grained supervision annotated by trained graduate experts. To rigorously validate ecological validity beyond synthetic benchmarks, we incorporate an empirical study with debate-naive participants, utilizing these authentic data to serve as a qualitative case study for mid-to-low proficiency populations. Across our evaluation protocol, our scoring model achieves accurate and stable scoring, outperforming prompt-based large language model evaluators and existing debate scoring methods, while mitigating common failure modes observed in current approaches. The code for DEFINED is available on GitHub at https://github.com/tzwo/DEFINED.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- A Large-Scale Dataset for Argument Quality Ranking: Construction and AnalysisShai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo 等AAAI 2020 · 被引用 148 次
- Evaluating Text Creativity across Diverse Domains: a Dataset and Large Language Model EvaluatorQian Cao, Xiting Wang, Yuzhuo Yuan, Yahui Liu 等ICLR 2026 · 被引用 11 次
- Orchid: Flexible and Data-Dependent Convolution for Sequence ModelingMahdi Karami, Ali GhodsiNeurIPS 2024 · 被引用 10 次
- InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for DebatingFuyu Wang, Jiangtong Li, Kun Zhu, Changjun JiangACL 2025 · 被引用 3 次
相关 Paper
- Automated Creativity Evaluation of Language Models Across Open-Ended TasksTan Min Sen, Zachary Choy Kit Chun, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya 等ACL 2026
- Analyze-Compose-Execute: A Dynamic Dialogue Framework for Multi-Agent DebateWenyuan Gu, Haowen Wang, Jiale Han, Xiang Li 等AAAI 2026
- Systematic Task Exploration with LLMs: A Study in Citation Text GenerationFurkan Sahinuç, Ilia Kuznetsov, Yufang Hou, Iryna GurevychACL 2024
- iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM InferenceWei Fan, JinYi Yoon, Bo JiAAAI 2026 · 被引用 5 次
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang 等NeurIPS 2025 · 被引用 2 次
