HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
Andrew Zhuoer Feng, Cunxiang Wang, Yu Luo, Lin Fan, Irene Zhou, Zikang Wang, Xiaotao Gu, Jie Tang, Hongning Wang, Minlie Huang
Abstract
Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of existing metrics. LLM's performance in thousand-words level and open-ended writing is inadequately assessed by traditional reference-based metrics or modern LLM-as-a-judge methods. We propose Tree-of-Writing (ToW), to resolve the implicit inconsistency often found when LLM-as-a-judge aggregates all sub-features in text evaluation. ToW incorporates a tree-structured workflow by explicitly modeling the aggregation weights of sub-features. We also present HowToBench, a large-scale Chinese writing benchmark encompassing 12 genres and 1302 instructions across three task categories: contextual completion, outline-guided writing, and open-ended generation. ToW successfully mitigates the biases, achieving a 0.93 Pearson correlation with human judgments. Furthermore, we detect that both overlap-based text generation metrics and popular LLM-as-a-judge practices are vulnerable to textual disturbances, while ToW is robust to them. We also uncover a negative correlation between input length and content-related scores in the Guide task, showcasing that it cannot be simply improved by input-side information piling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70e8002d-c662-4e08-8b00-033349591f88Builds on4
- On the Limitations of Reference-Free Evaluations of Generated TextDaniel Deutsch, Rotem Dror, Dan RothEMNLP 2022 · 23 citations
- DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and RewritingXuanming Zhang, Anthony Diaz, Zixun Chen, Qingyang Wu et al.EMNLP 2024 · 2 citations
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
- JudgeLM: Fine-tuned Large Language Models are Scalable JudgesLianghui Zhu, Xinggang Wang, Xinlong WangICLR 2025
Related papers
- AlignBench: Benchmarking Chinese Alignment of Large Language ModelsXiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang et al.ACL 2024 · 9 citations
- ToMBench: Benchmarking Theory of Mind in Large Language ModelsZhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen et al.ACL 2024 · 6 citations
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
- WaterBench: Towards Holistic Evaluation of Watermarks for Large Language ModelsShangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu et al.ACL 2024
