Counterfactual LLM-based Framework for Measuring Rhetorical Style
Jingyi Qiu, Hong Chen, Zongyi Li
摘要
The rise of AI has fueled growing concerns about "hype" in machine learning papers, yet a reliable way to quantify rhetorical style independently of substantive content has remained elusive. Because bold language can stem from either strong empirical results or mere rhetorical style, it is often difficult to distinguish between the two. To disentangle rhetorical style from substantive content, we introduce a counterfactual, LLM-based framework: multiple LLM rhetorical personas generate counterfactual writings from the same substantive content, an LLM judge compares them through pairwise evaluations, and the outcomes are aggregated using a Bradley-Terry model. Applying this method to 8,485 ICLR submissions sampled from 2017 to 2025, we generate more than 250,000 counterfactual writings and provide a large-scale quantification of rhetorical style in ML papers. We find that visionary framing significantly predicts downstream attention, including citations and media attention, even after controlling for peer-review evaluations. We also observe a sharp rise in rhetorical strength after 2023, and provide empirical evidence showing that this increase is largely driven by the adoption of LLM-based writing assistance. The reliability of our framework is validated by its robustness to the choice of personas and the high correlation between LLM judgments and human annotations. Our work demonstrates that LLMs can serve as instruments to measure and improve scientific evaluation. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Sentence-Level and Aspect-Level (Un)certainty in Science CommunicationsJiaxin Pei, David JurgensEMNLP 2021 · 被引用 17 次
- The Noisy Path from Source to Citation: Measuring How Scholars Engage with Past ResearchHong Chen, Misha Teplitskiy, David JurgensACL 2025
- Measuring scalar constructs in social science with LLMsHauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel 等EMNLP 2025
相关 Paper
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing StylesKimberly Le Truong, Riccardo Fogliato, Hoda Heidari, Steven WuEMNLP 2025 · 被引用 2 次
- Narrative License and Model Sycophancy in LLM Summaries of Scientific WorkCalvin Isch, Grace JenningsACL 2026
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance RatesGiuseppe Russo, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky 等CSCW 2025 · 被引用 10 次
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer ReviewsWeixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp 等ICML 2024 · 被引用 213 次
- LLM or Human? Perceptions of Trust and Quality in Research SummariesNil-Jana Akpinar, Sandeep Avula, Chia-Jung Lee, Brandon Dang 等CHI 2026 · 被引用 2 次
