NewsInterview: a Dataset and a Playground to Evaluate LLMs' Grounding Gap via Informational Interviews
Alexander Spangher, Michael Lu, Sriya Kalyan, Hyundong Justin Cho, Tenghao Huang, Weiyan Shi, Jonathan May
摘要
Large Language Models (LLMs) have demonstrated impressive capabilities in generating coherent text but often struggle with strategic dialogue. To address this gap, we focus on journalistic interviews. We curate a dataset of 40,000 two-person informational interviews from major news organizations in scenarios where human interviewers employ strategies to coax information from sources. We then try to mimic these activities with LLMs and find striking differences; models are less likely to use acknowledgments and more likely to rabbit-hole and not pivot to other topics. Real-izing that a fundamental deficit exists in LLM multi-turn planning and strategic thinking, we develop a realistic simulated environment, incorporating source personas and persuasive elements, in order to facilitate the development of agents with long-horizon rewards. Our experiments show that mimicry failures are not two-sided; when posing as a source, models adequately reflect human behavior in information sharing, making our simulation a realistic benchmark. Interviewer-LLMs, however, struggle with engaging persuasively, leading to sub-optimal information extraction across model size and capability. This simulated game lays the groundwork for future work in enhancing LLMs’ strategic dialogue capabilities. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- Optimizing Diversity and Quality through Base-Aligned Model CollaborationYichen Wang, Chenghao Yang, Tenghao Huang, Muhao Chen 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia 等EMNLP 2020 · 被引用 76 次
相关 Paper
- Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical StudyYuqi Zhu, Yi Zhong, Jintian Zhang, Ziheng Zhang 等AAAI 2026 · 被引用 3 次
- Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMsXuhui Zhou, Zhe Su, Tiwalayo Eisape, Hyunwoo Kim 等EMNLP 2024 · 被引用 5 次
- Towards Strategic Persuasion with Language ModelsZirui Cheng, Jiaxuan YouICLR 2026 · 被引用 11 次
- Interview: Large-scale Modeling of Media Dialog with Discourse Patterns and Knowledge GroundingBodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, Julian J. McAuleyEMNLP 2020 · 被引用 12 次
- Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsNicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas 等AAAI 2026 · 被引用 3 次
