From Chat Logs to Collective Insights: Aggregative Question Answering
Wentao Zhang, Woojeong Kim, Yuntian Deng
Abstract
Conversational agents powered by large language models (LLMs) are rapidly becoming integral to our daily interactions, generating unprecedented amounts of conversational data.Such datasets offer a powerful lens into societal interests, trending topics, and collective concerns.Yet, existing approaches typically treat these interactions as independent and miss critical insights that could emerge from aggregating and reasoning across large-scale conversation logs.In this paper, we introduce Aggregative Question Answering, a novel task requiring models to reason over thousands of user-chatbot interactions to answer aggregative queries, such as identifying emerging concerns among specific demographics.To enable research in this direction, we constructed WildChat-AQA, a benchmark comprising 6,027 aggregative questions derived from 182,330 real-world chatbot conversations.Experiments show that existing methods either struggle to reason effectively or incur prohibitive computational costs, underscoring the need for new approaches capable of extracting collective insights from large-scale conversational data.Figure 2: Overview of the WildChat-AQA dataset creation process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question AnsweringQingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha et al.EMNLP 2024 · 13 citations
Related papers
- QAConv: Question Answering on Informative ConversationsChien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung et al.ACL 2022 · 34 citations
- WideSearch: Benchmarking Agentic Broad Info-SeekingRyan Wong, Jiawei Wang, Junjie Zhao, Li Chen et al.ICLR 2026 · 66 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
- WildFeedback: Aligning LLMs With In-situ User Interactions And FeedbackTaiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin et al.ACL 2026 · 35 citations
