BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
Yoonseo Choi, Eunhye Kim, Hyunwoo Kim, Donghyun Park, Honggu Lee, Jin Young Kim, Juho Kim
摘要
If 100 people issue the same search query, they may have 100 different goals. While existing work on user-centric AI evaluation highlights the importance of aligning systems with fine-grained user intents, current search evaluation methods struggle to represent and assess this diversity. We introduce BloomIntent, a user-centric search evaluation method that uses user intents as the evaluation unit. BloomIntent first generates a set of plausible, fine-grained search intents grounded on taxonomies of user attributes and information-seeking intent types. Then, BloomIntent provides an automated evaluation of search results against each intent powered by large language models. To support practical analysis, BloomIntent clusters semantically similar intents and summarizes evaluation outcomes in a structured interface. With three technical evaluations, we showed that BloomIntent generated fine-grained, evaluable, and realistic intents and produced scalable assessments of intent-level satisfaction that achieved 72% agreement with expert evaluators. In a case study (N=4), we showed that BloomIntent supported search specialists in identifying intents for ambiguous queries, uncovering underserved user needs, and discovering actionable insights for improving search experiences. By shifting from query-level to intent-level evaluation, BloomIntent reimagines how search systems can be assessed—not only for performance but for their ability to serve a multitude of user goals.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation CriteriaAnnalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A. Metoyer 等CHI 2026 · 被引用 1 次
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and ConsiderationsKarla Felix Navarro, Eugene Syriani, Ian ArawjoCHI 2026 · 被引用 1 次
它引用的顶会 Paper20
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 被引用 153 次
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran 等UIST 2024 · 被引用 143 次
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis TestingIan Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg 等CHI 2024 · 被引用 141 次
相关 Paper
- BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic FeedbackHyunseo Kim, Sangam Lee, Kwangwook Seo, Dongha LeeICML 2026
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun 等EMNLP 2024 · 被引用 10 次
- Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom's TaxonomyFei Zhang, Zhe Zhao, Haibin Wen, Tianshuo Wei 等ACL 2026
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning 等ICML 2025
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 被引用 1 次
