CiteBench: A Benchmark for Scientific Citation Text Generation
Martin Funkquist, Ilia Kuznetsov, Yufang Hou, Iryna Gurevych
摘要
Science progresses by building upon the prior body of knowledge documented in scientific publications. The acceleration of research makes it hard to stay up-to-date with the recent developments and to summarize the evergrowing body of prior work. To address this, the task of citation text generation aims to produce accurate textual summaries given a set of papers-to-cite and the citing paper context. Due to otherwise rare explicit anchoring of cited documents in the citing paper, citation text generation provides an excellent opportunity to study how humans aggregate and synthesize textual knowledge from sources. Yet, existing studies are based upon widely diverging task definitions, which makes it hard to study this task systematically. To address this challenge, we propose CITEBENCH: a benchmark for citation text generation that unifies multiple diverse datasets and enables standardized evaluation of citation text generation models across task designs and domains. Using the new benchmark, we investigate the performance of multiple strong baselines, test their transferability between the datasets, and deliver new insights into the task definition and evaluation to guide future research in citation text generation. We make the code for CITEBENCH publicly available at https://github.com/ UKPLab/citebench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking ProcessMinjun Zhu, Yixuan Weng, Linyi Yang, Yue ZhangACL 2025 · 被引用 70 次
- Towards Verifiable Text Generation with Evolving Memory and Self-ReflectionHao Sun, Hengyi Cai, Bo Wang, Yingyan Hou 等EMNLP 2024 · 被引用 6 次
- L-CiteEval: A Suite for Evaluating Fidelity of Long-context ModelsZecheng Tang, Keyan Zhou, Juntao Li, Baibei Ji 等ACL 2025 · 被引用 5 次
- Improving Attributed Long-form Question Answering with Intent AwarenessXinran Zhao, Aakanksha Naik, Jay DeYoung, Joseph Chee Chang 等ICLR 2026 · 被引用 3 次
- Towards Verifiable Text Generation with Generative AgentBin Ji, Huijun Liu, Mingzhe Du, Shasha Li 等AAAI 2025 · 被引用 3 次
它引用的顶会 Paper7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation GenerationYifan Wang, Yiping Song, Shuai Li, Chaoran Cheng 等AAAI 2022 · 被引用 42 次
- Automatic Generation of Citation Texts in Scholarly Papers: A Pilot StudyXinyu Xing, Xiaosheng Fan, Xiaojun WanACL 2020 · 被引用 39 次
- CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited SupervisionYuning Mao, Ming Zhong, Jiawei HanEMNLP 2022 · 被引用 11 次
- Explaining Relationships Between Scientific DocumentsKelvin Luu, Xinyi Wu, Rik Koncel-Kedziorski, Kyle Lo 等ACL 2021
相关 Paper
- What Should I Cite? A RAG Benchmark for Academic Citation PredictionLeqi Zheng, Jiajun Zhang, Canzhi Chen, Chaokun Wang 等WWW 2026 · 被引用 2 次
- Systematic Task Exploration with LLMs: A Study in Citation Text GenerationFurkan Sahinuç, Ilia Kuznetsov, Yufang Hou, Iryna GurevychACL 2024
- Enhancing Scientific Papers Summarization with Citation GraphChenxin An, Ming Zhong, Yiran Chen, Danqing Wang 等AAAI 2021 · 被引用 47 次
- Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation PapersChen Tang, Shun Wang, Tomas Goldsack, Chenghua LinEMNLP 2023 · 被引用 5 次
- TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing PracticesGerrit Quaremba, Elizabeth Black, Denny Vrandecic, Elena SimperlICLR 2026 · 被引用 2 次
