WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
Zhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei, Xiaoqian Shen, Liangbing Zhao, Chun-Mei Feng, Mohamed Elhoseiny
摘要
Knowledge discovery and collection are intelligenceintensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style article generation by retrieving and synthesizing information from the internet. However, these methods primarily focus on text-only generation, overlooking the importance of multimodal content in enhancing informativeness and engagement. In this work, we introduce WikiAutoGen, a novel system for automated multimodal Wikipedia-style article generation. Unlike prior approaches, WikiAutoGen retrieves and integrates relevant images alongside text, enriching both the depth and visual appeal of generated content. To further improve factual accuracy and comprehensiveness, we propose a multiperspective self-reflection mechanism, which critically assesses retrieved content from diverse viewpoints to enhance reliability, breadth, and coherence, etc. Additionally, we introduce WikiSeek, a benchmark comprising Wikipedia articles with topics paired with both textual and image-based representations, designed to evaluate multimodal knowledge generation on more challenging topics. Experimental results show that WikiAutoGen outperforms previous methods by 8%-29% on our WikiSeek benchmark, producing more accurate, coherent, and visually enriched Wikipediastyle articles. Our code and examples are available at https://wikiautogen.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent CollaborationZhongyu Yang, Zuhao Yang, Shuo Zhan, Tan Yue 等CVPR 2026 · 被引用 5 次
- InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent CollaborationZhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An 等AAAI 2026 · 被引用 5 次
- XR: Cross-Modal Agents for Composed Image RetrievalZhongyu Yang, Wei Pang, Yingfang YuanWWW 2026 · 被引用 1 次
- WikiREVIEW: A Multi-Perspective Review Framework for Automatic Wiki-Style Article GenerationGuo-Biao Zhang, Zhijing Wu, Tian Lan, Ding-Yuan Liu 等AAAI 2026
- MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion RecognitionZhongyu Yang, Junhao Song, Siyang Song, Wei Pang 等EMNLP 2025
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen 等ICLR 2024 · 被引用 699 次
相关 Paper
- WikiMAG: A Multi-Agent Guided Framework for Generating Structured Wikipedia-like ArticlesXiuli Kang, Yinlong Xiao, Minghao Hu, Yuan Huang 等AAAI 2026
- Hierarchical Memory Organization for Wikipedia GenerationEugene J. Yu, Dawei Zhu, Yifan Song, Xiangyu Wong 等ACL 2025
- XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesDhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian 等WWW 2023 · 被引用 3 次
- A Suite of Generative Tasks for Multi-Level Multimodal Webpage UnderstandingAndrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown 等EMNLP 2023 · 被引用 2 次
- Descartes: Generating Short Descriptions of Wikipedia ArticlesMarija Sakota, Maxime Peyrard, Robert WestWWW 2023 · 被引用 6 次
