PhiloGPT: A Philology-Oriented Large Language Model for Ancient Chinese Manuscripts with Dunhuang as Case Study
Yuqing Zhang, Baoyi He, Yihan Chen, Hangqi Li, Han Yue, Shengyu Zhang, Huaiyong Dou, Junchi Yan, Zemin Liu, Yongquan Zhang, Fei Wu
摘要
Philology, the study of ancient manuscripts, demands years of professional training in extensive knowledge memorization and manual textual retrieval. Despite these requirements align closely with strengths of recent successful Large Language Models (LLMs), the scarcity of high-quality, specialized training data has hindered direct applications. To bridge this gap, we curated the PhiloCorpus-ZH, a rich collection of ancient Chinese texts spanning a millennium with 30 diverse topics, including firsthand folk copies. This corpus facilitated the development of PhiloGPT, the first LLM tailored for discovering ancient Chinese manuscripts. To effectively tackle complex philological tasks like restoration, attribution, and linguistic analysis, we introduced the PhiloCoP framework. Modeled on the analytical patterns of philologists, PhiloCoP enhances LLM's handling of historical linguistic peculiarities such as phonetic loans, polysemy, and syntactic inversions. We further integrated these tasks into the PhiloBenchmark, establishing a new standard for evaluating ancient Chinese LLMs addressing philology tasks. Deploying PhiloGPT in practical scenarios has enabled Dunhuang specialists to resolve philology tasks, such as identifying duplication of copied text and assisting archaeologists with text completion, demonstrating its potential in real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang 等ACM MM 2020 · 被引用 66 次
- OceanGPT: A Large Language Model for Ocean Science TasksZhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou 等ACL 2024 · 被引用 35 次
相关 Paper
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese CorporaZhihan Zhou, Daqian Shi, Rui Song, Lida Shi 等AAAI 2026 · 被引用 1 次
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li 等ACL 2026 · 被引用 1 次
- MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical StudiesYang Liu, Jiahuan Cao, Hiuyi Cheng, Yongxin Shi 等ACL 2025
- MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical DocumentsYijun Sheng, Shipeng Zhu, Ruijia Zuo, Na Nie 等CVPR 2026
- LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLPDanlu Chen, Freda Shi, Aditi Agarwal, Jacobo Myerston 等ACL 2024
