PhiloGPT: A Philology-Oriented Large Language Model for Ancient Chinese Manuscripts with Dunhuang as Case Study
Yuqing Zhang, Baoyi He, Yihan Chen, Hangqi Li, Han Yue, Shengyu Zhang, Huaiyong Dou, Junchi Yan, Zemin Liu, Yongquan Zhang, Fei Wu
Abstract
Philology, the study of ancient manuscripts, demands years of professional training in extensive knowledge memorization and manual textual retrieval. Despite these requirements align closely with strengths of recent successful Large Language Models (LLMs), the scarcity of high-quality, specialized training data has hindered direct applications. To bridge this gap, we curated the PhiloCorpus-ZH, a rich collection of ancient Chinese texts spanning a millennium with 30 diverse topics, including firsthand folk copies. This corpus facilitated the development of PhiloGPT, the first LLM tailored for discovering ancient Chinese manuscripts. To effectively tackle complex philological tasks like restoration, attribution, and linguistic analysis, we introduced the PhiloCoP framework. Modeled on the analytical patterns of philologists, PhiloCoP enhances LLM's handling of historical linguistic peculiarities such as phonetic loans, polysemy, and syntactic inversions. We further integrated these tasks into the PhiloBenchmark, establishing a new standard for evaluating ancient Chinese LLMs addressing philology tasks. Deploying PhiloGPT in practical scenarios has enabled Dunhuang specialists to resolve philology tasks, such as identifying duplication of copied text and assisting archaeologists with text completion, demonstrating its potential in real-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfef0f32-4118-46cd-85a2-5cf28c8fe8afCited by top-tier papers1
Ask how each one uses itBuilds on6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang et al.ACM MM 2020 · 66 citations
- OceanGPT: A Large Language Model for Ocean Science TasksZhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou et al.ACL 2024 · 35 citations
Related papers
- AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese CorporaZhihan Zhou, Daqian Shi, Rui Song, Lida Shi et al.AAAI 2026 · 1 citation
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li et al.ACL 2026 · 1 citation
- MCS-Bench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in Chinese Classical StudiesYang Liu, Jiahuan Cao, Hiuyi Cheng, Yongxin Shi et al.ACL 2025
- MCHDoc: A Comprehensive Benchmark for Reading Multi-Carrier Chinese Historical DocumentsYijun Sheng, Shipeng Zhu, Ruijia Zuo, Na Nie et al.CVPR 2026
- LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLPDanlu Chen, Freda Shi, Aditi Agarwal, Jacobo Myerston et al.ACL 2024
