Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
Li Zhou, Lutong Yu, Dongchu Xie, Shaohuan Cheng, Wenyan Li, Haizhou Li
摘要
Culture is a rich and dynamic domain that evolves across both geography and time. However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often overlooking the critical temporal dimensions. To bridge this gap, we introduce Hanfu-Bench, a novel, expert-curated multimodal dataset. Hanfu, a traditional garment spanning ancient Chinese dynasties, serves as a representative cultural heritage that reflects the profound temporal aspects of Chinese culture while remaining highly popular in Chinese contemporary society. Hanfu-Bench comprises two core tasks: cultural visual understanding and cultural image transcreation. The former task examines temporal-cultural feature recognition based on single-or multi-image inputs through multiplechoice visual question answering, while the latter focuses on transforming traditional attire into modern designs through cultural element inheritance and modern context adaptation. Our evaluation shows that closed VLMs perform comparably to non-experts on visual cutural understanding but fall short by 10% to human experts, while open VLMs lags further behind non-experts. For the transcreation task, multi-faceted human evaluation indicates that the best-performing model achieves a success rate of only 42%. Our benchmark provides an essential testbed, revealing significant challenges in this new direction of temporal cultural understanding and creative adaptation. 1 * Corresponding author † Equal contribution. 1 Following Jacovi et al. (2023), the Hanfu-Bench dataset is publicly available at lizhou21/Hanfu-Bench under the CC BY-NC-SA 4.0 License. The code details are freely available for reuse at hlt-cuhksz/TemporalCulture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- Broaden the Vision: Geo-Diverse Visual Commonsense ReasoningDa Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng 等EMNLP 2021 · 被引用 32 次
- Benchmarking Vision Language Models for Cultural UnderstandingShravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 等EMNLP 2024 · 被引用 26 次
相关 Paper
- CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationYuxuan Wang, Yijun Liu, Fei Yu, Chen Huang 等AAAI 2025 · 被引用 7 次
- Can MLLMs Understand the Deep Implication Behind Chinese Images?Chenhao Zhang, Xi Feng, Yuelin Bai, Xeron Du 等ACL 2025
- Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM CollaborationChaeHun Park, Yujin Baek, Jaeseok Kim, Yu-Jung Heo 等ACL 2025 · 被引用 17 次
- VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video ComprehensionXinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu 等ACL 2025
- STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language ModelsMahiro Ukai, Shuhei Kurita, Nakamasa InoueACM MM 2025
