OldVisOnline: Curating a Dataset of Historical Visualizations
Yu Zhang, Ruike Jiang, Liwenhan Xie, Yuheng Zhao, Can Liu, Tianhong Ding, Siming Chen, Xiaoru Yuan
摘要
With the increasing adoption of digitization, more and more historical visualizations created hundreds of years ago are accessible in digital libraries online. It provides a unique opportunity for visualization and history research. Meanwhile, there is no large-scale digital collection dedicated to historical visualizations. The visualizations are scattered in various collections, which hinders retrieval. In this study, we curate the first large-scale dataset dedicated to historical visualizations. Our dataset comprises 13K historical visualization images with corresponding processed metadata from seven digital libraries. In curating the dataset, we propose a workflow to scrape and process heterogeneous metadata. We develop a semi-automatic labeling approach to distinguish visualizations from other artifacts. Our dataset can be accessed with OldVisOnline, a system we have built to browse and label historical visualizations. We discuss our vision of usage scenarios and research opportunities with our dataset, such as textual criticism for historical visualizations. Drawing upon our experience, we summarize recommendations for future efforts to improve our dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DataWink: Reusing and Adapting SVG-Based Visualization Examples with Large Multimodal ModelsLiwenhan Xie, Yanna Lin, Can Liu, Huamin Qu 等IEEE VIS 2025 · 被引用 3 次
- ZuantuSet: A Collection of Historical Chinese Visualizations and IllustrationsXiyao Mei, Yu Zhang, Chaofan Yang, Rui Shi 等CHI 2025 · 被引用 1 次
- Reframing Pattern: A Comprehensive Approach to a Composite Visual VariableTingying He, Jason Dykes, Petra Isenberg, Tobias IsenbergIEEE VIS 2025
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Composition and Configuration Patterns in Multiple-View VisualizationsXi Chen, Wei Zeng, Yanna Lin, Hayder Mahdi Al-Maneea 等IEEE VIS 2020 · 被引用 138 次
- Exploring Visual Information Flows in InfographicsMin Lu, Chufeng Wang, Joel Lanir, Nanxuan Zhao 等CHI 2020 · 被引用 65 次
- What Makes a Data-GIF Understandable?Xinhuan Shu, Aoyu Wu, Junxiu Tang, Benjamin Bach 等IEEE VIS 2020 · 被引用 64 次
相关 Paper
- ComputableViz: Mathematical Operators as a Formalism for Visualisation Processing and AnalysisAoyu Wu, Wai Tong, Haotian Li, Dominik Moritz 等CHI 2022 · 被引用 8 次
- "I Came Across a Junk": Understanding Design Flaws of Data Visualization from the Public's PerspectiveXingyu Lan, Yu LiuIEEE VIS 2024 · 被引用 12 次
- Input Visualization: Collecting and Modifying Data with Visual RepresentationsNathalie Bressa, Jordan Louis, Wesley Willett, Samuel HuronCHI 2024 · 被引用 25 次
- Annotating Line Charts for Addressing DeceptionArlen Fan, Yuxin Ma, Michelle Mancenido, Ross MaciejewskiCHI 2022 · 被引用 44 次
- Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information RetrievalZheng Liu, Ze Liu, Zhengyang Liang, Junjie Zhou 等ACL 2025 · 被引用 9 次
