Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek Proverbs
John Pavlopoulos, Panos Louridas, Panagiotis Filos
摘要
Proverbs carry wisdom transferred orally from generation to generation. Based on the place they were recorded, this study introduces a publicly-available and machine-actionable dataset of more than one hundred thousand Greek proverb variants. By quantifying the spatial distribution of proverbs, we show that the most widespread proverbs come from the mainland while the least widespread proverbs come primarily from the islands. By focusing on the least dispersed proverbs, we present the most frequent tokens per location and undertake a benchmark in geographical attribution, using text classification and regression (text geocoding). Our results show that this is a challenging task for which specific locations can be attributed more successfully compared to others. The potential of our resource and benchmark is showcased by two novel applications. First, we extracted terms moving the regression prediction toward the four cardinal directions. Second, we leveraged conformal prediction to attribute 3,676 unregistered proverbs with statistically rigorous predictions of locations each of these proverbs was possibly registered in.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Classification with Valid and Adaptive CoverageYaniv Romano, Matteo Sesia, Emmanuel J. CandèsNeurIPS 2020 · 被引用 586 次
- Uncertainty Sets for Image Classifiers using Conformal PredictionAnastasios Nikolas Angelopoulos, Stephen Bates, Michael I. Jordan, Jitendra MalikICLR 2021 · 被引用 31 次
- ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language UnderstandingSayan Ghosh, Shashank SrivastavaACL 2022
相关 Paper
- AI Sees Your Location - But With A Bias Toward The Wealthy WorldJingyuan Huang, Jen-tse Huang, Ziyi Liu, Xiaoyuan Liu 等EMNLP 2025
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li 等EMNLP 2022 · 被引用 27 次
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等EMNLP 2022 · 被引用 7 次
- TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language ModelsJie He, Bo Peng, Yi Liao, Qun Liu 等ACL 2021
- Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test OraclesZihao Xu, Junchen Ding, Yiling Lou, Kun Zhang 等AAAI 2026 · 被引用 1 次
