Supporting Land Reuse of Former Open Pit Mining Sites using Text Classification and Active Learning
Christopher Schröder, Kim Bürgl, Yves Annanias, Andreas Niekler, Lydia Müller, Daniel Wiegreffe, Christian Bender, Christoph Mengs, Gerik Scheuermann, Gerhard Heyer
摘要
Open pit mines left many regions worldwide inhospitable or uninhabitable. Many sites are left behind in a hazardous or contaminated state, show remnants of waste, or have other restrictions imposed upon them, e.g., for the protection of human or nature. Such information has to be permanently managed in order to reuse those areas in the future. In this work we present and evaluate an automated workflow for supporting the post-mining management of former lignite open pit mines in the eastern part of Germany, where prior to any planned land reuse, aforementioned information has to be acquired to ensure the safety and validity of such an endeavor. Usually, this information is found in expert reports, either in the form of paper documents, or in the best case as digitized unstructured text-all of them in German language. However, due to the size and complexity of these documents, any inquiry is tedious and time-consuming, thereby slowing down or even obstructing the reuse of related areas. Since no training data is available, we employ active learning in order to perform multi-label sentence classification for two categories of restrictions and seven categories of topics. The final system integrates optical character recognition (OCR), active-learningbased text classification, and geographic information system visualization in order to effectively extract, query, and visualize this information for any area of interest. Active learning and text classification results are twofold: Whereas the restriction categories were reasonably accurate (>0.85 F1), the seven topicoriented categories seemed to be complex even for human annotators and achieved mediocre evaluation scores (<0.70 F1).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Natural Language Processing for Achieving Sustainable Development: the Case of Neural Labelling to Enhance Community ProfilingCostanza Conforti, Stephanie Hirmer, Dai Morgan, Marco Basaldella 等EMNLP 2020
相关 Paper
- Active Learning with Query Generation for Cost-Effective Text ClassificationYifan Yan, Sheng-Jun Huang, Shaoyi Chen, Meng Liao 等AAAI 2020 · 被引用 27 次
- CoMAL: Contrastive Active Learning for Multi-Label Text ClassificationCheng Peng, Haobo Wang, Ke Chen, Lidan Shou 等KDD 2024 · 被引用 2 次
- Learning Action Conditions from Instructional Manuals for Instruction UnderstandingTe-Lin Wu, Caiqi Zhang, Qingyuan Hu, Alexander Spangher 等ACL 2023 · 被引用 2 次
- Actively Supervised Clustering for Open Relation ExtractionJun Zhao, Yongxin Zhang, Qi Zhang, Tao Gui 等ACL 2023 · 被引用 2 次
- An Integer Linear Programming Framework for Mining Constraints from DataTao Meng, Kai-Wei ChangICML 2021 · 被引用 8 次
