Disambiguating Symbolic Expressions in Informal Documents
Dennis Müller, Cezary Kaliszyk
摘要
We propose the task of disambiguating symbolic expressions in informal STEM documents in the form of L A T E X files -that is, determining their precise semantics and abstract syntax tree -as a neural machine translation task. We discuss the distinct challenges involved and present a dataset with roughly 33,000 entries. We evaluated several baseline models on this dataset, which failed to yield even syntactically valid L A T E X before overfitting. Consequently, we describe a methodology using a transformer language model pre-trained on sources obtained from arxiv.org, which yields promising results despite the small size of the dataset. We evaluate our model using a plurality of dedicated techniques, taking the syntax and semantics of symbolic expressions into account.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Neural Machine Translation for Mathematical FormulaeFelix Petersen, Moritz Schubotz, André Greiner-Petter, Bela GippACL 2023 · 被引用 6 次
- Syntax-Aware Network for Handwritten Mathematical Expression RecognitionYe Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu 等CVPR 2022 · 被引用 74 次
- Generating Data for Symbolic Language with Large Language ModelsJiacheng Ye, Chengzu Li, Lingpeng Kong, Tao YuEMNLP 2023 · 被引用 8 次
- Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong BaselineWeikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie 等AAAI 2026 · 被引用 2 次
- KERMIT: Complementing Transformer Architectures with Encoders of Explicit Syntactic InterpretationsFabio Massimo Zanzotto, Andrea Santilli, Leonardo Ranaldi, Dario Onorati 等EMNLP 2020 · 被引用 45 次
