Detecting Layout Templates in Complex Multiregion Files
Gerardo Vitagliano, Lan Jiang, Felix Naumann
摘要
Spreadsheets are among the most commonly used file formats for data management, distribution, and analysis. Their widespread employment makes it easy to gather large collections of data, but their flexible canvas-based structure makes automated analysis difficult without heavy preparation. One of the common problems that practitioners face is the presence of multiple, independent regions in a single spreadsheet, possibly separated by repeated empty cells. We define such files as "multiregion" files. In collections of various spreadsheets, we can observe that some share the same layout. We present the Mondrian approach to automatically identify layout templates across multiple files and systematically extract the corresponding regions. Our approach is composed of three phases: first, each file is rendered as an image and inspected for elements that could form regions; then, using a clustering algorithm, the identified elements are grouped to form regions; finally, every file layout is represented as a graph and compared with others to find layout templates. We compare our method to state-of-the-art table recognition algorithms on two corpora of real-world enterprise spreadsheets. Our approach shows the best performances in detecting reliable region boundaries within each file and can correctly identify recurring layouts across files.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener 等VLDB 2023 · 被引用 9 次
- Efficient and Compact Spreadsheet Formula GraphsDixin Tang, Fanchao Chen, Christopher De Leon, Tana Wattanawaroon 等ICDE 2023
它引用的顶会 Paper3
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 被引用 98 次
- Valentine: Evaluating Matching Techniques for Dataset DiscoveryChristos Koutras, George Siachamis, Andra Ionescu, Kyriakos Psarakis 等ICDE 2021 · 被引用 87 次
- Pytheas: Pattern-based Table Discovery in CSV FilesChristina Christodoulakis, Eric B. Munson, Moshe Gabel, Angela Demke Brown 等VLDB 2020
相关 Paper
- Semantic table structure identification in spreadsheetsYakun Zhang, Xiao Lv, Haoyu Dong, Wensheng Dou 等ISSTA 2021 · 被引用 11 次
- SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention NetworkRan Jia, Qiyu Li, Zihan Xu, Xiaoyuan Jin 等AAAI 2023 · 被引用 3 次
- SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based ReflectionQin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu 等EMNLP 2025
- Learning to detect table clones in spreadsheetsYakun Zhang, Wensheng Dou, Jiaxin Zhu, Liang Xu 等ISSTA 2020 · 被引用 8 次
- Shape-Agnostic Table Overlap Discovery: A Maximum Common Subhypergraph ApproachGe Lee, Shixun Huang, Zhifeng Bao, Felix Naumann 等SIGMOD 2026 · 被引用 1 次
