Detecting Layout Templates in Complex Multiregion Files
Gerardo Vitagliano, Lan Jiang, Felix Naumann
Abstract
Spreadsheets are among the most commonly used file formats for data management, distribution, and analysis. Their widespread employment makes it easy to gather large collections of data, but their flexible canvas-based structure makes automated analysis difficult without heavy preparation. One of the common problems that practitioners face is the presence of multiple, independent regions in a single spreadsheet, possibly separated by repeated empty cells. We define such files as "multiregion" files. In collections of various spreadsheets, we can observe that some share the same layout. We present the Mondrian approach to automatically identify layout templates across multiple files and systematically extract the corresponding regions. Our approach is composed of three phases: first, each file is rendered as an image and inspected for elements that could form regions; then, using a clustering algorithm, the identified elements are grouped to form regions; finally, every file layout is represented as a graph and compared with others to find layout templates. We compare our method to state-of-the-art table recognition algorithms on two corpora of real-world enterprise spreadsheets. Our approach shows the best performances in detecting reliable region boundaries within each file and can correctly identify recurring layouts across files.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener et al.VLDB 2023 · 9 citations
- Efficient and Compact Spreadsheet Formula GraphsDixin Tang, Fanchao Chen, Christopher De Leon, Tana Wattanawaroon et al.ICDE 2023
Builds on3
- Finding Related Tables in Data Lakes for Interactive Data ScienceYi Zhang, Zachary G. IvesSIGMOD 2020 · 98 citations
- Valentine: Evaluating Matching Techniques for Dataset DiscoveryChristos Koutras, George Siachamis, Andra Ionescu, Kyriakos Psarakis et al.ICDE 2021 · 87 citations
- Pytheas: Pattern-based Table Discovery in CSV FilesChristina Christodoulakis, Eric B. Munson, Moshe Gabel, Angela Demke Brown et al.VLDB 2020
Related papers
- Semantic table structure identification in spreadsheetsYakun Zhang, Xiao Lv, Haoyu Dong, Wensheng Dou et al.ISSTA 2021 · 11 citations
- SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention NetworkRan Jia, Qiyu Li, Zihan Xu, Xiaoyuan Jin et al.AAAI 2023 · 3 citations
- SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based ReflectionQin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu et al.EMNLP 2025
- Learning to detect table clones in spreadsheetsYakun Zhang, Wensheng Dou, Jiaxin Zhu, Liang Xu et al.ISSTA 2020 · 8 citations
- Shape-Agnostic Table Overlap Discovery: A Maximum Common Subhypergraph ApproachGe Lee, Shixun Huang, Zhifeng Bao, Felix Naumann et al.SIGMOD 2026 · 1 citation
