Learning to detect table clones in spreadsheets
Yakun Zhang, Wensheng Dou, Jiaxin Zhu, Liang Xu, Zhiyong Zhou, Jun Wei, Dan Ye, Bo Yang
摘要
In order to speed up spreadsheet development productivity, end users can create a spreadsheet table by copying and modifying an existing one. These two tables share the similar computational semantics, and form a table clone. End users may modify the tables in a table clone, e.g., adding new rows and deleting columns, thus introducing structure changes into the table clone. Our empirical study on real-world spreadsheets shows that about 58.5% of table clones involve structure changes. However, existing table clone detection approaches in spreadsheets can only detect table clones with the same structures. Therefore, many table clones with structure changes cannot be detected.
We observe that, although the tables in a table clone may be modified, they usually share the similar structures and formats, e.g., headers, formulas and background colors. Based on this observation, we propose LTC (Learning to detect T able Clones), to automatically detect table clones with or without structure changes. LTC utilizes the structure and format information from labeled table clones and non table clones to train a binary classifier. LTC first identifies tables in spreadsheets, and then uses the trained binary classifier to judge whether every two tables can form a table clone. Our experiments on real-world spreadsheets from the EUSES and Enron corpora show that, LTC can achieve a precision of 97.8% and recall
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Semantic table structure identification in spreadsheetsYakun Zhang, Xiao Lv, Haoyu Dong, Wensheng Dou 等ISSTA 2021 · 被引用 11 次
- One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)Xu Yang, Shaowei Wang, Jiayuan Zhou, Wenhan ZhuFSE 2025 · 被引用 7 次
相关 Paper
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton 等ICML 2021 · 被引用 63 次
- NIL: large-scale detection of large-variance clonesTasuku Nakagawa, Yoshiki Higo, Shinji KusumotoFSE 2021 · 被引用 41 次
- Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table RepresentationsSibei Chen, Yeye He, Weiwei Cui, Ju Fan 等SIGMOD 2024 · 被引用 4 次
- SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention NetworkRan Jia, Qiyu Li, Zihan Xu, Xiaoyuan Jin 等AAAI 2023 · 被引用 3 次
- CC2Vec: Combining Typed Tokens with Contrastive Learning for Effective Code Clone DetectionShihan Dou, Yueming Wu, Haoxiang Jia, Yuhao Zhou 等FSE 2024 · 被引用 9 次
