ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
Sanghyu Yoon, Dongmin Kim, Suhee Yoon, Ye Seul Sim, Seungdong YOA, Hye-Seung Cho, Soonyoung Lee, Hankook Lee, Woohyung Lim
摘要
In tabular anomaly detection (AD), textual semantic context often carries critical signals, as the definition of an anomaly is closely tied to domain-specific context. However, existing benchmarks provide only raw data points without semantic context, overlooking rich textual metadata such as feature descriptions and domain knowledge that experts rely on in practice. This limitation restricts research flexibility and prevents models from fully leveraging domain knowledge for detection. ReTabAD addresses this gap by Restoring textual semantics to enable contextaware Tabular Anomaly Detection research. We provide (1) 20 carefully curated tabular datasets enriched with structured textual metadata, together with implementations of state-of-the-art AD algorithms-including classical, deep learning, and LLM-based approaches-and (2) a zero-shot LLM framework that leverages semantic context without task-specific training, establishing a strong baseline for future research. Furthermore, this work provides insights into the role and utility of textual metadata in AD through experiments and analysis. Results show that semantic context improves detection performance and enhances interpretability by supporting domain-aware reasoning. These findings establish ReTabAD as a benchmark for systematic exploration of context-aware AD. The resource is publicly released at https://yoonsanghyu.github.io/ReTabAD/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Classification-Based Anomaly Detection for General DataLiron Bergman, Yedid HoshenICLR 2020 · 被引用 412 次
- A Closer Look at AUROC and AUPRC under Class ImbalanceMatthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti 等NeurIPS 2024 · 被引用 191 次
- Neural Transformation Learning for Deep Anomaly Detection Beyond ImagesChen Qiu, Timo Pfrommer, Marius Kloft, Stephan Mandt 等ICML 2021 · 被引用 171 次
- Anomaly Detection for Tabular Data with Internal Contrastive LearningTom Shenkar, Lior WolfICLR 2022 · 被引用 127 次
相关 Paper
- AnoLLM: Large Language Models for Tabular Anomaly DetectionChe-Ping Tsai, Ganyu Teng, Phillip Wallis, Wei DingICLR 2025
- LLM as an Algorithmist: Enhancing Anomaly Detectors via Programmatic SynthesisHangting Ye, Jinmeng Li, He Zhao, Mingchen Zhuge 等ICLR 2026 · 被引用 2 次
- TAB: Unified Benchmarking of Time Series Anomaly Detection MethodsXiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu 等VLDB 2025 · 被引用 57 次
- ZeroED: Hybrid Zero-Shot Error Detection Through Large Language Model ReasoningWei Ni, Kaihang Zhang, Xiaoye Miao, Xiangyu Zhao 等ICDE 2025 · 被引用 5 次
- ICAD-LLM: One-for-All Anomaly Detection via In-Context Learning with Large Language ModelsZhongyuan Wu, Jingyuan Wang, Zexuan Cheng, Yilong Zhou 等AAAI 2026 · 被引用 1 次
