LLM as an Algorithmist: Enhancing Anomaly Detectors via Programmatic Synthesis
Hangting Ye, Jinmeng Li, He Zhao, Mingchen Zhuge, Dandan Guo, Yi Chang, Hongyuan Zha
Abstract
Existing anomaly detection (AD) methods for tabular data usually rely on some assumptions about anomaly patterns, leading to inconsistent performance in realworld scenarios. While Large Language Models (LLMs) show remarkable reasoning capabilities, their direct application to tabular AD is impeded by fundamental challenges, including difficulties in processing heterogeneous data and significant privacy risks. To address these limitations, we propose LLM-DAS, a novel framework that repositions the LLM from a "data processor" to an "algorithmist". Instead of being exposed to raw data, our framework leverages the LLM's ability to reason about algorithms. It analyzes a high-level description of a given detector to understand its intrinsic weaknesses and then generates detector-specific, dataagnostic Python code to synthesize "hard-to-detect" anomalies that exploit these vulnerabilities. This generated synthesis program, which is reusable across diverse datasets, is then instantiated to augment training data, systematically enhancing the detector's robustness by transforming the problem into a more discriminative two-class classification task. Extensive experiments on 36 TAD benchmarks show that LLM-DAS consistently boosts the performance of mainstream detectors. By bridging LLM reasoning with classic AD algorithms via programmatic synthesis, LLM-DAS offers a scalable, effective, and privacy-preserving approach to patching the logical blind spots of existing detectors. The source code is available at https://github.com/HangtingYe/LLM_DAS# .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c44c1e00-58b1-4ec9-a846-69230b0c22ffCited by top-tier papers1
Ask how each one uses itBuilds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Classification-Based Anomaly Detection for General DataLiron Bergman, Yedid HoshenICLR 2020 · 412 citations
- Learning and Evaluating Representations for Deep One-Class ClassificationKihyuk Sohn, Chun-Liang Li, Jinsung Yoon, Minho Jin et al.ICLR 2021 · 243 citations
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
Related papers
- AnoLLM: Large Language Models for Tabular Anomaly DetectionChe-Ping Tsai, Ganyu Teng, Phillip Wallis, Wei DingICLR 2025
- MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language ModelsXincheng Yao, Zefeng Qian, Chao Shi, Jiayang Song et al.CVPR 2026 · 2 citations
- ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly DetectionSanghyu Yoon, Dongmin Kim, Suhee Yoon, Ye Seul Sim et al.ICLR 2026 · 3 citations
- RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?Di Cao, Yong Liao, Xiuwei ShangEMNLP 2024 · 11 citations
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan et al.KDD 2026
