Forte : Finding Outliers with Representation Typicality Estimation
Debargha Ganguly, Warren Richard Morningstar, Andrew Seohwan Yu, Vipin Chaudhary
摘要
Generative models can now produce photorealistic synthetic data which is virtually indistinguishable from the real data used to train it. This is a significant evolution over previous models which could produce reasonable facsimiles of the training data, but ones which could be visually distinguished from the training data by human evaluation. Recent work on OOD detection has raised doubts that generative model likelihoods are optimal OOD detectors due to issues involving likelihood misestimation, entropy in the generative process, and typicality. We speculate that generative OOD detectors also failed because their models focused on the pixels rather than the semantic content of the data, leading to failures in near-OOD cases where the pixels may be similar but the information content is significantly different. We hypothesize that estimating typical sets using self-supervised learners leads to better OOD detectors. We introduce a novel approach that leverages representation learning, and informative summary statistics based on manifold estimation, to address all of the aforementioned issues. Our method outperforms other unsupervised approaches and achieves state-of-the art performance on well-established challenging benchmarks, and new synthetic data detection tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning TasksDebargha Ganguly, Vikash Singh, Sreehari Sankar, Biyao Zhang 等NeurIPS 2025 · 被引用 11 次
- Trust The TypicalDebargha Ganguly, Sreehari Sankar, Biyao Zhang, Vikash Singh 等ICLR 2026 · 被引用 3 次
- CausalRAG2: Hierarchical Causal Knowledge Graph Design for RAGNengbo Wang, Tuo Liang, Vikash Singh, Chaoda Song 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
相关 Paper
- Understanding Failures in Out-of-Distribution Detection with Deep Generative ModelsLily H. Zhang, Mark Goldstein, Rajesh RanganathICML 2021 · 被引用 129 次
- Diffusion-based Semantic Outlier Generation via Nuisance Awareness for Out-of-Distribution DetectionSuhee Yoon, Sanghyu Yoon, Ye Seul Sim, Sungik Choi 等AAAI 2025 · 被引用 3 次
- SSD: A Unified Framework for Self-Supervised Outlier DetectionVikash Sehwag, Mung Chiang, Prateek MittalICLR 2021 · 被引用 410 次
- Gradient-Based Novelty Detection Boosted by Self-Supervised Binary ClassificationJingbo Sun, Li Yang, Jiaxin Zhang, Frank Liu 等AAAI 2022 · 被引用 17 次
- Detecting Generated Images by Fitting Natural Image DistributionsYonggang Zhang, Jun Nie, Xinmei Tian, Mingming Gong 等NeurIPS 2025 · 被引用 9 次
