Forte : Finding Outliers with Representation Typicality Estimation
Debargha Ganguly, Warren Richard Morningstar, Andrew Seohwan Yu, Vipin Chaudhary
Abstract
Generative models can now produce photorealistic synthetic data which is virtually indistinguishable from the real data used to train it. This is a significant evolution over previous models which could produce reasonable facsimiles of the training data, but ones which could be visually distinguished from the training data by human evaluation. Recent work on OOD detection has raised doubts that generative model likelihoods are optimal OOD detectors due to issues involving likelihood misestimation, entropy in the generative process, and typicality. We speculate that generative OOD detectors also failed because their models focused on the pixels rather than the semantic content of the data, leading to failures in near-OOD cases where the pixels may be similar but the information content is significantly different. We hypothesize that estimating typical sets using self-supervised learners leads to better OOD detectors. We introduce a novel approach that leverages representation learning, and informative summary statistics based on manifold estimation, to address all of the aforementioned issues. Our method outperforms other unsupervised approaches and achieves state-of-the art performance on well-established challenging benchmarks, and new synthetic data detection tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67f7469b-6f30-419f-a9da-b865fcf573e5Cited by top-tier papers3
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning TasksDebargha Ganguly, Vikash Singh, Sreehari Sankar, Biyao Zhang et al.NeurIPS 2025 · 11 citations
- Trust The TypicalDebargha Ganguly, Sreehari Sankar, Biyao Zhang, Vikash Singh et al.ICLR 2026 · 3 citations
- CausalRAG2: Hierarchical Causal Knowledge Graph Design for RAGNengbo Wang, Tuo Liang, Vikash Singh, Chaoda Song et al.ICML 2026 · 2 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- Understanding Failures in Out-of-Distribution Detection with Deep Generative ModelsLily H. Zhang, Mark Goldstein, Rajesh RanganathICML 2021 · 129 citations
- Diffusion-based Semantic Outlier Generation via Nuisance Awareness for Out-of-Distribution DetectionSuhee Yoon, Sanghyu Yoon, Ye Seul Sim, Sungik Choi et al.AAAI 2025 · 3 citations
- SSD: A Unified Framework for Self-Supervised Outlier DetectionVikash Sehwag, Mung Chiang, Prateek MittalICLR 2021 · 410 citations
- Gradient-Based Novelty Detection Boosted by Self-Supervised Binary ClassificationJingbo Sun, Li Yang, Jiaxin Zhang, Frank Liu et al.AAAI 2022 · 17 citations
- Detecting Generated Images by Fitting Natural Image DistributionsYonggang Zhang, Jun Nie, Xinmei Tian, Mingming Gong et al.NeurIPS 2025 · 9 citations
