DeepLens: Interactive Out-of-distribution Data Detection in NLP Models
Da Song, Zhijie Wang, Yuheng Huang, Lei Ma, Tianyi Zhang
摘要
Machine Learning (ML) has been widely used in Natural Language Processing (NLP) applications. A fundamental assumption in ML is that training data and real-world data should follow a similar distribution. However, a deployed ML model may suffer from out-of-distribution (OOD) issues due to distribution shifts in the real-world data. Though many algorithms have been proposed to detect OOD data from text corpora, there is still a lack of interactive tool support for ML developers. In this work, we propose DeepLens, an interactive system that helps users detect and explore OOD issues in massive text corpora. Users can efficiently explore different OOD types in DeepLens with the help of a text clustering method. Users can also dig into a specific text by inspecting salient words highlighted through neuron activation analysis. In a within-subjects user study with 24 participants, participants using DeepLens were able to find nearly twice more types of OOD issues accurately with 22% more confidence compared with a variant of DeepLens that has no interaction or visualization support.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Can multi-label classification networks know what they don't know?Haoran Wang, Weitang Liu, Alex Bocchieri, Yixuan LiNeurIPS 2021 · 被引用 168 次
- Towards neural networks that provably know when they don't knowAlexander Meinke, Matthias HeinICLR 2020 · 被引用 151 次
- Provable Guarantees for Understanding Out-of-Distribution DetectionPeyman Morteza, Yixuan LiAAAI 2022 · 被引用 104 次
相关 Paper
- Types of Out-of-Distribution Texts and How to Detect ThemUdit Arora, William Huang, He HeEMNLP 2021
- Improving Out-of-Distribution Detection with Markov Logic NetworksKonstantin Kirchheim, Frank OrtmeierICML 2025
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry 等NeurIPS 2022 · 被引用 24 次
- Cats Are Not Fish: Deep Learning Testing Calls for Out-Of-Distribution AwarenessDavid Berend, Xiaofei Xie, Lei Ma, Lingjun Zhou 等ASE 2020 · 被引用 56 次
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 被引用 9 次
