ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text Classification
Bowen Wei, Ziwei Zhu
Abstract
Deep neural networks have achieved remarkable performance in various text-based tasks but often lack interpretability, making them less suitable for applications where transparency is critical. To address this, we propose ProtoLens, a novel prototype-based model that provides fine-grained, sub-sentence level interpretability for text classification. ProtoLens uses a Prototype-aware Span Extraction module to identify relevant text spans associated with learned prototypes and a Prototype Alignment mechanism to ensure prototypes are semantically meaningful throughout training. By aligning the prototype embeddings with human-understandable examples, ProtoLens provides interpretable predictions while maintaining competitive accuracy. Extensive experiments demonstrate that ProtoLens outperforms both prototype-based and non-interpretable baselines on multiple text classification benchmarks. Code and data are available at https://anonymous.4open.science/r/ProtoLens-CE0B/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d6c4cb5-78f9-4fb3-a6f0-272202492d2fCited by top-tier papers2
- Prototype Transformer: Towards Language Model Architectures Interpretable by DesignYordan Yordanov, Matteo Forasassi, Bayar Menzat, Ruizhi Wang et al.ICML 2026 · 1 citation
- SCOUT: Selective Coupling via Optimal Unbalanced Transport for Interpretable Text ClassificationJunhao Jia, Hanwen Zheng, Yueyi Wu, Huangwei Chen et al.ACL 2026 · 1 citation
Builds on8
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Visualizing Deep Networks by Optimizing with Integrated GradientsZhongang Qi, Saeed Khorram, Fuxin LiAAAI 2020 · 149 citations
- ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelSrishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin et al.NeurIPS 2022 · 55 citations
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
- Interpretable Image Classification with Adaptive Prototype-based Vision TransformersChiyu Ma, Jon Donnelly, Wenjun Liu, Soroush Vosoughi et al.NeurIPS 2024 · 48 citations
Related papers
- A Multi-Grained Self-Interpretable Symbolic-Neural Model For Single/Multi-Labeled Text ClassificationXiang Hu, Xinyu Kong, Kewei TuICLR 2023 · 2 citations
- ProtoTEx: Explaining Model Decisions with Prototype TensorsAnubrata Das, Chitrank Gupta, Venelin Kovatchev, Matthew Lease et al.ACL 2022
- Making Sense of LLM Decisions: A Prototype-based Framework for Explainable ClassificationBowen Wei, Mehrdad Fazli, Ziwei ZhuAAAI 2026
- ProtGNN: Towards Self-Explaining Graph Neural NetworksZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu et al.AAAI 2022 · 173 citations
- Prototype-Grounded Concept Models for Verifiable Concept AlignmentStefano Colamonaco, David Debot, Pietro Barbiero, Giuseppe MarraICML 2026
