A General Framework for Producing Interpretable Semantic Text Embeddings
Yiqun Sun, Qiang Huang, Yixuan Tang, Anthony Kum Hoe Tung, Jun Yu
Abstract
Semantic text embedding is essential to many tasks in Natural Language Processing (NLP). While black-box models are capable of generating high-quality embeddings, their lack of interpretability limits their use in tasks that demand transparency. Recent approaches have improved interpretability by leveraging domain-expert-crafted or LLM-generated questions, but these methods rely heavily on expert input or well-prompt design, which restricts their generalizability and ability to generate discriminative questions across a wide range of tasks. To address these challenges, we introduce CQG-MBQA (Contrastive Question Generation -Multi-task Binary Question Answering), a general framework for producing interpretable semantic text embeddings across diverse tasks. Our framework systematically generates highly discriminative, low cognitive load yes/no questions through the CQG method and answers them efficiently with the MBQA model, resulting in interpretable embeddings in a cost-effective manner. We validate the effectiveness and interpretability of CQG-MBQA through extensive experiments and ablation studies, demonstrating that it delivers embedding quality comparable to many advanced black-box models while maintaining inherently interpretability. Additionally, CQG-MBQA outperforms other interpretable text embedding methods across various downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Bayesian Concept Bottleneck Models with LLM PriorsJean Feng, Avni Kothari, Lucas Zier, Chandan Singh et al.NeurIPS 2025 · 23 citations
- Interpretable Text Embeddings and Text Similarity Explanation: A SurveyJuri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó et al.EMNLP 2025 · 3 citations
- Very Efficient Listwise Multimodal Reranking for Long DocumentsYiqun Sun, Pengfei Wei, Lawrence HsiehICML 2026 · 1 citation
Builds on12
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li et al.AAAI 2024 · 12 citations
- Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question AnsweringZhanghao Hu, Hanqi Yan, Qinglin Zhu, Zhenyi Shen et al.ACL 2025
- Let the CAT out of the bag: Contrastive Attributed explanations for TextSaneem A. Chemmengath, Amar Prakash Azad, Ronny Luss, Amit DhurandharEMNLP 2022 · 6 citations
- Concept Bottleneck Large Language ModelsChung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei WengICLR 2025
