Answer is All You Need: Instruction-following Text Embedding via Answering the Question
Letian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa, Gaowen Liu, Zihan Wang, Jingbo Shang
摘要
This work aims to build a text embedder that can capture characteristics of texts specified by user instructions. Despite its tremendous potential to deploy user-oriented embeddings, none of previous approaches provides a concrete solution for it. This paper offers a new viewpoint, which treats the instruction as a question about the input text and encodes the expected answers to obtain the representation accordingly. Intuitively, texts with the same (implicit) semantics would share similar answers following the instruction, thus leading to more similar embeddings. Specifically, we propose INBEDDER that instantiates this embed-via-answering idea by only fine-tuning language models on abstractive question answering tasks. INBEDDER demonstrates significantly improved instruction-following capabilities according to our proposed instruction awareness tests and instruction robustness tests, when applied to both large language models (LLMs) (e.g., llama-2-7b) and smaller encoder-based LMs (e.g., roberta-large). Additionally, our qualitative analysis of clustering outcomes, achieved by applying different instructions to the same corpus, demonstrates a high degree of interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Think Then Embed: Generative Context Improves Multimodal EmbeddingXuanming Cui, Jianpeng Cheng, Hong-You Chen, Satya Narayan Shukla 等ICLR 2026 · 被引用 41 次
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello 等NeurIPS 2024 · 被引用 26 次
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi 等ACL 2025 · 被引用 13 次
- Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive RefinementYu-Che Tsai, Kuan-Yu Chen, Yuan-Chi Li, Yuan-Hao Chen 等ICLR 2026 · 被引用 11 次
- Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingZilong Wang, Zifeng Wang, Long T. Le, Huaixiu Steven Zheng 等ICLR 2025 · 被引用 7 次
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
相关 Paper
- Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsJiuding Sun, Chantal Shaib, Byron C. WallaceICLR 2024 · 被引用 75 次
- Do LLMs "know" internally when they follow instructions?Juyeon Heo, Christina Heinze-Deml, Oussama Elachqar, Kwan Ho Ryan Chan 等ICLR 2025
- BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMsChaoyuan Shen, Chi Zhang, Chengliang Chai, Jiacheng Wang 等VLDB 2026 · 被引用 2 次
- ClusterLLM: Large Language Models as a Guide for Text ClusteringYuwei Zhang, Zihan Wang, Jingbo ShangEMNLP 2023 · 被引用 43 次
- LLaMA-Excitor: General Instruction Tuning via Indirect Feature InteractionBo Zou, Chao Yang, Yu Qiao, Chengbin Quan 等CVPR 2024 · 被引用 5 次
