TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection
Siddhant Garg, Thuy Vu, Alessandro Moschitti
Abstract
We propose TANDA, an effective technique for fine-tuning pre-trained Transformer models for natural language tasks. Specifically, we first transfer a pre-trained model into a model for a general task by fine-tuning it with a large and highquality dataset. We then perform a second fine-tuning step to adapt the transferred model to the target domain. We demonstrate the benefits of our approach for answer sentence selection, which is a well-known inference task in Question Answering. We built a large scale dataset to enable the transfer step, exploiting the Natural Questions dataset. Our approach establishes the state of the art on two well-known benchmarks, WikiQA and TREC-QA, achieving MAP scores of 92% and 94.3%, respectively, which largely outperform the previous highest scores of 83.4% and 87.5%, obtained in very recent work. We empirically show that TANDA generates more stable and robust models reducing the effort required for selecting optimal hyper-parameters. Additionally, we show that the transfer step of TANDA makes the adaptation step more robust to noise. This enables a more effective use of noisy datasets for fine-tuning. Finally, we also confirm the positive impact of TANDA in an industrial setting, using domain specific datasets subject to different types of noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 783f7490-6adf-4343-947a-c56d9e8d7bbbCited by top-tier papers19
- Open-Retrieval Conversational Question AnsweringChen Qu, Liu Yang, Cen Chen, Minghui Qiu et al.SIGIR 2020 · 84 citations
- CABINET: Content Relevance-based Noise Reduction for Table Question AnsweringSohan Patnaik, Heril Changwal, Milan Aggarwal, Sumit Bhatia et al.ICLR 2024 · 34 citations
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal ReasoningRujun Han, Xiang Ren, Nanyun PengEMNLP 2021 · 31 citations
- Answer Summarization for Technical Queries: Benchmark and New ApproachChengran Yang, Bowen Xu, Ferdian Thung, Yucen Shi et al.ASE 2022 · 12 citations
- Efficient Re-ranking with Cross-encoders via Early ExitFrancesco Busolin, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando et al.SIGIR 2025 · 9 citations
Related papers
- Knowledge Transfer from Answer Ranking to Answer GenerationMatteo Gabburo, Rik Koncel-Kedziorski, Siddhant Garg, Luca Soldaini et al.EMNLP 2022 · 4 citations
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut et al.ACL 2020 · 168 citations
- Topic Transferable Table Question AnsweringSaneem A. Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen et al.EMNLP 2021
- Recursive Tree-Structured Self-Attention for Answer Sentence SelectionKhalil Mrini, Emilia Farcas, Ndapa NakasholeACL 2021
- FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text modelsRakesh Chada, Pradeep NatarajanEMNLP 2021 · 36 citations
