Beyond Single Representations: Multi-Model Embedding Fusion for Stable Text Classification
Jiho Gwak, Yuchul Jung
摘要
Embedding fusion has become a widely adopted technique for enhancing performance across various NLP tasks. While prior research suggests that different layers of language models encode distinct representations and that pooling strategies influence performance, there is a lack of systematic analysis regarding the empirical efficacy of these differences or the impact of combining embeddings from multiple models. This study provides a rigorous, empirical evaluation of layer-wise fusion strategies to determine their actual contribution to classification performance. Our findings reveal that the effectiveness of individual layers is more dependent on dataset characteristics than on the model architecture itself. Furthermore, we demonstrate that fusing embeddings from multiple models yields more robust and consistent representations across tasks, with the influence of any single model diminishing as the number of integrated models increases. Notably, experiments on low-resource datasets show that embedding fusion provides particularly significant gains when training data is scarce, highlighting its robustness and adaptability in data-constrained environments. We also analyze the trade-off between performance gains and computational overhead, and discuss which fusion configurations provide the best balance between stability and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- On the Worst Prompt Performance of Large Language ModelsBowen Cao, Deng Cai, Zhisong Zhang, Yuexian Zou 等NeurIPS 2024 · 被引用 65 次
- Prompt-Based Meta-Learning For Few-shot Text ClassificationHaoxing Zhang, Xiaofeng Zhang, Haibo Huang, Lei YuEMNLP 2022 · 被引用 34 次
相关 Paper
- Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERTRik van Noord, Antonio Toral, Johan BosEMNLP 2020 · 被引用 22 次
- Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence LearningXuebo Liu, Longyue Wang, Derek F. Wong, Liang Ding 等ICLR 2021 · 被引用 11 次
- Analysis of Multi-Source Language Training in Cross-Lingual TransferSeong Hoon Lim, Taejun Yun, Jinhyeon Kim, Jihun Choi 等ACL 2024 · 被引用 1 次
- FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input RepresentationsLukas Lange, Heike Adel, Jannik Strötgen, Dietrich KlakowEMNLP 2021 · 被引用 6 次
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang 等NeurIPS 2024 · 被引用 94 次
