TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domains
Yidan Sun, Mengying Zhu, Feiyue Chen, Yangyang Wu, Xiaolei Dan, Mengyuan Yang, Xiaolin Zheng, Shenglin Ben
摘要
Large language models (LLMs) have demonstrated impressive performance in text generation tasks; however, their embedding spaces often suffer from the isotropy problem, resulting in poor discrimination of domain-specific terminology, particularly in legal and financial contexts. This weakness in terminology-level representation can severely hinder downstream tasks such as legal judgment prediction or financial risk analysis, where subtle semantic distinctions are critical. To address this problem, we propose TermGPT, a multi-level contrastive fine-tuning framework designed for terminology adaptation. We first construct a sentence graph to capture semantic and structural relations, and generate semantically consistent yet discriminative positive and negative samples based on contextual and topological cues. We then devise a multi-level contrastive learning approach at both the sentence and token levels, enhancing global contextual understanding and fine-grained terminology discrimination. To support robust evaluation, we construct the first financial terminology dataset derived from official regulatory documents. Experiments show that TermGPT outperforms existing baselines in term discrimination tasks within the finance and legal domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama 等NeurIPS 2022 · 被引用 349 次
- Contrastive Learning with Adversarial Perturbations for Conditional Text GenerationSeanie Lee, Dong Bok Lee, Sung Ju HwangICLR 2021 · 被引用 117 次
- Frequency-Aware Contrastive Learning for Neural Machine TranslationTong Zhang, Wei Ye, Baosong Yang, Long Zhang 等AAAI 2022 · 被引用 35 次
- Following the Autoregressive Nature of LLM Embeddings via Compression and AlignmentJingcheng Deng, Zhongtao Jiang, Liang Pang, Zihao Wei 等EMNLP 2025
相关 Paper
- CoEvo: Coevolution of LLM and Retrieval Model for Domain-Specific Information RetrievalAng Li, Yiquan Wu, Yinghao Hu, Lizhi Qing 等EMNLP 2025
- CLER: Improving Multimodal Financial Reasoning by Cross-MLLM Error ReflectionShuangyan Deng, Zhongsheng Wang, Rui Mao, Ciprian Doru Giurcaneanu 等AAAI 2026
- UTAG: Leveraging LLM as a Unified Embedding Generator for Text-Attributed GraphsMingqian Ding, Jianjun Li, Zhiyuan Ma, Liwei Zhang 等WWW 2026
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined DataZhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen 等NeurIPS 2025 · 被引用 14 次
