Cross-Modal Fine-Tuning: Align then Refine
Junhong Shen, Liam Li, Lucio M. Dery, Corey Staten, Mikhail Khodak, Graham Neubig, Ameet Talwalkar
Abstract
Fine-tuning large-scale pretrained models has led to tremendous progress in well-studied modalities such as vision and NLP. However, similar gains have not been observed in many other modalities due to a lack of relevant pretrained models. In this work, we propose ORCA, a general cross-modal fine-tuning framework that extends the applicability of a single large-scale pretrained model to diverse modalities. ORCA adapts to a target task via an align-then-refine workflow: given the target input, ORCA first learns an embedding network that aligns the embedded feature distribution with the pretraining modality. The pretrained model is then fine-tuned on the embedded data to exploit the knowledge shared across modalities. Through extensive experiments, we show that ORCA obtains state-of-the-art results on 3 benchmarks containing over 60 datasets from 12 modalities, outperforming a wide range of hand-designed, Au-toML, general-purpose, and task-specific methods. We highlight the importance of data alignment via a series of ablation studies and demonstrate ORCA's utility in data-limited regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7bcd107c-e20c-4558-a43c-be71771662c1Cited by top-tier papers19
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai et al.ICML 2024 · 442 citations
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
- CALF: Aligning LLMs for Time Series Forecasting via Cross-modal Fine-TuningPeiyuan Liu, Hang Guo, Tao Dai, Naiqi Li et al.AAAI 2025 · 117 citations
- Multiple Physics Pretraining for Spatiotemporal Surrogate ModelsMichael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana et al.NeurIPS 2024 · 97 citations
- Tag-LLM: Repurposing General-Purpose LLMs for Specialized DomainsJunhong Shen, Neil A. Tenenholtz, James Brian Hall, David Alvarez-Melis et al.ICML 2024 · 60 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu et al.ICLR 2021 · 3,911 citations
Related papers
- Release the Powers of Prompt Tuning: Cross-Modality Prompt TransferNingyuan Zhang, Jie Lu, Keqiuyin Li, Zhen Fang et al.ICLR 2025
- Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality GenerationLincan Cai, Shuang Li, Wenxuan Ma, Jingxuan Kang et al.ICML 2024 · 4 citations
- Learning Modality Knowledge Alignment for Cross-Modality TransferWenxuan Ma, Shuang Li, Lincan Cai, Jingxuan KangICML 2024 · 10 citations
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
- Exploring Conditions for Diffusion Models in Robotic ControlHeeseong Shin, Byeongho Heo, Dongyoon Han, Seungryong Kim et al.CVPR 2026
