Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
Mahdi Karami, Ali Ghodsi
摘要
In the rapidly evolving field of deep learning, the demand for models that are both expressive and computationally efficient has never been more critical. This paper introduces Orchid, a novel architecture designed to address the quadratic complexity of traditional attention mechanisms without compromising the ability to capture long-range dependencies and in-context learning. At the core of this architecture lies a new data-dependent global convolution layer, which contextually adapts its kernel conditioned on input sequence using a dedicated conditioning neural network. We design two simple conditioning networks that maintain shift equivariance in our data-dependent convolution operation. The dynamic nature of the proposed convolution kernel grants Orchid high expressivity while maintaining quasilinear scalability for long sequences. We evaluate the proposed model across multiple domains, including language modeling and image classification, to highlight its performance and generality. Our experiments demonstrate that this architecture not only outperforms traditional attention-based architectures such as BERT and Vision Transformers with smaller model sizes, but also extends the feasible sequence length beyond the limitations of the dense attention layers. This achievement represents a significant step towards more efficient and scalable deep learning models for sequence modeling. The code is available at https://github.com/Karami-m/orchid.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate ScenariosTongzhou Yu, Mingjia Li, Hong Qian, Wenkai Wang 等KDD 2026 · 被引用 1 次
- Best of Both Worlds: Advantages of Hybrid Graph Sequence ModelsAli Behrouz, Ali Parviz, Mahdi Karami, Clayton Sanford 等ICML 2025
- Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and BeyondCostin-Andrei Oncescu, Sanket Purandare, Stratos Idreos, Sham M. KakadeICLR 2025
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
相关 Paper
- What Makes Convolutional Models Great on Long Sequence Modeling?Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen 等ICLR 2023 · 被引用 20 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
- Efficient Representation Learning via Adaptive Context PoolingChen Huang, Walter Talbott, Navdeep Jaitly, Joshua M. SusskindICML 2022 · 被引用 10 次
- Core Context Aware Transformers for Long Context Language ModelingYaofo Chen, Zeng You, Shuhai Zhang, Haokun Li 等ICML 2025
- Monarch Mixer: A Simple Sub-Quadratic GEMM-Based ArchitectureDaniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson 等NeurIPS 2023 · 被引用 80 次
