Conquering the CNN Over-Parameterization Dilemma: A Volterra Filtering Approach for Action Recognition
Siddharth Roheda, Hamid Krim
Abstract
The importance of inference in Machine Learning (ML) has led to an explosive number of different proposals, particularly in Deep Learning. In an attempt to reduce the complexity of Convolutional Neural Networks, we propose a Volterra filter-inspired Network architecture. This architecture introduces controlled non-linearities in the form of interactions between the delayed input samples of data. We propose a cascaded implementation of Volterra Filtering so as to significantly reduce the number of parameters required to carry out the same classification task as that of a conventional Neural Network. We demonstrate an efficient parallel implementation of this Volterra Neural Network (VNN), along with its remarkable performance while retaining a relatively simpler and potentially more tractable structure. Furthermore, we show a rather sophisticated adaptation of this network to nonlinearly fuse the RGB (spatial) information and the Optical Flow (temporal) information of a video sequence for action recognition. The proposed approach is evaluated on UCF-101 and HMDB-51 datasets for action recognition, and is shown to outperform state of the art CNN approaches. The code-base for our paper is available on request (Fill THIS form).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e824db23-dbad-4f81-8332-c4a89a952500Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Convolution Goes Higher-Order: A Biologically Inspired Mechanism Empowers Image ClassificationSimone Azeglio, Olivier Marre, Peter Neri, Ulisse FerrariNeurIPS 2025 · 5 citations
- A Slow-I-Fast-P Architecture for Compressed Video Action RecognitionJiapeng Li, Ping Wei, Yongchi Zhang, Nanning ZhengACM MM 2020 · 49 citations
- Video Modeling With Correlation NetworksHeng Wang, Du Tran, Lorenzo Torresani, Matt FeiszliCVPR 2020
- Two-Stream Action Recognition-Oriented Video Super-ResolutionHaochen Zhang, Dong Liu, Zhiwei XiongICCV 2019 · 58 citations
- VA-RED2: Video Adaptive Redundancy ReductionBowen Pan, Rameswar Panda, Camilo Luciano Fosco, Chung-Ching Lin et al.ICLR 2021 · 20 citations
