Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition
Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L Mishara
Abstract
Speech emotion recognition (SER) hinges on front-ends that expose informative time-frequency (TF) structure from raw speech. Classical short-time Fourier and wavelet transforms impose fixed resolution trade-offs, while prior ”superlet” variants rely on integer orders and hand-tuned hyperparameters. We revisit TF analysis from first principles and formulate a learnable continuum of superlet transforms. Starting from DC-corrected analytic Morlet wavelets, we define superlets as multiplicative ensembles of wavelet responses and realize learnable fractional orders via softmax-normalized weights over discrete orders, computed as a logdomain geometric mean. We establish admissibility (zero mean) and continuity in order and frequency, and characterize approximate analyticity by bounding negative-frequency leakage as a function of an effective cycle parameter. Building on these results, we introduce the Learnable Fractional Superlet Transform (LFST), a fully differentiable front-end that jointly optimizes (i) a monotone, logspaced frequency grid, (ii) frequency-dependent base cycles, and (iii) learnable fractional-order weights, all trained end-to-end. LFST further includes a learnable asymmetric hard-thresholding (LAHT) module that promotes sparse, denoised TF activations while preserving transients; we provide sufficient conditions for boundedness and stability under mild cycle and grid constraints. To exploit LFST for SER, we design a compact Spectro-Temporal Emotion Encoder (STEE), achieving strong performance with a parameter budget that is orders of magnitude smaller than large self-supervised models, at the cost of additional frontend computation compared to STFT- or LEAF-based baselines. STEE consumes two-channel TF maps, magnitude S and phase-congruency κ, through a compact multi-scale stack with residual temporal and depthwise-frequency blocks, Adaptive FiLM gating, axial (time-axis) self-attention, global attentive pooling, and a lightweight classifier. The full LFST+STEE system is trained in a standard train-validate-test regime using focal loss with optional class rebalancing, and is validated on IEMOCAP, EMO-DB, and the private NSPL-CRISE dataset under standard protocols. By unifying a principled, learnable TF transform with a compact encoder, LFST+STEE replaces ad hoc front-ends with a mathematically grounded alternative that is differentiable, stable, and adaptable to data, enabling systematic ablations over frequency grids, cycle schedules, and fractional orders within a single end-to-end model. The source code of this paper is shared on the GitHub repository: https://github.com/alaaNfissi/LFST-for-SER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- LEAF: A Learnable Frontend for Audio ClassificationNeil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco TagliasacchiICLR 2021 · 181 citations
Related papers
- SEBSFormer: A Spectral-Enhanced Bi-Stream Transformer for Robust EEG DecodingLin Zhang, Shikui Tu, Lei XuAAAI 2026
- Signal Enhancement via Multi-view Dynamic Representation and Alignment-aware FusionZikun Jin, Yuhua Qian, Xinyan Liang, Jiaqian Zhang et al.AAAI 2026
- ASTDF-Net: Attention-Based Spatial-Temporal Dual-Stream Fusion Network for EEG-Based Emotion RecognitionPeiliang Gong, Ziyu Jia, Pengpai Wang, Yueying Zhou et al.ACM MM 2023 · 38 citations
- Unsupervised Domain Adaptation Integrating Transformer and Mutual Information for Cross-Corpus Speech Emotion RecognitionShiqing Zhang, Ruixin Liu, Yijiao Yang, Xiaoming Zhao et al.ACM MM 2022 · 19 citations
- RECTOR: Masked Region-Channel-Temporal Modeling for Affective and Cognitive Representation LearningJinhan Liu, Mahsa ShoaranICML 2026
