Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
Hongfei Du, Jiacheng Shi, Sidi Lu, Gang Zhou, Ashley Gao
Abstract
Integrating large language models (LLMs) into text-to-speech (TTS) systems has improved speech expressiveness, yet interpretable emotional control remains challenging. Existing approaches primarily rely on external conditioning or global activation steering, offering limited insight into the internal representations underlying emotional control. In this work, we analyze emotion-related variation in the semantic hidden states of LLM-based TTS models using sparse autoencoders (SAEs) to identify sparse latent features. Our analysis shows that emotional variation is distributed across multiple sparse latent features, while intervening on a small subset enables interpretable emotion control. Building on this observation, we introduce a feature-level intervention framework for bidirectional emotion induction and suppression without modifying backbone parameters. We further show that distinct latent features are associated with specific acoustic attributes (e.g., pitch), suggesting that emotional expression arises from coordinated latent contributions rather than a single global shift. Empirically, steering these sparse latent features achieves comparable or superior emotion induction and suppression performance relative to global steering and existing TTS baselines. GitHub-Demo
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d2ff4c5-fa1d-40dd-bbd5-d0505c7a28b7Builds on11
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations CategoriesTianlong Wang, Xianfeng Jiao, Yinghao Zhu, Zhongzhi Chen et al.WWW 2025 · 64 citations
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou et al.AAAI 2026 · 63 citations
Related papers
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani et al.EMNLP 2025 · 1 citation
- Uncovering Sentiment Analysis Circuit in Large Language ModelShichen Li, Zhouyang Wang, Zhongqing Wang, Peifeng LiACL 2026
- CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation SteeringSiyi Wang, Shihong Tan, Siyi Liu, Hong Jia et al.ICML 2026
- Semi-Supervised Generative Modeling for Controllable Speech SynthesisRaza Habib, Soroosh Mariooryad, Matt Shannon, Eric Battenberg et al.ICLR 2020 · 48 citations
- Sparse Latents Steer Retrieval-Augmented GenerationChunlei Xin, Shuheng Zhou, Huijia Zhu, Weiqiang Wang et al.ACL 2025 · 2 citations
