Text-Free Image-to-Speech Synthesis Using Learned Segmental Units
Wei-Ning Hsu, David Harwath, Tyler Miller, Christopher Song, James R. Glass
Abstract
In this paper we present the first model for directly synthesizing fluent, naturalsounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision. Instead, we connect the image captioning module and the speech synthesis module with a set of discrete, sub-word speech units that are discovered with a self-supervised visual grounding task. We conduct experiments on the Flickr8k spoken caption dataset in addition to a novel corpus of spoken audio captions collected for the popular MSCOCO dataset, demonstrating that our generated captions also capture diverse visual semantics of the images they describe. We investigate several different intermediate speech representations, and empirically find that the representation must satisfy several important properties in order to work well. a person in a blue jacket is on a snowboard on a snow covered slope a snowboarder is snowboarding on the side of the mountain a snowboarder is snowboarding on the side of the mountain Same unit sequence, different speakers Different unit sequences, same speaker * Equal contribution † The author performed the work while at MIT, and is now at Facebook AI Research Self-Supervised Learning for Speech and Audio Processing Workshop @ NeurIPS 2020.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66fc98a1-e67c-4c27-a72b-91e35bac0a03Cited by top-tier papers8
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsGuangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen et al.ICML 2024 · 92 citations
- ConceptBeam: Concept Driven Target Speech ExtractionYasunori Ohishi, Marc Delcroix, Tsubasa Ochiai, Shoko Araki et al.ACM MM 2022 · 18 citations
- TranSpeech: Speech-to-Speech Translation With Bilateral PerturbationRongjie Huang, Jinglin Liu, Huadai Liu, Yi Ren et al.ICLR 2023 · 17 citations
- Video-Guided Curriculum Learning for Spoken Video GroundingYan Xia, Zhou Zhao, Shangwei Ye, Yang Zhao et al.ACM MM 2022 · 7 citations
Builds on8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja et al.ICCV 2019 · 113 citations
Related papers
- SEAR: Semantically-grounded Audio RepresentationsRajat Hebbar, Digbalay Bose, Shrikanth NarayananACM MM 2023
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go et al.EMNLP 2025
- Bridging the Gap between Vision and Language Domains for Improved Image CaptioningFenglin Liu, Xian Wu, Shen Ge, Xiaoyu Zhang et al.ACM MM 2020 · 13 citations
- Improving Image Captioning with Better Use of CaptionZhan Shi, Xu Zhou, Xipeng Qiu, Xiaodan ZhuACL 2020 · 84 citations
- Direct Speech-to-Speech Translation With Discrete UnitsAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu et al.ACL 2022 · 235 citations
