Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended Labeling
Pol van Rijn, Silvan Mertes, Kathrin Janowski, Katharina Weitz, Nori Jacoby, Elisabeth André
Abstract
Speech is a natural interface for humans to interact with robots. Yet, aligning a robot’s voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots and tested them on a limited number of robots and existing voices. Here, we develop a robot-voice creation tool followed by large-scale behavioral human experiments (N=2,505). First, participants collectively tune robotic voices to match 175 robot images using an adaptive human-in-the-loop pipeline. Then, participants describe their impression of the robot or their matched voice using another human-in-the-loop paradigm for open-ended labeling. The elicited taxonomy is then used to rate robot attributes and to predict the best voice for an unseen robot. We offer a web interface to aid engineers in customizing robot voices, demonstrating the synergy between cognitive science and machine learning for engineering tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d5b9383-e3db-4422-9981-9756bd3242abCited by top-tier papers2
- Futuring Social Assemblages: How Enmeshing AIs into Social Life Challenges the Individual and the InterpersonalLingqing Wang, Yingting Gao, Chidimma Lois Anyi, Ashok K. GoelCHI 2026 · 2 citations
- Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with PeopleDun-Ming Huang, Pol van Rijn, Ilia Sucholutsky, Raja Marjieh et al.ACL 2024
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech SynthesisRafael Valle, Kevin J. Shih, Ryan Prenger, Bryan CatanzaroICLR 2021 · 133 citations
- Gibbs Sampling with PeoplePeter M. C. Harrison, Raja Marjieh, Federico Adolfi, Pol van Rijn et al.NeurIPS 2020 · 81 citations
Related papers
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech InteractionXiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou et al.ICLR 2026 · 1 citation
- Investigating Effect of Altered Auditory Feedback on Self-Representation, Subjective Operator Experience, and Task Performance in Teleoperation of a Social RobotNami Ogawa, Jun Baba, Junya NakanishiCHI 2024 · 8 citations
- SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian et al.ACL 2025 · 2 citations
- The R-U-A-Robot Dataset: Helping Avoid Chatbot Deception by Detecting User Questions About Human or Non-Human IdentityDavid Gros, Yu Li, Zhou YuACL 2021
- Weirding Haptics: In-Situ Prototyping of Vibrotactile Feedback in Virtual Reality through VocalizationDonald Degraen, Bruno Fruchard, Frederik Smolders, Emmanouil Potetsianakis et al.UIST 2021 · 41 citations
