Hear Me Out: A Study on the Use of the Voice Modality for Crowdsourced Relevance Assessments
Nirmal Roy, Agathe Balayn, David Maxwell, Claudia Hauff
Abstract
The creation of relevance assessments by human assessors (often nowadays crowdworkers) is a vital step when building IR test collections. Prior works have investigated assessor quality & behaviour, and tooling to support assessors in their task. We have few insights though into the impact of a document's presentation modality on assessor efficiency and effectiveness. Given the rise of voice-based interfaces, we investigate whether it is feasible for assessors to judge the relevance of text documents via a voice-based interface. We ran a user study (n = 49) on a crowdsourcing platform where participants judged the relevance of short and long documents- sampled from the TREC Deep Learning corpus-presented to them either in the text or voice modality. We found that: (i) participants are equally accurate in their judgements across both the text and voice modality; (ii) with increased document length it takes partic- ipants significantly longer (for documents of length > 120 words it takes almost twice as much time) to make relevance judgements in the voice condition; and (iii) the ability of assessors to ignore stimuli that are not relevant (i.e., inhibition) impacts the assessment quality in the voice modality-assessors with higher inhibition are significantly more accurate than those with lower inhibition. Our results indicate that we can reliably leverage the voice modality as a means to effectively collect relevance labels from crowdworkers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbf4ff0a-4dac-4d3d-ae5b-ced6949e22edBuilds on3
- Designing Voice Interfaces: Back to the (Curriculum) BasicsChristine Murad, Cosmin MunteanuCHI 2020 · 31 citations
- "Hi! I am the Crowd Tasker" Crowdsourcing through Digital Voice AssistantsDanula Hettiachchi, Zhanna Sarsenbayeva, Fraser Allison, Niels van Berkel et al.CHI 2020 · 22 citations
- Karamad: A Voice-based Crowdsourcing Platform for Underserved PopulationsShan M. Randhawa, Tallal Ahmad, Jay Chen, Agha Ali RazaCHI 2021 · 19 citations
Related papers
- Choice of Voices: A Large-Scale Evaluation of Text-to-Speech Voice Quality for Long-Form ContentJulia Cambre, Jessica Colnago, Jim Maddock, Janice Y. Tsai et al.CHI 2020 · 65 citations
- Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance JudgmentsKevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca DemartiniWWW 2022 · 7 citations
- Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsJüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak et al.SIGIR 2026 · 1 citation
- Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-CheckingKevin Roitero, Dustin Wright, Michael Soprano, Isabelle Augenstein et al.SIGIR 2025 · 5 citations
- Learning to Rank with Multi-Criteria LLM-Judge AnnotationsNaghmeh Farzi, Laura DietzSIGIR 2026
