From Naturalness to Norms: Interactional Cultural Competence for SpeechLMs
T. Y. S. S. Santosh
摘要
Spoken language models (SpeechLMs) are increasingly real-time conversational actors. Yet many culturally consequential aspects of spoken interaction are not primarily lexical. Across sociolinguistics, linguistic anthropology, and conversation analysis, meaning emerges through how talk is produced and coordinated-prosody, timing, turn-taking, overlap, backchannels, and repair-within situated speech events. A transcript can be semantically correct yet interactionally inappropriate because many culture-bearing signals are audible and sequential rather than textual. This position paper argues for a speech-first view of cultural competence as interactional competence: the ability of a spoken agent to participate appropriately in event-situated interaction with locally normative conduct, while allowing plural acceptable realizations. Here, appropriate does not imply generic human-likeness; in many applications, the desired behavior may instead be constrained, neutral, predictable, or tool-like under an application-specific interaction contract. We synthesize social-science foundations into a theory-derived taxonomy of culture-bearing signals in speech, identify interactional phenomena where transcript correctness fails to predict appropriateness, and ground the agenda in today's SpeechLM stacks and evaluation practice. We propose an evaluation framing that complements WER/MOS and broad capability suites by making speech events and interaction contracts explicit, diagnosing where modern pipelines lose interactional cues, and treating cultural appropriateness as a norm-conditioned target rather than generic "naturalness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking DynamicsSiddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang 等ICLR 2025
- Research Borderlands: Analysing Writing Across Research CulturesShaily Bhatt, Tal August, Maria AntoniakACL 2025
- SocialCC: Interactive Evaluation for Cultural Competence in Language AgentsJincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. MengACL 2025 · 被引用 8 次
- LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken DialoguesAmir Ivry, Shinji WatanabeICML 2026 · 被引用 4 次
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang 等ICLR 2026 · 被引用 143 次
