今日论文合集:eess.AS音频处理3篇。

本文经arXiv每日学术速递授权转载


eess.AS音频处理

【1】 Dialogue Understandability: Why are we streaming movies with subtitles?
标题:对话可理解性:为什么我们要播放带有字幕的电影?
链接:https://arxiv.org/abs/2403.15336
作者:Helard Becerra,Alessandro Ragano,Diptasree Debnath,Asad Ullah,Crisron Rudolf Lucas,Martin Walsh,Andrew Hines
摘要:观看启用字幕的电影和电视节目不仅仅取决于可听度或语音清晰度。与技术进步、电影制作和社会行为有关的各种不断变化的因素挑战着我们的感知和理解。本研究旨在正式和背景下,这些影响因素更广泛和新颖的术语称为对话的可理解性。我们提出了一个工作定义为对话的可理解性是一个听众的能力,遵循的故事,而不需要不适当的认知努力或浓度,影响他们的体验质量(QoE)。本文确定,描述和分类的因素,影响对话的可理解性映射他们的QoE框架,媒体流的生命周期,和利益相关者。然后,我们探索文献中可用的测量工具,并将它们与可能使用的因素联系起来。这些工具的成熟度和适用性进行了评估,在一组试点实验。最后,我们反思了仍然需要填补的空白,我们可以测量什么,什么不可以,未来的主观实验,以及新的研究趋势,可以帮助我们充分验证对话的可理解性。
摘要:Watching movies and TV shows with subtitles enabled is not simply down to audibility or speech intelligibility. A variety of evolving factors related to technological advances, cinema production and social behaviour challenge our perception and understanding. This study seeks to formalise and give context to these influential factors under a wider and novel term referred to as Dialogue Understandability. We propose a working definition for Dialogue Understandability being a listener's capacity to follow the story without undue cognitive effort or concentration being required that impacts their Quality of Experience (QoE). The paper identifies, describes and categorises the factors that influence Dialogue Understandability mapping them over the QoE framework, a media streaming lifecycle, and the stakeholders involved. We then explore available measurement tools in the literature and link them to the factors they could potentially be used for. The maturity and suitability of these tools is evaluated over a set of pilot experiments. Finally, we reflect on the gaps that still need to be filled, what we can measure and what not, future subjective experiments, and new research trends that could help us to fully characterise Dialogue Understandability.

【2】 Crowdsourced Multilingual Speech Intelligibility Testing
标题:众包多语言语音可懂度测试
链接:https://arxiv.org/abs/2403.14817
作者:Laura Lechler,Kamil Wojcicki
摘要:随着生成音频功能的出现,越来越需要快速评估其对语音清晰度的影响。除了现有的实验室措施,这是昂贵的,并没有很好的规模,有相对较少的工作,众包的可理解性评估。标准和建议尚待确定,公开提供的多语种测试材料也缺乏。为了应对这一挑战,我们提出了一种众包可懂度评估的方法。我们详细介绍了测试设计,收集和公开发布的多语种语音数据,我们的早期实验的结果。
摘要:With the advent of generative audio features, there is an increasing need for rapid evaluation of their impact on speech intelligibility. Beyond the existing laboratory measures, which are expensive and do not scale well, there has been comparatively little work on crowdsourced assessment of intelligibility. Standards and recommendations are yet to be defined, and publicly available multilingual test materials are lacking. In response to this challenge, we propose an approach for a crowdsourced intelligibility assessment. We detail the test design, the collection and public release of the multilingual speech data, and the results of our early experiments.

【3】 Visually Grounded Speech Models have a Mutual Exclusivity Bias
标题:基于视觉的语音模型存在相互排他性偏差
链接:https://arxiv.org/abs/2403.13922
作者:Leanne Nortje,Dan Oneaţă,Yevgen Matusevych,Herman Kamper
备注:Accepted to TACL, pre-MIT Press publication version
摘要:当儿童学习新单词时,他们会使用诸如互斥性(ME)偏见等约束条件:一个新单词映射到一个新对象,而不是熟悉的对象。这种偏差已经在计算上进行了研究,但仅在使用离散单词表示作为输入的模型中,忽略了口语单词的高度可变性。我们调查ME偏见的背景下,视觉接地语音模型,从自然图像和连续语音音频学习。具体来说,我们训练一个熟悉的单词模型,并通过要求它在一个新的单词和一个熟悉的对象之间进行选择来测试它的ME偏见。为了模拟先前的声学和视觉知识,我们使用预训练的语音和视觉网络进行了几种初始化策略的实验。我们的研究结果揭示了不同初始化方法的ME偏差,在具有更多先验知识(特别是视觉知识)的模型中具有更强的偏差。额外的测试证实了我们的结果的鲁棒性,即使在考虑不同的损失函数。
摘要:When children learn new words, they employ constraints such as the mutual exclusivity (ME) bias: a novel word is mapped to a novel object rather than a familiar one. This bias has been studied computationally, but only in models that use discrete word representations as input, ignoring the high variability of spoken words. We investigate the ME bias in the context of visually grounded speech models that learn from natural images and continuous speech audio. Concretely, we train a model on familiar words and test its ME bias by asking it to select between a novel and a familiar object when queried with a novel word. To simulate prior acoustic and visual knowledge, we experiment with several initialisation strategies using pretrained speech and vision networks. Our findings reveal the ME bias across the different initialisation approaches, with a stronger bias in models with more prior (in particular, visual) knowledge. Additional tests confirm the robustness of our results, even when different loss functions are considered.

机器翻译由腾讯交互翻译提供,仅供参考