今日论文合集:cs.SD语音4篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】An Investigation of Incorporating Mamba for Speech Enhancement
标题:使用曼巴语增强言语功能的研究
链接:https://arxiv.org/abs/2405.06573
作者:Rong Chao,Wen-Huang Cheng,Moreno La Quatra,Sabato Marco Siniscalchi,Chao-Han Huck Yang,Szu-Wei Fu,Yu Tsao
摘要:这项工作的目的是研究一个可扩展的状态空间模型(SSM),曼巴,语音增强(SE)的任务。我们利用一个基于Mamba的回归模型来表征语音信号,并建立一个SE系统上的Mamba,称为SEMamba。我们通过将Mamba作为基本和高级SE系统的核心模型进行集成,并利用信号级距离和面向度量的损失函数来探索Mamba的属性。SEMamba展示了有希望的结果,并在VoiceBank-DEMAND数据集上获得了3.55的PESQ分数。当与感知对比度拉伸技术相结合时,所提出的SEMamba产生了一个新的最先进的PESQ得分为3.69。
摘要:This work aims to study a scalable state-space model (SSM), Mamba, for the speech enhancement (SE) task. We exploit a Mamba-based regression model to characterize speech signals and build an SE system upon Mamba, termed SEMamba. We explore the properties of Mamba by integrating it as the core model in both basic and advanced SE systems, along with utilizing signal-level distances as well as metric-oriented loss functions. SEMamba demonstrates promising results and attains a PESQ score of 3.55 on the VoiceBank-DEMAND dataset. When combined with the perceptual contrast stretching technique, the proposed SEMamba yields a new state-of-the-art PESQ score of 3.69.


【2】 Look Once to Hear: Target Speech Hearing with Noisy Examples
标题:看一次听:目标言语听力有噪音的例子
链接:https://arxiv.org/abs/2405.06289
作者:Bandhav Veluri,Malek Itani,Tuochao Chen,Takuya Yoshioka,Shyamnath Gollakota
备注:Honorable mention at CHI 2024
摘要:在拥挤的环境中,人类大脑可以专注于目标说话者的讲话,因为他们事先知道他们的声音。我们介绍了一种新的智能可听系统,实现了这一能力,使目标语音听觉忽略所有干扰语音和噪声,但目标扬声器。一种简单的方法是要求一个干净的语音示例来注册目标说话者。然而,这与可听应用领域并不一致,因为在现实世界的场景中获得干净的示例是具有挑战性的,从而产生独特的用户界面问题。我们提出了第一个注册界面,其中佩戴者看着目标扬声器几秒钟,以捕获目标扬声器的单个,短,高噪声,双耳的例子。该噪声示例用于在存在干扰扬声器和噪声的情况下进行登记和随后的语音提取。我们的系统实现了7.01 dB的信号质量改善,使用不到5秒的嘈杂的注册音频,并可以在6.24毫秒的嵌入式CPU上处理8毫秒的音频块。我们的用户研究证明了在以前看不见的室内和室外多径环境中对真实世界的静态和移动扬声器的泛化。最后,与干净的示例相比,我们的噪声示例注册界面不会导致性能下降,同时方便且用户友好。退一步讲,本文朝着用人工智能增强人类听觉感知迈出了重要的一步。我们提供代码和数据:https://github.com/vb000/LookOnceToHear。
摘要:In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achieves this capability, enabling target speech hearing to ignore all interfering speech and noise, but the target speaker. A naive approach is to require a clean speech example to enroll the target speaker. This is however not well aligned with the hearable application domain since obtaining a clean example is challenging in real world scenarios, creating a unique user interface problem. We present the first enrollment interface where the wearer looks at the target speaker for a few seconds to capture a single, short, highly noisy, binaural example of the target speaker. This noisy example is used for enrollment and subsequent speech extraction in the presence of interfering speakers and noise. Our system achieves a signal quality improvement of 7.01 dB using less than 5 seconds of noisy enrollment audio and can process 8 ms of audio chunks in 6.24 ms on an embedded CPU. Our user studies demonstrate generalization to real-world static and mobile speakers in previously unseen indoor and outdoor multipath environments. Finally, our enrollment interface for noisy examples does not cause performance degradation compared to clean examples, while being convenient and user-friendly. Taking a step back, this paper takes an important step towards enhancing the human auditory perception with artificial intelligence. We provide code and data at: https://github.com/vb000/LookOnceToHear.


【3】 Muting Whisper: A Universal Acoustic Adversarial Attack on Speech  Foundation Models
标题:静音低语:对语音基础模型的通用声学对抗攻击
链接:https://arxiv.org/abs/2405.06134
作者:Vyas Raina,Rao Ma,Charles McGhee,Kate Knill,Mark Gales
摘要:最近,像Whisper这样的大型语音基础模型的发展使得它们在许多自动语音识别(ASR)应用中得到了广泛的使用。这些系统在它们的词汇表中加入了“特殊标记”,例如$\texttt{}$,以指导它们的语言生成过程。然而,我们证明了这些令牌可以被对抗性攻击用来操纵模型的行为。我们提出了一种简单而有效的方法来学习Whisper的$\texttt{ }$令牌的通用声学实现,该令牌在预先添加到任何语音信号时,鼓励模型忽略语音,只转录特殊令牌,有效地“静音”模型。我们的实验表明,相同的通用0.64秒对抗性音频片段可以成功地使超过97%的语音样本的目标Whisper ASR模型静音。此外,我们发现这种通用的对抗性音频片段经常转移到新的数据集和任务中。总的来说,这项工作证明了Whisper模型对“静音”对抗性攻击的脆弱性,这种攻击在现实环境中可能带来风险和潜在利益:例如,攻击可以用来绕过语音审核系统,或者相反,攻击也可以用来保护私人语音数据。
摘要:Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications. These systems incorporate `special tokens' in their vocabulary, such as $\texttt{}$, to guide their language generation process. However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior. We propose a simple yet effective method to learn a universal acoustic realization of Whisper's $\texttt{}$ token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively `muting' the model. Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97\% of speech samples. Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks. Overall this work demonstrates the vulnerability of Whisper models to `muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data.

【4】 Sound training platform applied to astronomy
标题:应用于天文学的完善训练平台
链接:https://arxiv.org/abs/2405.06042
作者:Natasha Bertaina Lucero,Johanna Casado,Beatriz García,Gonzalo Cayo备注:4 pages, 4 figures, preprint of the III Workshop on Astronomy Beyond the Common Senses for Accessibility and Inclusion
摘要:天文学和数据声化之间的融合代表了宇宙信息方法和分析的重大进步。通过超越天文学数据分析的视觉排他性,创新项目开发了超越视觉表现的软件,将数据转换为听觉和触觉显示。然而,已经证明,这种新颖的技术需要专门的训练,特别是对于音频格式数据。这项工作描述了一个平台的初步发展,旨在通过声化提供天文学数据分析培训。将这些工具纳入天文学教育和研究开辟了新的视野,有助于以更具包容性和多感官的方式参与空间科学探索。
摘要:The convergence between astronomy and data sonification represents a significant advancement in the approach and analysis of cosmic information. By surpassing the visual exclusivity in data analysis in astronomy, innovative projects have developed software that goes beyond visual representation, transforming data into auditory and tactile displays. However, it has been evidenced that this novel technique requires specialized training, particularly for audio format data. This work describes the initial development of a platform aimed at providing training for data analysis in astronomy through sonification. The integration of these tools in astronomical education and research opens new horizons, facilitating a more inclusive and multisensory participation in the exploration of space science.

eess.AS音频处理
【1】 An Investigation of Incorporating Mamba for Speech Enhancement
标题:使用曼巴语增强言语功能的研究
链接:https://arxiv.org/abs/2405.06573
作者:Rong Chao,Wen-Huang Cheng,Moreno La Quatra,Sabato Marco Siniscalchi,Chao-Han Huck Yang,Szu-Wei Fu,Yu Tsao
摘要:这项工作的目的是研究一个可扩展的状态空间模型(SSM),曼巴,语音增强(SE)的任务。我们利用一个基于Mamba的回归模型来表征语音信号,并建立一个SE系统上的Mamba,称为SEMamba。我们通过将Mamba作为基本和高级SE系统的核心模型进行集成,并利用信号级距离和面向度量的损失函数来探索Mamba的属性。SEMamba展示了有希望的结果,并在VoiceBank-DEMAND数据集上获得了3.55的PESQ分数。当与感知对比度拉伸技术相结合时,所提出的SEMamba产生了一个新的最先进的PESQ得分为3.69。
摘要:This work aims to study a scalable state-space model (SSM), Mamba, for the speech enhancement (SE) task. We exploit a Mamba-based regression model to characterize speech signals and build an SE system upon Mamba, termed SEMamba. We explore the properties of Mamba by integrating it as the core model in both basic and advanced SE systems, along with utilizing signal-level distances as well as metric-oriented loss functions. SEMamba demonstrates promising results and attains a PESQ score of 3.55 on the VoiceBank-DEMAND dataset. When combined with the perceptual contrast stretching technique, the proposed SEMamba yields a new state-of-the-art PESQ score of 3.69.


【2】 Look Once to Hear: Target Speech Hearing with Noisy Examples
标题:看一次听:目标言语听力有噪音的例子
链接:https://arxiv.org/abs/2405.06289
作者:Bandhav Veluri,Malek Itani,Tuochao Chen,Takuya Yoshioka,Shyamnath Gollakota
备注:Honorable mention at CHI 2024
摘要:在拥挤的环境中,人类大脑可以专注于目标说话者的讲话,因为他们事先知道他们的声音。我们介绍了一种新的智能可听系统,实现了这一能力,使目标语音听觉忽略所有干扰语音和噪声,但目标扬声器。一种简单的方法是要求一个干净的语音示例来注册目标说话者。然而,这与可听应用领域并不一致,因为在现实世界的场景中获得干净的示例是具有挑战性的,从而产生独特的用户界面问题。我们提出了第一个注册界面,其中佩戴者看着目标扬声器几秒钟,以捕获目标扬声器的单个,短,高噪声,双耳的例子。该噪声示例用于在存在干扰扬声器和噪声的情况下进行登记和随后的语音提取。我们的系统实现了7.01 dB的信号质量改善,使用不到5秒的嘈杂的注册音频,并可以在6.24毫秒的嵌入式CPU上处理8毫秒的音频块。我们的用户研究证明了在以前看不见的室内和室外多径环境中对真实世界的静态和移动扬声器的泛化。最后,与干净的示例相比,我们的噪声示例注册界面不会导致性能下降,同时方便且用户友好。退一步讲,本文朝着用人工智能增强人类听觉感知迈出了重要的一步。我们提供代码和数据:https://github.com/vb000/LookOnceToHear。
摘要:In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achieves this capability, enabling target speech hearing to ignore all interfering speech and noise, but the target speaker. A naive approach is to require a clean speech example to enroll the target speaker. This is however not well aligned with the hearable application domain since obtaining a clean example is challenging in real world scenarios, creating a unique user interface problem. We present the first enrollment interface where the wearer looks at the target speaker for a few seconds to capture a single, short, highly noisy, binaural example of the target speaker. This noisy example is used for enrollment and subsequent speech extraction in the presence of interfering speakers and noise. Our system achieves a signal quality improvement of 7.01 dB using less than 5 seconds of noisy enrollment audio and can process 8 ms of audio chunks in 6.24 ms on an embedded CPU. Our user studies demonstrate generalization to real-world static and mobile speakers in previously unseen indoor and outdoor multipath environments. Finally, our enrollment interface for noisy examples does not cause performance degradation compared to clean examples, while being convenient and user-friendly. Taking a step back, this paper takes an important step towards enhancing the human auditory perception with artificial intelligence. We provide code and data at: https://github.com/vb000/LookOnceToHear.

【3】 Lost in Transcription: Identifying and Quantifying the Accuracy Biases  of Automatic Speech Recognition Systems Against Disfluent Speech
标题:迷失在转录中:识别和量化自动语音识别系统针对不流利语音的准确性偏差
链接:https://arxiv.org/abs/2405.06150
作者:Dena Mujtaba,Nihar R. Mahapatra,Megan Arney,J. Scott Yaruss,Hope Gerlach-Houck,Caryn Herring,Jia Bin
备注:Accepted to NAACL 2024
摘要:自动语音识别(ASR)系统在教育、医疗保健、就业和移动技术中越来越普遍,但在包容性方面面临着重大挑战,特别是对于全球8000万口吃者来说。这些系统通常无法准确地解释偏离典型流畅性的语音模式,导致关键的可用性问题和误解。这项研究评估了六种领先的ASR,分析了它们在口吃者语音样本的真实数据集和来自广泛使用的LibriSpeech基准的合成数据集上的表现。该合成数据集被独特地设计为包含各种口吃事件,能够深入分析每个ASR对不流利语音的处理。我们的综合评估包括单词错误率(WER),字符错误率(CER)和成绩单的语义准确性等指标。结果表明,一致的和统计上显着的准确性偏差,所有的ASR对不流利的讲话,表现在显着的句法和语义不准确的transmittance。这些发现突出了当前ASR技术的关键差距,强调了有效的偏见缓解策略的必要性。解决这种偏见不仅是为了提高口吃者对技术的可用性,也是为了确保他们公平和包容地参与快速发展的数字环境。
摘要:Automatic speech recognition (ASR) systems, increasingly prevalent in education, healthcare, employment, and mobile technology, face significant challenges in inclusivity, particularly for the 80 million-strong global community of people who stutter. These systems often fail to accurately interpret speech patterns deviating from typical fluency, leading to critical usability issues and misinterpretations. This study evaluates six leading ASRs, analyzing their performance on both a real-world dataset of speech samples from individuals who stutter and a synthetic dataset derived from the widely-used LibriSpeech benchmark. The synthetic dataset, uniquely designed to incorporate various stuttering events, enables an in-depth analysis of each ASR's handling of disfluent speech. Our comprehensive assessment includes metrics such as word error rate (WER), character error rate (CER), and semantic accuracy of the transcripts. The results reveal a consistent and statistically significant accuracy bias across all ASRs against disfluent speech, manifesting in significant syntactical and semantic inaccuracies in transcriptions. These findings highlight a critical gap in current ASR technologies, underscoring the need for effective bias mitigation strategies. Addressing this bias is imperative not only to improve the technology's usability for people who stutter but also to ensure their equitable and inclusive participation in the rapidly evolving digital landscape.

【4】 Muting Whisper: A Universal Acoustic Adversarial Attack on Speech  Foundation Models
标题:静音低语:对语音基础模型的通用声学对抗攻击
链接:https://arxiv.org/abs/2405.06134
作者:Vyas Raina,Rao Ma,Charles McGhee,Kate Knill,Mark Gales
摘要:最近,像Whisper这样的大型语音基础模型的发展使得它们在许多自动语音识别(ASR)应用中得到了广泛的使用。这些系统在它们的词汇表中加入了“特殊标记”,例如$\texttt{}$,以指导它们的语言生成过程。然而,我们证明了这些令牌可以被对抗性攻击用来操纵模型的行为。我们提出了一种简单而有效的方法来学习Whisper的$\texttt{ }$令牌的通用声学实现,该令牌在预先添加到任何语音信号时,鼓励模型忽略语音,只转录特殊令牌,有效地“静音”模型。我们的实验表明,相同的通用0.64秒对抗性音频片段可以成功地使超过97%的语音样本的目标Whisper ASR模型静音。此外,我们发现这种通用的对抗性音频片段经常转移到新的数据集和任务中。总的来说,这项工作证明了Whisper模型对“静音”对抗性攻击的脆弱性,这种攻击在现实环境中可能带来风险和潜在利益:例如,攻击可以用来绕过语音审核系统,或者相反,攻击也可以用来保护私人语音数据。
摘要:Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications. These systems incorporate `special tokens' in their vocabulary, such as $\texttt{}$, to guide their language generation process. However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior. We propose a simple yet effective method to learn a universal acoustic realization of Whisper's $\texttt{}$ token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively `muting' the model. Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97\% of speech samples. Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks. Overall this work demonstrates the vulnerability of Whisper models to `muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data.

【5】 Sound training platform applied to astronomy
标题:应用于天文学的完善训练平台
链接:https://arxiv.org/abs/2405.06042
作者:Natasha Bertaina Lucero,Johanna Casado,Beatriz García,Gonzalo Cayo
备注:4 pages, 4 figures, preprint of the III Workshop on Astronomy Beyond the Common Senses for Accessibility and Inclusion
摘要:天文学和数据声化之间的融合代表了宇宙信息方法和分析的重大进步。通过超越天文学数据分析的视觉排他性,创新项目开发了超越视觉表现的软件,将数据转换为听觉和触觉显示。然而,已经证明,这种新颖的技术需要专门的训练,特别是对于音频格式数据。这项工作描述了一个平台的初步开发,旨在通过声化提供天文学数据分析培训。将这些工具纳入天文学教育和研究开辟了新的视野,有助于以更具包容性和多感官的方式参与空间科学探索。
摘要:The convergence between astronomy and data sonification represents a significant advancement in the approach and analysis of cosmic information. By surpassing the visual exclusivity in data analysis in astronomy, innovative projects have developed software that goes beyond visual representation, transforming data into auditory and tactile displays. However, it has been evidenced that this novel technique requires specialized training, particularly for audio format data. This work describes the initial development of a platform aimed at providing training for data analysis in astronomy through sonification. The integration of these tools in astronomical education and research opens new horizons, facilitating a more inclusive and multisensory participation in the exploration of space science.


机器翻译由腾讯交互翻译提供,仅供参考