今日论文合集:cs.SD语音3篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Boosting keyword spotting through on-device learnable user speech  characteristics
标题:通过设备上可学习的用户语音特征增强关键词识别
链接:https://arxiv.org/abs/2403.07802
作者:Cristian Cioflan,Lukas Cavigelli,Luca Benini
备注:5 pages, 3 tables, 2 figures. Accepted as a full paper by the tinyML Research Symposium 2024摘要:用于始终在线的TinyML约束应用程序的关键字定位系统需要现场调优,以提高离线训练分类器在不可见的推理条件下部署时的准确性。适应目标用户的语音特性需要许多域内样本,这些样本在现实世界的场景中通常不可用。此外,当前的设备上学习技术依赖于计算密集型和内存消耗型骨干更新方案,不适合始终在线的电池供电设备。在这项工作中,我们提出了一种新的设备上学习架构,由一个预先训练的骨干和一个用户感知的嵌入学习用户的语音特征。如此生成的特征被融合并用于对输入话语进行分类。对于由看不见的说话者产生的域偏移,我们基于Google Speech Commands数据集的35类问题,通过用户预测的廉价更新,测量错误率从30.1%降低到24.3%。此外,我们证明了我们提出的架构在样本和类稀缺的学习条件下的Few-Shot学习能力。随着23.7 kparameters和1 MFLOP每个epoch所需的设备上的培训,我们的系统是可行的TinyML应用程序,旨在电池供电的微控制器。
摘要:Keyword spotting systems for always-on TinyML-constrained applications require on-site tuning to boost the accuracy of offline trained classifiers when deployed in unseen inference conditions. Adapting to the speech peculiarities of target users requires many in-domain samples, often unavailable in real-world scenarios. Furthermore, current on-device learning techniques rely on computationally intensive and memory-hungry backbone update schemes, unfit for always-on, battery-powered devices. In this work, we propose a novel on-device learning architecture, composed of a pretrained backbone and a user-aware embedding learning the user's speech characteristics. The so-generated features are fused and used to classify the input utterance. For domain shifts generated by unseen speakers, we measure error rate reductions of up to 19% from 30.1% to 24.3% based on the 35-class problem of the Google Speech Commands dataset, through the inexpensive update of the user projections. We moreover demonstrate the few-shot learning capabilities of our proposed architecture in sample- and class-scarce learning conditions. With 23.7 kparameters and 1 MFLOP per epoch required for on-device training, our system is feasible for TinyML applications aimed at battery-powered microcontrollers.


【2】 Multichannel Long-Term Streaming Neural Speech Enhancement for Static  and Moving Speakers
标题:静止和运动说话人的多通道长期流神经语音增强
链接:https://arxiv.org/abs/2403.07675
作者:Changsheng Quan,Xiaofei Li
摘要:在这项工作中,我们扩展了我们以前提出的离线SpatialNet的长期流多通道语音增强在静态和移动扬声器的情况下。SpatialNet利用空间信息,如语音的空间/转向方向,用于区分目标语音和干扰,并取得了优异的性能。SpatialNet的核心是一个窄带自注意模块,用于学习空间向量的时间动态。对于长时间的流语音增强,我们提出用在线网络代替离线自注意网络,在线网络具有线性推理复杂度,同时保持学习长期信息的能力。三种变体是基于(i)掩蔽的自注意,(ii)保留,一种具有线性推理复杂性的自注意变体,以及(iii)Mamba,一种基于结构化状态空间的RNN类网络。此外,我们还研究了不同网络的长度外推能力,即对比训练信号长得多的信号进行测试,并提出了一种短信号训练加长信号微调的策略,在有限的训练时间内大大提高了网络的长度外推能力。总体而言,建议的在线SpatialNet实现了出色的语音增强性能的长音频流,静态和移动扬声器。拟议的方法将在https://github.com/Audio-WestlakeU/NBSS上开源。
摘要:In this work, we extend our previously proposed offline SpatialNet for long-term streaming multichannel speech enhancement in both static and moving speaker scenarios. SpatialNet exploits spatial information, such as the spatial/steering direction of speech, for discriminating between target speech and interferences, and achieved outstanding performance. The core of SpatialNet is a narrow-band self-attention module used for learning the temporal dynamic of spatial vectors. Towards long-term streaming speech enhancement, we propose to replace the offline self-attention network with online networks that have linear inference complexity w.r.t signal length and meanwhile maintain the capability of learning long-term information. Three variants are developed based on (i) masked self-attention, (ii) Retention, a self-attention variant with linear inference complexity, and (iii) Mamba, a structured-state-space-based RNN-like network. Moreover, we investigate the length extrapolation ability of different networks, namely test on signals that are much longer than training signals, and propose a short-signal training plus long-signal fine-tuning strategy, which largely improves the length extrapolation ability of the networks within limited training time. Overall, the proposed online SpatialNet achieves outstanding speech enhancement performance for long audio streams, and for both static and moving speakers. The proposed method will be open-sourced in https://github.com/Audio-WestlakeU/NBSS.


【3】 Gender-ambiguous voice generation through feminine speaking style  transfer in male voices
标题:男性嗓音中女性说话风格迁移的性别歧义语音生成
链接:https://arxiv.org/abs/2403.07661
作者:Maria Koutsogiannaki,Shafel Mc Dowall,Ioannis Agiomyrgiannakis备注:5 pages, 1 figure, 3 tables
摘要:最近,在负责任的人工智能的保护伞下,人们努力开发性别模糊的合成语音,用一个声音代表性别谱中的所有个体。然而,研究工作完全忽略了说话风格,尽管二元和非二元人群之间存在差异。在这项工作中,我们合成性别模糊的语音相结合的男性扬声器的音色与一个女性扬声器的语音方式,使用语音变形和音高向男女边界转移。主观评价表明,变形的样本,传达女性的讲话风格的模糊性高于那些经历纯音高变换,这表明说话风格可以是一个促成因素,在创建性别模糊的讲话。据我们所知,这是第一个明确使用说话风格的转移来创建性别模糊的声音的研究。
摘要:Recently, and under the umbrella of Responsible AI, efforts have been made to develop gender-ambiguous synthetic speech to represent with a single voice all individuals in the gender spectrum. However, research efforts have completely overlooked the speaking style despite differences found among binary and non-binary populations. In this work, we synthesise gender-ambiguous speech by combining the timbre of a male speaker with the manner of speech of a female speaker using voice morphing and pitch shifting towards the male-female boundary. Subjective evaluations indicate that the ambiguity of the morphed samples that convey the female speech style is higher than those that undergo pure pitch transformations suggesting that the speaking style can be a contributing factor in creating gender-ambiguous speech. To our knowledge, this is the first study that explicitly uses the transfer of the speaking style to create gender-ambiguous voices.


eess.AS音频处理
【1】 Beyond the Labels: Unveiling Text-Dependency in Paralinguistic Speech  Recognition Datasets
标题:超越标签:揭示副语言语音识别数据集中的文本依赖
链接:https://arxiv.org/abs/2403.07767
作者:Jan Pešán,Santosh Kesiraju,Lukáš Burget,Jan ''Honza'' Černocký
摘要:认知负荷和情感等副语言特征越来越被认为是语音识别研究的关键领域,通常通过CLSE和IEMOCAP等专业数据集进行检查。然而,这些数据集的完整性很少被仔细检查文本依赖性。本文批判性地评估了普遍的假设,即在这些数据集上训练的机器学习模型真正学会识别非语言特征,而不仅仅是捕获词汇特征。通过研究这些数据集中的词汇重叠和测试机器学习模型的性能,我们发现了显著的文本依赖的特征标记。我们的研究结果表明,一些机器学习模型,特别是像HuBERT这样的大型预训练模型,可能会无意中关注词汇特征,而不是预期的非语言特征。该研究呼吁研究界重新评估现有数据集和方法的可靠性,确保机器学习模型真正学习它们旨在识别的内容。
摘要:Paralinguistic traits like cognitive load and emotion are increasingly recognized as pivotal areas in speech recognition research, often examined through specialized datasets like CLSE and IEMOCAP. However, the integrity of these datasets is seldom scrutinized for text-dependency. This paper critically evaluates the prevalent assumption that machine learning models trained on such datasets genuinely learn to identify paralinguistic traits, rather than merely capturing lexical features. By examining the lexical overlap in these datasets and testing the performance of machine learning models, we expose significant text-dependency in trait-labeling. Our results suggest that some machine learning models, especially large pre-trained models like HuBERT, might inadvertently focus on lexical characteristics rather than the intended paralinguistic features. The study serves as a call to action for the research community to reevaluate the reliability of existing datasets and methodologies, ensuring that machine learning models genuinely learn what they are designed to recognize.


【2】 Gender-ambiguous voice generation through feminine speaking style  transfer in male voices
标题:男性嗓音中女性说话风格迁移的性别歧义语音生成
链接:https://arxiv.org/abs/2403.07661
作者:Maria Koutsogiannaki,Shafel Mc Dowall,Ioannis Agiomyrgiannakis
备注:5 pages, 1 figure, 3 tables
摘要:最近,在负责任的人工智能的保护伞下,人们努力开发性别模糊的合成语音,用一个声音代表性别谱中的所有个体。然而,研究工作完全忽略了说话风格,尽管二元和非二元人群之间存在差异。在这项工作中,我们合成性别模糊的语音相结合的男性扬声器的音色与一个女性扬声器的语音方式,使用语音变形和音高向男女边界转移。主观评价表明,变形的样本,传达女性的讲话风格的模糊性高于那些经历纯音高变换,这表明说话风格可以是一个促成因素,在创建性别模糊的讲话。据我们所知,这是第一个明确使用说话风格的转移来创建性别模糊的声音的研究。
摘要:Recently, and under the umbrella of Responsible AI, efforts have been made to develop gender-ambiguous synthetic speech to represent with a single voice all individuals in the gender spectrum. However, research efforts have completely overlooked the speaking style despite differences found among binary and non-binary populations. In this work, we synthesise gender-ambiguous speech by combining the timbre of a male speaker with the manner of speech of a female speaker using voice morphing and pitch shifting towards the male-female boundary. Subjective evaluations indicate that the ambiguity of the morphed samples that convey the female speech style is higher than those that undergo pure pitch transformations suggesting that the speaking style can be a contributing factor in creating gender-ambiguous speech. To our knowledge, this is the first study that explicitly uses the transfer of the speaking style to create gender-ambiguous voices.

【3】 On HRTF Notch Frequency Prediction Using Anthropometric Features and  Neural Networks
标题:基于人体测量特征和神经网络的HRTF切迹频率预测研究
链接:https://arxiv.org/abs/2403.07579
作者:Lior Arbel,Ishwarya Ananthabhotla,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
摘要:高保真空间音频通常在使用个性化的头部相关传递函数(HRTF)产生时表现更好。然而,直接获取HRTF是麻烦的,并且需要专门的设备。因此,许多个性化方法从容易获得的耳廓、头部和躯干的人体测量特征来估计HRTF特征。已知第一HRTF陷波频率(N1)是仰角定位中的主导特征,并且因此是用于HRTF个性化的有用特征。本文介绍了使用神经模型预测N1频率从耳廓人体测量。预测分别在三个数据库上进行,模拟和测量,然后在数据库之间进行域混合。该模型成功地预测N1频率为个别数据库和一些数据库之间的域混合。预测误差更好或与以前报道的那些相当,当在大型数据库中获取时显示出显着的改善,并且具有更大的输出范围。
摘要:High fidelity spatial audio often performs better when produced using a personalized head-related transfer function (HRTF). However, the direct acquisition of HRTFs is cumbersome and requires specialized equipment. Thus, many personalization methods estimate HRTF features from easily obtained anthropometric features of the pinna, head, and torso. The first HRTF notch frequency (N1) is known to be a dominant feature in elevation localization, and thus a useful feature for HRTF personalization. This paper describes the prediction of N1 frequency from pinna anthropometry using a neural model. Prediction is performed separately on three databases, both simulated and measured, and then by domain mixing in-between the databases. The model successfully predicts N1 frequency for individual databases and by domain mixing between some databases. Prediction errors are better or comparable to those previously reported, showing significant improvement when acquired over a large database and with a larger output range.

【4】 Boosting keyword spotting through on-device learnable user speech  characteristics
标题:通过设备上可学习的用户语音特征增强关键词识别
链接:https://arxiv.org/abs/2403.07802
作者:Cristian Cioflan,Lukas Cavigelli,Luca Benini
备注:5 pages, 3 tables, 2 figures. Accepted as a full paper by the tinyML Research Symposium 2024
摘要:用于始终在线的TinyML约束应用程序的关键字定位系统需要现场调优,以提高离线训练分类器在不可见的推理条件下部署时的准确性。适应目标用户的语音特性需要许多域内样本,这些样本在现实世界的场景中通常不可用。此外,当前的设备上学习技术依赖于计算密集型和内存消耗型骨干更新方案,不适合始终在线的电池供电设备。在这项工作中,我们提出了一种新的设备上学习架构,由一个预先训练的骨干和一个用户感知的嵌入学习用户的语音特征。如此生成的特征被融合并用于对输入话语进行分类。对于由看不见的说话者产生的域偏移,我们基于Google Speech Commands数据集的35类问题,通过用户预测的廉价更新,测量错误率从30.1%降低到24.3%。此外,我们证明了我们提出的架构在样本和类稀缺的学习条件下的Few-Shot学习能力。随着23.7 kparameters和1 MFLOP每个epoch所需的设备上的培训,我们的系统是可行的TinyML应用程序,旨在电池供电的微控制器。
摘要:Keyword spotting systems for always-on TinyML-constrained applications require on-site tuning to boost the accuracy of offline trained classifiers when deployed in unseen inference conditions. Adapting to the speech peculiarities of target users requires many in-domain samples, often unavailable in real-world scenarios. Furthermore, current on-device learning techniques rely on computationally intensive and memory-hungry backbone update schemes, unfit for always-on, battery-powered devices. In this work, we propose a novel on-device learning architecture, composed of a pretrained backbone and a user-aware embedding learning the user's speech characteristics. The so-generated features are fused and used to classify the input utterance. For domain shifts generated by unseen speakers, we measure error rate reductions of up to 19% from 30.1% to 24.3% based on the 35-class problem of the Google Speech Commands dataset, through the inexpensive update of the user projections. We moreover demonstrate the few-shot learning capabilities of our proposed architecture in sample- and class-scarce learning conditions. With 23.7 kparameters and 1 MFLOP per epoch required for on-device training, our system is feasible for TinyML applications aimed at battery-powered microcontrollers.


【5】 Multichannel Long-Term Streaming Neural Speech Enhancement for Static  and Moving Speakers
标题:静止和运动说话人的多通道长期流神经语音增强
链接:https://arxiv.org/abs/2403.07675
作者:Changsheng Quan,Xiaofei Li
摘要:在这项工作中,我们扩展了我们以前提出的离线SpatialNet的长期流多通道语音增强在静态和移动扬声器的情况下。SpatialNet利用空间信息,如语音的空间/转向方向,用于区分目标语音和干扰,并取得了优异的性能。SpatialNet的核心是一个窄带自注意模块,用于学习空间向量的时间动态。对于长时间的流语音增强,我们提出用在线网络代替离线自注意网络,在线网络具有线性推理复杂度,同时保持学习长期信息的能力。三种变体是基于(i)掩蔽的自注意,(ii)保留,一种具有线性推理复杂性的自注意变体,以及(iii)Mamba,一种基于结构化状态空间的RNN类网络。此外,我们还研究了不同网络的长度外推能力,即对比训练信号长得多的信号进行测试,并提出了一种短信号训练加长信号微调的策略,在有限的训练时间内大大提高了网络的长度外推能力。总体而言,建议的在线SpatialNet实现了出色的语音增强性能的长音频流,静态和移动扬声器。拟议的方法将在https://github.com/Audio-WestlakeU/NBSS上开源。
摘要:In this work, we extend our previously proposed offline SpatialNet for long-term streaming multichannel speech enhancement in both static and moving speaker scenarios. SpatialNet exploits spatial information, such as the spatial/steering direction of speech, for discriminating between target speech and interferences, and achieved outstanding performance. The core of SpatialNet is a narrow-band self-attention module used for learning the temporal dynamic of spatial vectors. Towards long-term streaming speech enhancement, we propose to replace the offline self-attention network with online networks that have linear inference complexity w.r.t signal length and meanwhile maintain the capability of learning long-term information. Three variants are developed based on (i) masked self-attention, (ii) Retention, a self-attention variant with linear inference complexity, and (iii) Mamba, a structured-state-space-based RNN-like network. Moreover, we investigate the length extrapolation ability of different networks, namely test on signals that are much longer than training signals, and propose a short-signal training plus long-signal fine-tuning strategy, which largely improves the length extrapolation ability of the networks within limited training time. Overall, the proposed online SpatialNet achieves outstanding speech enhancement performance for long audio streams, and for both static and moving speakers. The proposed method will be open-sourced in https://github.com/Audio-WestlakeU/NBSS.


机器翻译由腾讯交互翻译提供,仅供参考