cs.SD语音,共计3篇,eess.AS音频处理,共计4篇
1.cs.SD语音:
【1】 DT-SV: A Transformer-based Time-domain Approach for Speaker Verification
标题:DT-SV:一种基于Transformer的说话人确认的时域方法
链接:https://arxiv.org/abs/2205.13249
作者:Nan Zhang,Jianzong Wang,Zhenhou Hong,Chendong Zhao,Xiaoyang Qu,Jing Xiao机构:Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)摘要:说话人验证(SV)旨在确定说话人在测试话语中的身份是否与参考语音相同。在过去几年中,使用深度神经网络提取SV系统中的说话人嵌入已成为主流。近年来,不同的注意机制和Transformer网络在SV领域得到了广泛的研究。然而,直接在SV中使用原始Transformer可能会在输出特性上产生帧级信息浪费,这可能会导致对扬声器嵌入的容量和辨别力的限制。因此,我们提出了一种通过Transformer结构来推导话语级说话人嵌入的方法,该结构使用一种新的损失函数diffluence loss来集成不同Transformer层的特征信息。其中,分流损失旨在将帧级特征聚合为话语级表示,并且可以方便地将其集成到转换器中。此外,我们还介绍了一种可学习的mel-fbank能量特征提取器,称为时域特征提取器,它比标准mel-fbank提取器更精确、更高效地计算mel-fbank特征。将分流损失和时域特征提取相结合,提出了一种新的基于Transformer的时域SV模型(DT-SV),具有更快的训练速度和更高的精度。实验表明,与其他模型相比,我们提出的模型可以获得更好的性能。摘要:Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embeddings using deep neural networks for SV systems has gone mainstream. Recently, different attention mechanisms and Transformer networks have been explored widely in SV fields. However, utilizing the original Transformer in SV directly may have frame-level information waste on output features, which could lead to restrictions on capacity and discrimination of speaker embeddings. Therefore, we propose an approach to derive utterance-level speaker embeddings via a Transformer architecture that uses a novel loss function named diffluence loss to integrate the feature information of different Transformer layers. Therein, the diffluence loss aims to aggregate frame-level features into an utterance-level representation, and it could be integrated into the Transformer expediently. Besides, we also introduce a learnable mel-fbank energy feature extractor named time-domain feature extractor that computes the mel-fbank features more precisely and efficiently than the standard mel-fbank extractor. Combining Diffluence loss and Time-domain feature extractor, we propose a novel Transformer-based time-domain SV model (DT-SV) with faster training speed and higher accuracy. Experiments indicate that our proposed model can achieve better performance in comparison with other models.
【2】 Urban Rhapsody: Large-scale exploration of urban soundscapes
标题:城市狂想曲:对城市声音景观的大规模探索
链接:https://arxiv.org/abs/2205.13064
作者:Joao Rulff,Fabio Miranda,Maryam Hosseini,Marcos Lage,Mark Cartwright,Graham Dove,Juan Bello,Claudio T. Silva机构:New York University,University of Illinois at Chicago,Universidade Federal Fluminense,New Jersey Institute of Technology, Broadway, NYC, Drilling, Large engine, Powered saw, Audio recordings, Heavy-construction model, Construction model, (c), (b)备注:Accepted at EuroVis 2022. Source code available at: this https URL摘要:噪声是城市环境中主要的生活质量问题之一。除了烦恼之外,噪音还对公众健康和教育绩效产生负面影响。虽然可以部署低成本传感器以高时间分辨率监测环境噪声水平,但它们产生的数据量和这些数据的复杂性构成了重大的分析挑战。解决这些挑战的一种方法是通过机器监听技术,该技术用于提取特征,试图对噪声源进行分类,并了解城市噪声状况的时间模式。然而,城市环境中大量的噪声源和标记数据的稀缺性使得几乎不可能创建具有足够大词汇表的分类模型,以捕捉城市声景观的真实动态。在本文中,我们首先确定了尚未开发的城市声景勘探领域中的一组要求。为了满足需求并应对已确定的挑战,我们提出了城市狂想曲,这是一个结合了最先进的音频表示、机器学习和视觉分析的框架,允许用户以交互方式创建分类模型,了解城市噪音模式,并快速检索和标记音频摘录,以创建一个大型高精度城市录音注释数据库。我们通过领域专家使用在纽约市部署一种独一无二的传感器网络五年期间生成的数据进行的案例研究,展示了该工具的实用性。摘要:Noise is one of the primary quality-of-life issues in urban environments. In addition to annoyance, noise negatively impacts public health and educational performance. While low-cost sensors can be deployed to monitor ambient noise levels at high temporal resolutions, the amount of data they produce and the complexity of these data pose significant analytical challenges. One way to address these challenges is through machine listening techniques, which are used to extract features in attempts to classify the source of noise and understand temporal patterns of a city's noise situation. However, the overwhelming number of noise sources in the urban environment and the scarcity of labeled data makes it nearly impossible to create classification models with large enough vocabularies that capture the true dynamism of urban soundscapes In this paper, we first identify a set of requirements in the yet unexplored domain of urban soundscape exploration. To satisfy the requirements and tackle the identified challenges, we propose Urban Rhapsody, a framework that combines state-of-the-art audio representation, machine learning, and visual analytics to allow users to interactively create classification models, understand noise patterns of a city, and quickly retrieve and label audio excerpts in order to create a large high-precision annotated database of urban sound recordings. We demonstrate the tool's utility through case studies performed by domain experts using data generated over the five-year deployment of a one-of-a-kind sensor network in New York City.
【3】 Joint Training of Speech Enhancement and Self-supervised Model for Noise-robust ASR
标题:语音增强和自监督模型联合训练的抗噪ASR
链接:https://arxiv.org/abs/2205.13293
作者:Qiu-Shi Zhu,Jie Zhang,Zi-Qiang Zhang,Li-Rong Dai机构: FundamentalResearch Funds for the Central Universities and the Leading Plan of CAS(XDC080 10 200), University of Sci-ence and Technology of China (USTC)备注:submitted to IEEE/ACM TASLP. arXiv admin note: text overlap with arXiv:2201.08930摘要:语音增强(SE)通常需要作为前端来改善噪声环境中的语音质量,而由于语音失真,增强后的语音可能不是自动语音识别(ASR)系统的最佳语音。另一方面,研究表明,自监督预训练能够利用大量未标记的噪声数据,这对ASR的噪声鲁棒性非常有利。然而,SE和自我监督的预训练(最佳)整合的潜力仍不清楚。为了找到一个合适的组合并减少SE引起的语音失真的影响,本文提出了一种SE模块和自监督模型的联合预训练方法。首先,在预训练阶段,将原始噪声波形或SE获得的波形输入到自监督模型中,学习上下文表示,其中量化的干净语音作为目标。其次,我们提出了一种双注意融合方法来融合带噪语音和增强语音的特征,可以补偿单独使用单个模块所造成的信息损失。由于灵活利用干净/有噪声/增强的分支,所提出的方法是一些现有抗噪声ASR模型的推广,例如增强的wav2vec2.0。最后,在合成和真实噪声数据集上的实验结果表明,所提出的联合训练方法可以在各种噪声环境下提高ASR性能,从而具有更强的噪声鲁棒性。摘要:Speech enhancement (SE) is usually required as a front end to improve the speech quality in noisy environments, while the enhanced speech might not be optimal for automatic speech recognition (ASR) systems due to speech distortion. On the other hand, it was shown that self-supervised pre-training enables the utilization of a large amount of unlabeled noisy data, which is rather beneficial for the noise robustness of ASR. However, the potential of the (optimal) integration of SE and self-supervised pre-training still remains unclear. In order to find an appropriate combination and reduce the impact of speech distortion caused by SE, in this paper we therefore propose a joint pre-training approach for the SE module and the self-supervised model. First, in the pre-training phase the original noisy waveform or the waveform obtained by SE is fed into the self-supervised model to learn the contextual representation, where the quantified clean speech acts as the target. Second, we propose a dual-attention fusion method to fuse the features of noisy and enhanced speeches, which can compensate the information loss caused by separately using individual modules. Due to the flexible exploitation of clean/noisy/enhanced branches, the proposed method turns out to be a generalization of some existing noise-robust ASR models, e.g., enhanced wav2vec2.0. Finally, experimental results on both synthetic and real noisy datasets show that the proposed joint training approach can improve the ASR performance under various noisy settings, leading to a stronger noise robustness.
2.eess.AS音频处理:
【1】 Joint Training of Speech Enhancement and Self-supervised Model for Noise-robust ASR
标题:语音增强和自监督模型联合训练的抗噪ASR
链接:https://arxiv.org/abs/2205.13293
作者:Qiu-Shi Zhu,Jie Zhang,Zi-Qiang Zhang,Li-Rong Dai机构: FundamentalResearch Funds for the Central Universities and the Leading Plan of CAS(XDC080 10 200), University of Sci-ence and Technology of China (USTC)备注:submitted to IEEE/ACM TASLP. arXiv admin note: text overlap with arXiv:2201.08930摘要:语音增强(SE)通常需要作为前端来改善噪声环境中的语音质量,而由于语音失真,增强后的语音可能不是自动语音识别(ASR)系统的最佳语音。另一方面,研究表明,自监督预训练能够利用大量未标记的噪声数据,这对ASR的噪声鲁棒性非常有利。然而,SE和自我监督的预训练(最佳)整合的潜力仍不清楚。为了找到一个合适的组合并减少SE引起的语音失真的影响,本文提出了一种SE模块和自监督模型的联合预训练方法。首先,在预训练阶段,将原始噪声波形或SE获得的波形输入到自监督模型中,学习上下文表示,其中量化的干净语音作为目标。其次,我们提出了一种双注意融合方法来融合带噪语音和增强语音的特征,可以补偿单独使用单个模块所造成的信息损失。由于灵活利用干净/有噪声/增强的分支,所提出的方法是一些现有抗噪声ASR模型的推广,例如增强的wav2vec2.0。最后,在合成和真实噪声数据集上的实验结果表明,所提出的联合训练方法可以在各种噪声环境下提高ASR性能,从而具有更强的噪声鲁棒性。摘要:Speech enhancement (SE) is usually required as a front end to improve the speech quality in noisy environments, while the enhanced speech might not be optimal for automatic speech recognition (ASR) systems due to speech distortion. On the other hand, it was shown that self-supervised pre-training enables the utilization of a large amount of unlabeled noisy data, which is rather beneficial for the noise robustness of ASR. However, the potential of the (optimal) integration of SE and self-supervised pre-training still remains unclear. In order to find an appropriate combination and reduce the impact of speech distortion caused by SE, in this paper we therefore propose a joint pre-training approach for the SE module and the self-supervised model. First, in the pre-training phase the original noisy waveform or the waveform obtained by SE is fed into the self-supervised model to learn the contextual representation, where the quantified clean speech acts as the target. Second, we propose a dual-attention fusion method to fuse the features of noisy and enhanced speeches, which can compensate the information loss caused by separately using individual modules. Due to the flexible exploitation of clean/noisy/enhanced branches, the proposed method turns out to be a generalization of some existing noise-robust ASR models, e.g., enhanced wav2vec2.0. Finally, experimental results on both synthetic and real noisy datasets show that the proposed joint training approach can improve the ASR performance under various noisy settings, leading to a stronger noise robustness.
【2】 Audio Data Augmentation for Acoustic-to-articulatory Speech Inversion using Bidirectional Gated RNNs
标题:基于双向门控RNN的声学发音反转音频数据增强
链接:https://arxiv.org/abs/2205.13086
作者:Yashish M. Siriwardena,Ahmed Adel Attia,Ganesh Sivaraman,Carol Espy-Wilson机构:University of Maryland College park, MD, USA, Pindrop, GA, USA摘要:通过增加训练数据的可变性,数据增强在改善深度学习模型的性能方面具有很好的前景。在之前开发抗噪声声学发音语音反演系统的工作中,我们已经证明了噪声增强对于提高含噪语音中语音反演性能的重要性。在这项工作中,我们比较和对比了不同的数据增强方法,并展示了这种技术如何提高发音反转的性能,不仅在噪声语音上,而且在干净的语音数据上。我们还提出了一种双向选通递归神经网络来代替以前使用的前馈神经网络作为语音反演系统。该反演系统使用mel倒谱系数(MFCC)作为输入声学特征,六个声道变量(TVs)作为输出发音特征。该系统的性能是通过计算美国Wisc上估计的电视与实际电视之间的相关性来衡量的。X射线微束数据库。对于干净的语音数据,与基线噪声鲁棒系统相比,所提出的语音反演系统的相关性相对提高了5%。当预训练的模型适应测试集中每个看不见的说话人时,平均相关度又提高了6%。摘要:Data augmentation has proven to be a promising prospect in improving the performance of deep learning models by adding variability to training data. In previous work with developing a noise robust acoustic-to-articulatory speech inversion system, we have shown the importance of noise augmentation to improve the performance of speech inversion in noisy speech. In this work, we compare and contrast different ways of doing data augmentation and show how this technique improves the performance of articulatory speech inversion not only on noisy speech, but also on clean speech data. We also propose a Bidirectional Gated Recurrent Neural Network as the speech inversion system instead of the previously used feed forward neural network. The inversion system uses mel-frequency cepstral coefficients (MFCCs) as the input acoustic features and six vocal tract-variables (TVs) as the output articulatory features. The Performance of the system was measured by computing the correlation between estimated and actual TVs on the U. Wisc. X-ray Microbeam database. The proposed speech inversion system shows a 5% relative improvement in correlation over the baseline noise robust system for clean speech data. The pre-trained model, when adapted to each unseen speaker in the test set, improves the average correlation by another 6%.
【3】 DT-SV: A Transformer-based Time-domain Approach for Speaker Verification
标题:DT-SV:一种基于Transformer的说话人确认的时域方法
链接:https://arxiv.org/abs/2205.13249
作者:Nan Zhang,Jianzong Wang,Zhenhou Hong,Chendong Zhao,Xiaoyang Qu,Jing Xiao机构:Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China备注:Accepted by IJCNN2022 (The 2022 International Joint Conference on Neural Networks)摘要:说话人验证(SV)旨在确定说话人在测试话语中的身份是否与参考语音相同。在过去几年中,使用深度神经网络提取SV系统中的说话人嵌入已成为主流。近年来,不同的注意机制和Transformer网络在SV领域得到了广泛的研究。然而,直接在SV中使用原始Transformer可能会在输出特性上产生帧级信息浪费,这可能会导致对扬声器嵌入的容量和辨别力的限制。因此,我们提出了一种通过Transformer结构来推导话语级说话人嵌入的方法,该结构使用一种新的损失函数diffluence loss来集成不同Transformer层的特征信息。其中,分流损失旨在将帧级特征聚合为话语级表示,并且可以方便地将其集成到转换器中。此外,我们还介绍了一种可学习的mel-fbank能量特征提取器,称为时域特征提取器,它比标准mel-fbank提取器更精确、更高效地计算mel-fbank特征。将分流损失和时域特征提取相结合,提出了一种新的基于Transformer的时域SV模型(DT-SV),具有更快的训练速度和更高的精度。实验表明,与其他模型相比,我们提出的模型可以获得更好的性能。摘要:Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embeddings using deep neural networks for SV systems has gone mainstream. Recently, different attention mechanisms and Transformer networks have been explored widely in SV fields. However, utilizing the original Transformer in SV directly may have frame-level information waste on output features, which could lead to restrictions on capacity and discrimination of speaker embeddings. Therefore, we propose an approach to derive utterance-level speaker embeddings via a Transformer architecture that uses a novel loss function named diffluence loss to integrate the feature information of different Transformer layers. Therein, the diffluence loss aims to aggregate frame-level features into an utterance-level representation, and it could be integrated into the Transformer expediently. Besides, we also introduce a learnable mel-fbank energy feature extractor named time-domain feature extractor that computes the mel-fbank features more precisely and efficiently than the standard mel-fbank extractor. Combining Diffluence loss and Time-domain feature extractor, we propose a novel Transformer-based time-domain SV model (DT-SV) with faster training speed and higher accuracy. Experiments indicate that our proposed model can achieve better performance in comparison with other models.
【4】 Urban Rhapsody: Large-scale exploration of urban soundscapes
标题:城市狂想曲:对城市声音景观的大规模探索
链接:https://arxiv.org/abs/2205.13064
作者:Joao Rulff,Fabio Miranda,Maryam Hosseini,Marcos Lage,Mark Cartwright,Graham Dove,Juan Bello,Claudio T. Silva机构:New York University,University of Illinois at Chicago,Universidade Federal Fluminense,New Jersey Institute of Technology, Broadway, NYC, Drilling, Large engine, Powered saw, Audio recordings, Heavy-construction model, Construction model, (c), (b)备注:Accepted at EuroVis 2022. Source code available at: this https URL摘要:噪声是城市环境中主要的生活质量问题之一。除了烦恼之外,噪音还对公众健康和教育绩效产生负面影响。虽然可以部署低成本传感器以高时间分辨率监测环境噪声水平,但它们产生的数据量和这些数据的复杂性构成了重大的分析挑战。解决这些挑战的一种方法是通过机器监听技术,该技术用于提取特征,试图对噪声源进行分类,并了解城市噪声状况的时间模式。然而,城市环境中大量的噪声源和标记数据的稀缺性使得几乎不可能创建具有足够大词汇表的分类模型,以捕捉城市声景观的真实动态。在本文中,我们首先确定了尚未开发的城市声景勘探领域中的一组要求。为了满足需求并应对已确定的挑战,我们提出了城市狂想曲,这是一个结合了最先进的音频表示、机器学习和视觉分析的框架,允许用户以交互方式创建分类模型,了解城市噪音模式,并快速检索和标记音频摘录,以创建一个大型高精度城市录音注释数据库。我们通过领域专家使用在纽约市部署一种独一无二的传感器网络五年期间生成的数据进行的案例研究,展示了该工具的实用性。摘要:Noise is one of the primary quality-of-life issues in urban environments. In addition to annoyance, noise negatively impacts public health and educational performance. While low-cost sensors can be deployed to monitor ambient noise levels at high temporal resolutions, the amount of data they produce and the complexity of these data pose significant analytical challenges. One way to address these challenges is through machine listening techniques, which are used to extract features in attempts to classify the source of noise and understand temporal patterns of a city's noise situation. However, the overwhelming number of noise sources in the urban environment and the scarcity of labeled data makes it nearly impossible to create classification models with large enough vocabularies that capture the true dynamism of urban soundscapes In this paper, we first identify a set of requirements in the yet unexplored domain of urban soundscape exploration. To satisfy the requirements and tackle the identified challenges, we propose Urban Rhapsody, a framework that combines state-of-the-art audio representation, machine learning, and visual analytics to allow users to interactively create classification models, understand noise patterns of a city, and quickly retrieve and label audio excerpts in order to create a large high-precision annotated database of urban sound recordings. We demonstrate the tool's utility through case studies performed by domain experts using data generated over the five-year deployment of a one-of-a-kind sensor network in New York City.
机器翻译,仅供参考