cs.SD语音,共计2篇,eess.AS音频处理,共计5篇
1.cs.SD语音:
【1】 MIMII DG: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection for Domain Generalization Task
标题:MIMII DG:工业机械故障调查和检查领域概化任务的声音数据集
链接:https://arxiv.org/abs/2205.13879
作者:Kota Dohi,Tomoya Nishida,Harsh Purohit,Ryo Tanabe,Takashi Endo,Masaaki Yamamoto,Yuki Nikaido,Yohei Kawaguchi机构:Research and Development Group, Hitachi, Ltd., -, Higashi-koigakubo, Kokubunji, Tokyo ,-, Japan摘要:我们提出了一个机器声音数据集来测试异常声音检测(ASD)的领域泛化技术。为了处理由于难以检测或过于频繁的域转移而导致的性能下降,首选域泛化技术。然而,目前可用的数据集很难评估这些技术,例如,导致域移动的参数(域移动参数)的值数量有限。在本文中,我们提出了第一个用于领域泛化技术的ASD数据集,称为MIMII DG。该数据集由五种机器类型和每种机器类型的三种域转移场景组成。我们为源域中的域移位参数准备了至少两个值。此外,我们还引入了很难注意到的域转移。使用两个基线系统的实验结果表明,该数据集再现了域转移场景,有助于对域泛化技术进行基准测试。摘要:We present a machine sound dataset to benchmark domain generalization techniques for anomalous sound detection (ASD). To handle performance degradation caused by domain shifts that are difficult to detect or too frequent to adapt, domain generalization techniques are preferred. However, currently available datasets have difficulties in evaluating these techniques, such as limited number of values for parameters that cause domain shifts (domain shift parameters). In this paper, we present the first ASD dataset for the domain generalization techniques, called MIMII DG. The dataset consists of five machine types and three domain shift scenarios for each machine type. We prepared at least two values for the domain shift parameters in the source domain. Also, we introduced domain shifts that can be difficult to notice. Experimental results using two baseline systems indicate that the dataset reproduces the domain shift scenarios and is useful for benchmarking domain generalization techniques.
【2】 Adversarial attacks and defenses in Speaker Recognition Systems: A survey
标题:说话人识别系统中的对抗性攻击与防御
链接:https://arxiv.org/abs/2205.13685
作者:Jiahe Lan,Rui Zhang,Zheng Yan,Jie Wang,Yu Chen,Ronghui Hou机构:State Key Laboratory on Integrated Services Networks, School of Cyber Engineering, Xidian University, China, Department of Communications and Networking, Aalto University, Finland备注:38pages, 2 figures, 2 tables. Journal of Systems Architecture,2022摘要:由于远程控制的易用性和经济友好的特性,说话人识别在智能家居和智能助理等许多应用场景中变得非常流行。SRSs的快速发展离不开机器学习特别是神经网络的发展。然而,之前的研究表明,机器学习模型在图像领域容易受到对抗性攻击,这激发了研究人员探索说话人识别系统(SRS)中的对抗性攻击和防御。遗憾的是,现有文献缺乏对这一主题的全面回顾。在本文中,我们通过对SRS中的对抗性攻击和防御进行全面调查来填补这一空白。我们首先介绍SRS的基本知识和与对抗性攻击相关的概念。然后,我们分别提出了两套标准来评估SRS中攻击方法和防御方法的性能。然后,我们提供了现有攻击方法和防御方法的分类,并使用我们提出的标准对它们进行了进一步的审查。最后,基于我们的回顾,我们发现了一些尚未解决的问题,并进一步指明了一些未来的方向,以推动SRSs安全的研究。摘要:Speaker recognition has become very popular in many application scenarios, such as smart homes and smart assistants, due to ease of use for remote control and economic-friendly features. The rapid development of SRSs is inseparable from the advancement of machine learning, especially neural networks. However, previous work has shown that machine learning models are vulnerable to adversarial attacks in the image domain, which inspired researchers to explore adversarial attacks and defenses in Speaker Recognition Systems (SRS). Unfortunately, existing literature lacks a thorough review of this topic. In this paper, we fill this gap by performing a comprehensive survey on adversarial attacks and defenses in SRSs. We first introduce the basics of SRSs and concepts related to adversarial attacks. Then, we propose two sets of criteria to evaluate the performance of attack methods and defense methods in SRSs, respectively. After that, we provide taxonomies of existing attack methods and defense methods, and further review them by employing our proposed criteria. Finally, based on our review, we find some open issues and further specify a number of future directions to motivate the research of SRSs security.
2.eess.AS音频处理:
【1】 Speaker-conditioning Single-channel Target Speaker Extraction using Conformer-based Architectures
标题:基于一致性结构的说话人条件单声道目标说话人提取
链接:https://arxiv.org/abs/2205.13851
作者:Ragini Sinha,Marvin Tammen,Christian Rollwage,Simon Doclo机构:Fraunhofer Institute for Digital Media Technology IDMT, Oldenburg Branch for Hearing, Speech and Audio Technology HSA, Germany, Department of Medical Physics and Acoustics and Cluster of Excellence Hearing,all, University of Oldenburg, Germany备注:submitted to IWAENC 2022摘要:目标说话人提取旨在利用目标说话人的辅助信息,从多个说话人的混合中提取目标说话人。在本文中,我们考虑一个完整的时域目标说话人提取系统,该系统由说话人嵌入网络和说话人分离网络组成,这两个网络在端到端的学习过程中联合训练。我们提出了两种不同的基于卷积增强变换器(conformer)的说话人分离网络结构。第一种架构使用一致性块和外部前馈块的堆栈(一致性FFN),而第二种架构使用时间卷积网络(TCN)和一致性块的堆栈(TCN一致性)。对2个说话人混合、3个说话人混合和2个说话人的噪声混合的实验结果表明,与Conformer FFN和基于TCN的基线系统相比,TCN Conformer显著提高了目标说话人的提取性能。摘要:Target speaker extraction aims at extracting the target speaker from a mixture of multiple speakers exploiting auxiliary information about the target speaker. In this paper, we consider a complete time-domain target speaker extraction system consisting of a speaker embedder network and a speaker separator network which are jointly trained in an end-to-end learning process. We propose two different architectures for the speaker separator network which are based on the convolutional augmented transformer (conformer). The first architecture uses stacks of conformer and external feed-forward blocks (Conformer-FFN), while the second architecture uses stacks of temporal convolutional network (TCN) and conformer blocks (TCN-Conformer). Experimental results for 2-speaker mixtures, 3-speaker mixtures, and noisy mixtures of 2-speakers show that among the proposed separator networks, the TCN-Conformer significantly improves the target speaker extraction performance compared to the Conformer-FFN and a TCN-based baseline system.
【2】 Acoustic-to-articulatory Speech Inversion with Multi-task Learning
标题:基于多任务学习的声-清语音反转
链接:https://arxiv.org/abs/2205.13755
作者:Yashish M. Siriwardena,Ganesh Sivaraman,Carol Espy-Wilson机构:University of Maryland College park, MD, USA, Pindrop, GA, USA摘要:多任务学习(MTL)框架在自动语音识别(ASR)和语音情感识别等不同的语音相关任务中被证明是有效的。本文提出了一个MTL框架,通过同时学习声学到音素的映射作为一个共享任务来执行声学到发音语音的反转。我们使用Haskins生产率比较(HPRC)数据库,该数据库包含电磁关节造影(EMA)数据和相应的语音转录。该系统的性能是通过计算从声学到发音语音转换任务的估计和实际声道变量(TV)之间的相关性来衡量的。所提出的基于MTL的双向选通递归神经网络(RNN)模型学习将输入声学特征映射到九个TV,同时优于仅进行声学到发音反转的基线模型。摘要:Multi-task learning (MTL) frameworks have proven to be effective in diverse speech related tasks like automatic speech recognition (ASR) and speech emotion recognition. This paper proposes a MTL framework to perform acoustic-to-articulatory speech inversion by simultaneously learning an acoustic to phoneme mapping as a shared task. We use the Haskins Production Rate Comparison (HPRC) database which has both the electromagnetic articulography (EMA) data and the corresponding phonetic transcriptions. Performance of the system was measured by computing the correlation between estimated and actual tract variables (TVs) from the acoustic to articulatory speech inversion task. The proposed MTL based Bidirectional Gated Recurrent Neural Network (RNN) model learns to map the input acoustic features to nine TVs while outperforming the baseline model trained to perform only acoustic to articulatory inversion.
【3】 An enhanced Conv-TasNet model for speech separation using a speaker distance-based loss function
标题:基于说话人距离损失函数的改进Conv-TasNet语音分离模型
链接:https://arxiv.org/abs/2205.13657
作者:Jose A. Arango-Sánchez,Julián D. Arias-Londoño机构:Intelligent Information Systems Lab - In,Lab, Universidad de Antioquia, Medellín, Colombia., GAPS-SSR, ETSIT-Universidad Politécnica de Madrid, Madrid, Spain.摘要:这项工作使用预先训练好的深度学习模型来解决西班牙语中的语音分离问题。与许多语音处理任务一样,使用不同于英语的其他语言的大型数据库也很少。因此,本研究以Conv-TasNet模型为基准,探索不同的训练策略。对于最佳的训练策略,实现了9.9 dB的尺度不变信号失真比(SI-SDR)度量值。然后,通过实验,我们发现说话人相似度与模型性能之间存在反比关系,因此提出了一种改进的ConvTasNet体系结构。增强的Conv-TasNet模型使用预先训练的语音嵌入在代价函数中添加说话人之间的余弦相似项,从而产生10.6 dB的SI-SDR。最后,介绍了有关模型实时部署的一些见解和缺点。摘要:This work addresses the problem of speech separation in the Spanish Language using pre-trained deep learning models. As with many speech processing tasks, large databases in other languages different from English are scarce. Therefore this work explores different training strategies using the Conv-TasNet model as a benchmark. A scale-invariant signal distortion ratio (SI-SDR) metric value of 9.9 dB was achieved for the best training strategy. Then, experimentally, we identified an inverse relationship between the speakers' similarity and the model's performance, so an improved ConvTasNet architecture was proposed. The enhanced Conv-TasNet model uses pre-trained speech embeddings to add a between-speakers cosine similarity term in the cost function, yielding an SI-SDR of 10.6 dB. Lastly, some insights and drawbacks about the real-time deployment of the model.
【4】 MIMII DG: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection for Domain Generalization Task
标题:MIMII DG:工业机械故障调查和检查领域概化任务的声音数据集
链接:https://arxiv.org/abs/2205.13879
作者:Kota Dohi,Tomoya Nishida,Harsh Purohit,Ryo Tanabe,Takashi Endo,Masaaki Yamamoto,Yuki Nikaido,Yohei Kawaguchi机构:Research and Development Group, Hitachi, Ltd., -, Higashi-koigakubo, Kokubunji, Tokyo ,-, Japan摘要:我们提出了一个机器声音数据集来测试异常声音检测(ASD)的领域泛化技术。为了处理由于难以检测或过于频繁的域转移而导致的性能下降,首选域泛化技术。然而,目前可用的数据集很难评估这些技术,例如,导致域移动的参数(域移动参数)的值数量有限。在本文中,我们提出了第一个用于领域泛化技术的ASD数据集,称为MIMII DG。该数据集由五种机器类型和每种机器类型的三种域转移场景组成。我们为源域中的域移位参数准备了至少两个值。此外,我们还引入了很难注意到的域转移。使用两个基线系统的实验结果表明,该数据集再现了域转移场景,有助于对域泛化技术进行基准测试。摘要:We present a machine sound dataset to benchmark domain generalization techniques for anomalous sound detection (ASD). To handle performance degradation caused by domain shifts that are difficult to detect or too frequent to adapt, domain generalization techniques are preferred. However, currently available datasets have difficulties in evaluating these techniques, such as limited number of values for parameters that cause domain shifts (domain shift parameters). In this paper, we present the first ASD dataset for the domain generalization techniques, called MIMII DG. The dataset consists of five machine types and three domain shift scenarios for each machine type. We prepared at least two values for the domain shift parameters in the source domain. Also, we introduced domain shifts that can be difficult to notice. Experimental results using two baseline systems indicate that the dataset reproduces the domain shift scenarios and is useful for benchmarking domain generalization techniques.
【5】 Adversarial attacks and defenses in Speaker Recognition Systems: A survey
标题:说话人识别系统中的对抗性攻击与防御
链接:https://arxiv.org/abs/2205.13685
作者:Jiahe Lan,Rui Zhang,Zheng Yan,Jie Wang,Yu Chen,Ronghui Hou机构:State Key Laboratory on Integrated Services Networks, School of Cyber Engineering, Xidian University, China, Department of Communications and Networking, Aalto University, Finland备注:38pages, 2 figures, 2 tables. Journal of Systems Architecture,2022摘要:由于远程控制的易用性和经济友好的特性,说话人识别在智能家居和智能助理等许多应用场景中变得非常流行。SRSs的快速发展离不开机器学习特别是神经网络的发展。然而,之前的研究表明,机器学习模型在图像领域容易受到对抗性攻击,这激发了研究人员探索说话人识别系统(SRS)中的对抗性攻击和防御。遗憾的是,现有文献缺乏对这一主题的全面回顾。在本文中,我们通过对SRS中的对抗性攻击和防御进行全面调查来填补这一空白。我们首先介绍SRS的基本知识和与对抗性攻击相关的概念。然后,我们分别提出了两套标准来评估SRS中攻击方法和防御方法的性能。然后,我们提供了现有攻击方法和防御方法的分类,并使用我们提出的标准对它们进行了进一步的审查。最后,基于我们的回顾,我们发现了一些尚未解决的问题,并进一步指明了一些未来的方向,以推动SRSs安全的研究。摘要:Speaker recognition has become very popular in many application scenarios, such as smart homes and smart assistants, due to ease of use for remote control and economic-friendly features. The rapid development of SRSs is inseparable from the advancement of machine learning, especially neural networks. However, previous work has shown that machine learning models are vulnerable to adversarial attacks in the image domain, which inspired researchers to explore adversarial attacks and defenses in Speaker Recognition Systems (SRS). Unfortunately, existing literature lacks a thorough review of this topic. In this paper, we fill this gap by performing a comprehensive survey on adversarial attacks and defenses in SRSs. We first introduce the basics of SRSs and concepts related to adversarial attacks. Then, we propose two sets of criteria to evaluate the performance of attack methods and defense methods in SRSs, respectively. After that, we provide taxonomies of existing attack methods and defense methods, and further review them by employing our proposed criteria. Finally, based on our review, we find some open issues and further specify a number of future directions to motivate the research of SRSs security.
机器翻译,仅供参考