今天跟大家分享一篇语音相关的论文合集:cs.SD语音10篇,eess.AS音频处理10篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 Streaming Intended Query Detection using E2E Modeling for Continued Conversation
标题:基于E2E建模的连续会话流目标查询检测
链接:https://arxiv.org/abs/2208.13322
作者:Shuo-yiin Chang,Guru Prakash,Zelin Wu,Qiao Liang,Tara N. Sainath,Bo Li,Adam Stambler,Shyam Upadhyay,Manaal Faruqui,Trevor Strohman备注:5 pages, Interspeech 2022摘要:在支持语音的应用中,为了关注www.example.com,通常使用预先确定的热门词汇来激活设备,query.However每次说出查询后都跟随一个热门词汇,这会给后续会话带来认知负担。为了避免重复热点词,我们提出了一种流媒体端到端(E2E)意向查询检测器,它识别指向设备的话语,并过滤掉其他非指向设备的话语。该方法将意图查询检测器引入到E2E模型中,将语音识别流水线的不同部分合并为一个神经网络www.example.comE2Enetwork.The模型中,并基于早期的部分识别结果快速地进行意图查询检测,这对于减少系统的延迟和提高系统的响应速度是非常重要的。实验结果表明,与独立的意图查询检测器相比,E2E方法在等错误率(EER)上的检测准确率提高了22%,延迟时间提高了600ms。在我们的实验中,所提出的模型在用户开始说话后的1.4秒的中位延迟内检测用户是否以8.7%的EER与设备说话。摘要:In voice-enabled applications, a predetermined hotword isusually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each timeintroduces a cognitive burden in continued conversations. Toavoid repeating a hotword, we propose a streaming end-to-end(E2E) intended query detector that identifies the utterancesdirected towards the device and filters out other utterancesnot directed towards device. The proposed approach incor-porates the intended query detector into the E2E model thatalready folds different components of the speech recognitionpipeline into one neural network.The E2E modeling onspeech decoding and intended query detection also allows us todeclare a quick intended query detection based on early partialrecognition result, which is important to decrease latencyand make the system responsive. We demonstrate that theproposed E2E approach yields a 22% relative improvement onequal error rate (EER) for the detection accuracy and 600 mslatency improvement compared with an independent intendedquery detector. In our experiment, the proposed model detectswhether the user is talking to the device with a 8.7% EERwithin 1.4 seconds of median latency after user starts speaking.
【2】 Turn-Taking Prediction for Natural Conversational Speech
标题:自然会话语音的话轮转换预测
链接:https://arxiv.org/abs/2208.13321
作者:Shuo-yiin Chang,Bo Li,Tara N. Sainath,Chao Zhang,Trevor Strohman,Qiao Liang,Yanzhang He机构:Google Inc., U.S.A备注:5 pages, Interspeech 2022摘要:虽然流式语音助理系统已经在许多应用中使用,但是该系统典型地集中于假定来自单个语音查询的输入没有犹豫或不流畅的不自然的一次性交互。然而,一个普通的会话话语除了不流利之外,还经常涉及带有话轮转换的多个询问。这些不流利包括停顿思考、犹豫、单词加长、填充停顿和重复短语。这使得对会话语音(包括具有多个查询的会话语音)进行语音识别成为一项具有挑战性的任务。为了更好地对会话交互进行建模,关键的是区分不流利和查询结束,以便允许用户针对不流利保持发言权,同时在用户已经结束讲话时使系统尽可能快地做出响应。本文提出了一种建立在端到端语音识别器之上的话轮转换预测器。我们的最佳系统是通过联合优化ASR任务和检测用户何时暂停思考或结束讲话而获得的。实验结果表明,该方法在预测真实话轮转换时的召回率和准确率分别达到97%和85%以上,预测延迟仅为100 ms。摘要:While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice query without hesitation or disfluency. However, a common conversational utterance often involves multiple queries with turn-taking, in addition to disfluencies. These disfluencies include pausing to think, hesitations, word lengthening, filled pauses and repeated phrases. This makes doing speech recognition with conversational speech, including one with multiple queries, a challenging task. To better model the conversational interaction, it is critical to discriminate disfluencies and end of query in order to allow the user to hold the floor for disfluencies while having the system respond as quickly as possible when the user has finished speaking. In this paper, we present a turntaking predictor built on top of the end-to-end (E2E) speech recognizer. Our best system is obtained by jointly optimizing for ASR task and detecting when the user is paused to think or finished speaking. The proposed approach demonstrates over 97% recall rate and 85% precision rate on predicting true turn-taking with only 100 ms latency on a test set designed with 4 types of disfluencies inserted in conversational utterances.
【3】 Computing with Hypervectors for Efficient Speaker Identification
标题:高效说话人辨认的超向量计算
链接:https://arxiv.org/abs/2208.13285
作者:Ping-Chen Huang,Denis Kleyko,Jan M. Rabaey,Bruno A. Olshausen,Pentti Kanerva机构:Redwood Center of Theoretical Neuroscience, University of California, Berkeley, USA, Intelligent Systems Lab, Research Institutes of Sweden, Sweden, Berkeley Wireless Research Center, University of California, Berkeley, USA摘要:本文提出了一种基于高维随机向量计算的说话人识别方法。它的优点是简单和快速。仅使用www.example.com的1.02k活动参数和128分钟的训练数据,我们在1,251个说话者的VoxCeleb1数据集上获得了31%和52%的前1和前5得分。这与CNN模型形成对比,CNN模型需要数百万个参数和数量级更高的计算复杂性,而在互信息中测量的区分能力中仅获得2 × $的增益。使用广义学习矢量量化(GLVQ)的额外92秒训练将分数提高到48%和67%。经过训练的分类器在5.7ms内对1秒的语音进行分类。所有处理都在基于标准CPU的机器上完成。摘要:We introduce a method to identify speakers by computing with high-dimensional random vectors. Its strengths are simplicity and speed. With only 1.02k active parameters and a 128-minute pass through the training data we achieve Top-1 and Top-5 scores of 31% and 52% on the VoxCeleb1 dataset of 1,251 speakers. This is in contrast to CNN models requiring several million parameters and orders of magnitude higher computational complexity for only a 2$\times$ gain in discriminative power as measured in mutual information. An additional 92 seconds of training with Generalized Learning Vector Quantization (GLVQ) raises the scores to 48% and 67%. A trained classifier classifies 1 second of speech in 5.7 ms. All processing was done on standard CPU-based machines.
【4】 Towards Disentangled Speech Representations
标题:走向无纠缠的言语表征
链接:https://arxiv.org/abs/2208.13191
作者:Cal Peyser,Ronny Huang Andrew Rosenberg Tara N. Sainath,Michael Picheny,Kyunghyun Cho机构:Cho, Center for Data Science, New York University, New York City, USA, Google Inc., U.S.A摘要:在许多言语任务的方法设计中,音频表征的仔细构建已经成为一个主要特征。这种方法越来越强调“解纠缠,”其中表示只包含语音信号中与转录相关的部分,而丢弃不相关的信息.本文基于ASR和TTS的联合建模,构建了一个表征学习任务,并试图学习一种将语音信号中与转录相关的部分与不相关的部分分开的音频表征。我们提出的经验证据表明,成功地找到这样一个代表是绑在训练中固有的随机性。然后,我们观察到这些期望的、解纠缠的优化问题的解具有独特的统计特性。最后,我们表明,在训练过程中实施这些属性,相对于我们的联合建模任务,平均提高了24.5%的WER。这些观察激发了学习有效音频表示的新颖方法。摘要:The careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks. Increasingly, such approaches have emphasized "disentanglement", where a representation contains only parts of the speech signal relevant to transcription while discarding irrelevant information. In this paper, we construct a representation learning task based on joint modeling of ASR and TTS, and seek to learn a representation of audio that disentangles that part of the speech signal that is relevant to transcription from that part which is not. We present empirical evidence that successfully finding such a representation is tied to the randomness inherent in training. We then make the observation that these desired, disentangled solutions to the optimization problem possess unique statistical properties. Finally, we show that enforcing these properties during training improves WER by 24.5% relative on average for our joint modeling task. These observations motivate a novel approach to learning effective audio representations.
【5】 Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks
标题:从合成数据训练文本到语音系统:一种实用的口音转移任务方法
链接:https://arxiv.org/abs/2208.13183
作者:Lev Finkelstein,Heiga Zen,Norman Casagrande,Chun-an Chan,Ye Jia,Tom Kenter,Alexey Petelin,Jonathan Shen,Vincent Wan,Yu Zhang,Yonghui Wu,Rob Clark机构:Google LLC, †DeepMind备注:To be published in Interspeech 2022摘要:文本到语音(TTS)合成中的传输任务——其中一组说话者的语音的一个或多个方面被传输到最初不具有这些方面的另一组说话者——仍然是一项具有挑战性的任务。其中一个挑战是,具有高质量传输能力的模型可能存在稳定性问题,这使得它们不适用于面向用户的关键任务。本文证明,通过训练鲁棒TTS系统,可以获得高质量的传输任务,数据由设计用于高质量传输任务的鲁棒性较差的TTS系统生成;特别地,在为重音转移设计的Tacotron模型的输出上训练CHiVE-BERT单语TTS系统。虽然用这种方法不可避免地会有一些质量损失,但实验结果表明,用这种方法在合成数据上训练的模型可以产生显示重音转移的高质量音频,同时保留说话者的特征,例如说话风格。摘要:Transfer tasks in text-to-speech (TTS) synthesis - where one or more aspects of the speech of one set of speakers is transferred to another set of speakers that do not feature these aspects originally - remains a challenging task. One of the challenges is that models that have high-quality transfer capabilities can have issues in stability, making them impractical for user-facing critical tasks. This paper demonstrates that transfer can be obtained by training a robust TTS system on data generated by a less robust TTS system designed for a high-quality transfer task; in particular, a CHiVE-BERT monolingual TTS system is trained on the output of a Tacotron model designed for accent transfer. While some quality loss is inevitable with this approach, experimental results show that the models trained on synthetic data this way can produce high quality audio displaying accent transfer, while preserving speaker characteristics such as speaking style.
【6】 SA: Sliding attack for synthetic speech detection with resistance to clipping and self-splicing
标题:SA:抗剪裁和自拼接的合成语音检测滑动攻击
链接:https://arxiv.org/abs/2208.13066
作者:Deng JiaCheng,Dong Li,Yan Diqun,Wang Rangding,Zeng Jiaming机构:Department of Computer Science, Ningbo University, Ningbo , China, A R T I C L E I N F O备注:12 pages, Neurocomputing摘要:深度神经网络容易受到敌对例子的影响,这些例子会用不可察觉的扰动误导模型。在音频中,尽管对抗性攻击的例子在白盒设置和黑盒设置上取得了令人难以置信的攻击成功率,但大多数现有的对抗性攻击受到输入长度的限制。一个更实际的场景是对抗性的例子必须被剪切或自拼接并输入到黑盒模型中。因此,有必要探索如何在不同的输入长度设置下提高可转移性。本文以合成语音检测任务为例,考虑了两种有代表性的SOTA模型。通过对剪切或自拼接后样本输入模型得到的梯度进行分析,发现相同样本值的片段在不同模型中的梯度是相似的。受此启发,我们提出了一种新的对抗性攻击方法——滑动攻击。具体地说,我们使每个采样点知道不同位置的梯度,这可以模拟具有不同输入长度的敌对样本被输入到黑盒模型的情况。因此,代替在梯度计算的每次迭代中直接使用当前梯度,我们经历以下三个步骤。首先,我们使用滑动窗口提取不同长度的子段。然后,我们用来自相邻域的数据扩充子段。最后,我们将子片段馈入不同的模型中以获得聚合梯度来更新对抗范例。实验结果表明,该方法能显著提高剪切或自拼接后的对抗性实例的可移植性。此外,该方法还可以增强基于不同特征的模型之间的可移植性。摘要:Deep neural networks are vulnerable to adversarial examples that mislead models with imperceptible perturbations. In audio, although adversarial examples have achieved incredible attack success rates on white-box settings and black-box settings, most existing adversarial attacks are constrained by the input length. A More practical scenario is that the adversarial examples must be clipped or self-spliced and input into the black-box model. Therefore, it is necessary to explore how to improve transferability in different input length settings. In this paper, we take the synthetic speech detection task as an example and consider two representative SOTA models. We observe that the gradients of fragments with the same sample value are similar in different models via analyzing the gradients obtained by feeding samples into the model after cropping or self-splicing. Inspired by the above observation, we propose a new adversarial attack method termed sliding attack. Specifically, we make each sampling point aware of gradients at different locations, which can simulate the situation where adversarial examples are input to black-box models with varying input lengths. Therefore, instead of using the current gradient directly in each iteration of the gradient calculation, we go through the following three steps. First, we extract subsegments of different lengths using sliding windows. We then augment the subsegments with data from the adjacent domains. Finally, we feed the sub-segments into different models to obtain aggregate gradients to update adversarial examples. Empirical results demonstrate that our method could significantly improve the transferability of adversarial examples after clipping or self-splicing. Besides, our method could also enhance the transferability between models based on different features.
【7】 Sub-mW Neuromorphic SNN audio processing applications with Rockpool and Xylo
标题:采用Rockpool和Xylo的亚毫瓦神经形态SNN音频处理应用
链接:https://arxiv.org/abs/2208.12991
摘要:尖峰神经网络(SNN)为时间信号处理提供了一种高效的计算机制,尤其是在与低功耗SNN推理ASIC耦合时。SNN历来难以配置,缺乏为任意任务找到解决方案的通用方法。近年来,梯度下降优化方法已经越来越容易地应用于SNNs。因此,SNN和SNN推理处理器提供了一个良好的平台,用于在没有云依赖性的能量受限环境中进行商业低功耗信号处理。然而,到目前为止,这些方法还没有被工业中的ML工程师所使用,需要研究生水平的培训才能成功地配置单个SNN应用。在这里,我们展示了一个方便的高级流水线,用于设计、训练和部署任意时间信号处理应用到亚mW SNN推理硬件。我们采用一种新的简单的SNN结构,使用突触时间常数金字塔来提取时间尺度范围内的信号特征。我们在一个环境音频分类任务上演示了该架构,该任务以流模式部署到Xylo SNN推理处理器。我们的应用在低功耗(〈4 μ W推理功率)下实现了高准确度(98%)和低延迟(100ms)。我们的方法使得具有一般神经网络背景的ML工程师可以培训和部署SNN应用,而不需要具有尖峰神经网络的特定经验。我们希望我们的方法使神经形态硬件和SNN成为商业低功耗和边缘信号处理应用的有吸引力的选择。摘要:Spiking Neural Networks (SNNs) provide an efficient computational mechanism for temporal signal processing, especially when coupled with low-power SNN inference ASICs. SNNs have been historically difficult to configure, lacking a general method for finding solutions for arbitrary tasks. In recent years, gradient-descent optimization methods have been applied to SNNs with increasing ease. SNNs and SNN inference processors therefore offer a good platform for commercial low-power signal processing in energy constrained environments without cloud dependencies. However, to date these methods have not been accessible to ML engineers in industry, requiring graduate-level training to successfully configure a single SNN application. Here we demonstrate a convenient high-level pipeline to design, train and deploy arbitrary temporal signal processing applications to sub-mW SNN inference hardware. We apply a new straightforward SNN architecture designed for temporal signal processing, using a pyramid of synaptic time constants to extract signal features at a range of temporal scales. We demonstrate this architecture on an ambient audio classification task, deployed to the Xylo SNN inference processor in streaming mode. Our application achieves high accuracy (98%) and low latency (100ms) at low power (<4muW inference power). Our approach makes training and deploying SNN applications available to ML engineers with general NN backgrounds, without requiring specific prior experience with spiking NNs. We intend for our approach to make Neuromorphic hardware and SNNs an attractive choice for commercial low-power and edge signal processing applications.
【8】 Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation
标题:跨语言低资源ASR评估的数据分割策略研究
链接:https://arxiv.org/abs/2208.12888
作者:Zoey Liu,Justin Spence,Emily Prud'hommeaux机构:Boston College, University of California, Davis, Emily Prud’hommeaux摘要:许多自动语音识别(ASR)数据集包括单个预定义的测试集,该测试集由其语音从未出现在训练集中的一个或多个说话者组成。然而,这种“保留说话人”的数据划分策略对于说话人数量非常少的数据集可能并不理想。本研究以最少的ASR训练资源,对五种语言的十种不同的数据分割方法进行了研究。我们发现:(1)模型的性能变化很大,取决于选择哪个扬声器进行测试;(2)所有保持的说话者的平均字错误率(WER)不仅与多个随机分裂上的平均WER相当,而且与任何给定的单个随机分裂相当;(3)当数据被启发式地或敌对地分割时,WER通常也是可比较的;(4)无论数据分割如何,发声持续时间和强度是相对更具有预测性的可变性因素。这些结果表明,广泛使用的ASR数据划分的保持说话人方法可能产生不反映模型对未见数据或说话人的性能的结果。在数据稀疏的情况下,随机分割可以产生更可靠和可推广的估计。摘要:Many automatic speech recognition (ASR) data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. This "hold-speaker(s)-out" data partitioning strategy, however, may not be ideal for data sets in which the number of speakers is very small. This study investigates ten different data split methods for five languages with minimal ASR training resources. We find that (1) model performance varies greatly depending on which speaker is selected for testing; (2) the average word error rate (WER) across all held-out speakers is comparable not only to the average WER over multiple random splits but also to any given individual random split; (3) WER is also generally comparable when the data is split heuristically or adversarially; (4) utterance duration and intensity are comparatively more predictive factors of variability regardless of the data split. These results suggest that the widely used hold-speakers-out approach to ASR data partitioning can yield results that do not reflect model performance on unseen data or speakers. Random splits can yield more reliable and generalizable estimates when facing data sparsity.
【9】 Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization
标题:基于Transformer的目标说话人语音活动检测及其与端到端神经网络二值化的结合
链接:https://arxiv.org/abs/2208.13085
作者:Dongmei Wang,Xiong Xiao,Naoyuki Kanda,Takuya Yoshioka,Jian Wu机构:Microsoft, One Microsoft Way, Redmond, WA, USA摘要:本文提出了一种基于Transformers的目标说话人语音激活检测(TS-VAD)的说话人二值化模型。为了克服原始TS-VAD模型不能处理任意数目说话人的缺点,我们研究了使用具有可变长度时间和说话人维度的输入张量的模型结构。将Transformer层应用于扬声器轴,以使模型输出对提供给TS-VAD模型的扬声器简档的阶不敏感。在这些扬声器方式的Transformer层之间散布时间方式的顺序层,以允许捕获输入语音信号的时间和交叉扬声器相关性。我们还扩展了一个基于端到端神经网络的二分法模型(EEND-EDA),用基于变换的TS-VAD代替其基于点积的说话人检测层。在VoxConverse上的实验结果表明,使用Transformers进行交叉说话人建模将TS-VAD的二值化错误率(DER)降低了10.9%,实现了4.74%的新的最新(SOTA)DER。此外,与具有相似模型大小的原始EEND-EDA相比,我们的扩展EEND-EDA在CALLHOME数据集上将DER降低了6.9%,在广泛使用的训练数据设置下实现了11.18%的新SOTA DER。摘要:This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of being unable to handle an arbitrary number of speakers, we investigate model architectures that use input tensors with variable-length time and speaker dimensions. Transformer layers are applied to the speaker axis to make the model output insensitive to the order of the speaker profiles provided to the TS-VAD model. Time-wise sequential layers are interspersed between these speaker-wise transformer layers to allow the temporal and cross-speaker correlations of the input speech signal to be captured. We also extend a diarization model based on end-to-end neural diarization with encoder-decoder based attractors (EEND-EDA) by replacing its dot-product-based speaker detection layer with the transformer-based TS-VAD. Experimental results on VoxConverse show that using the transformers for the cross-speaker modeling reduces the diarization error rate (DER) of TS-VAD by 10.9%, achieving a new state-of-the-art (SOTA) DER of 4.74%. Also, our extended EEND-EDA reduces DER by 6.9% on the CALLHOME dataset relative to the original EEND-EDA with a similar model size, achieving a new SOTA DER of 11.18% under a widely used training data setting.
【10】 Speech Emotion Recognition using Supervised Deep Recurrent System for Mental Health Monitoring
标题:用于心理健康监测的有监督深度递归系统语音情感识别
链接:https://arxiv.org/abs/2208.12812
作者:Nelly Elsayed,Zag ElSayed,Navid Asadizanjani,Murat Ozer,Ahmed Abdelgawad,Magdy Bayoumi
机构:School of Information Technology, University of Cincinnati, Ohio, United States, Dep. of Electrical & Computer Engineering, University of Florida, Florida, United States, School of Engineering and Technology, Central Michigan University, Michigan, United States备注:6 pages, 5 figures, 3 tables, under reviewing process in the IEEE WFIoT2022摘要:了解人类行为和监测心理健康对维护社区和社会安全至关重要。由于在COVID-19大流行期间,由于不受控制的心理健康,心理健康问题有所增加,因此早期发现心理问题至关重要。如今,智能虚拟个人助理(IVA)的使用在全球范围内增加。个人使用他们的语音来控制这些设备以满足请求并获取不同的服务。提出了一种基于门控递归神经网络和卷积神经网络的深度学习模型,用于从语音中理解人类情感,以改善其IVA服务和监测其心理健康。摘要:Understanding human behavior and monitoring mental health are essential to maintaining the community and society's safety. As there has been an increase in mental health problems during the COVID-19 pandemic due to uncontrolled mental health, early detection of mental issues is crucial. Nowadays, the usage of Intelligent Virtual Personal Assistants (IVA) has increased worldwide. Individuals use their voices to control these devices to fulfill requests and acquire different services. This paper proposes a novel deep learning model based on the gated recurrent neural network and convolution neural network to understand human emotion from speech to improve their IVA services and monitor their mental health.
【1】 Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization
标题:基于Transformer的目标说话人语音活动检测及其与端到端神经网络二值化的结合
链接:https://arxiv.org/abs/2208.13085
* 与cs.SD语音【9】为同一篇
作者:Dongmei Wang,Xiong Xiao,Naoyuki Kanda,Takuya Yoshioka,Jian Wu机构:Microsoft, One Microsoft Way, Redmond, WA, USA摘要:本文提出了一种基于Transformers的目标说话人语音激活检测(TS-VAD)的说话人二值化模型。为了克服原始TS-VAD模型不能处理任意数目说话人的缺点,我们研究了使用具有可变长度时间和说话人维度的输入张量的模型结构。将Transformer层应用于扬声器轴,以使模型输出对提供给TS-VAD模型的扬声器简档的阶不敏感。在这些扬声器方式的Transformer层之间散布时间方式的顺序层,以允许捕获输入语音信号的时间和交叉扬声器相关性。我们还扩展了一个基于端到端神经网络的二分法模型(EEND-EDA),用基于变换的TS-VAD代替其基于点积的说话人检测层。在VoxConverse上的实验结果表明,使用Transformers进行交叉说话人建模将TS-VAD的二值化错误率(DER)降低了10.9%,实现了4.74%的新的最新(SOTA)DER。此外,与具有相似模型大小的原始EEND-EDA相比,我们的扩展EEND-EDA在CALLHOME数据集上将DER降低了6.9%,在广泛使用的训练数据设置下实现了11.18%的新SOTA DER。摘要:This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of being unable to handle an arbitrary number of speakers, we investigate model architectures that use input tensors with variable-length time and speaker dimensions. Transformer layers are applied to the speaker axis to make the model output insensitive to the order of the speaker profiles provided to the TS-VAD model. Time-wise sequential layers are interspersed between these speaker-wise transformer layers to allow the temporal and cross-speaker correlations of the input speech signal to be captured. We also extend a diarization model based on end-to-end neural diarization with encoder-decoder based attractors (EEND-EDA) by replacing its dot-product-based speaker detection layer with the transformer-based TS-VAD. Experimental results on VoxConverse show that using the transformers for the cross-speaker modeling reduces the diarization error rate (DER) of TS-VAD by 10.9%, achieving a new state-of-the-art (SOTA) DER of 4.74%. Also, our extended EEND-EDA reduces DER by 6.9% on the CALLHOME dataset relative to the original EEND-EDA with a similar model size, achieving a new SOTA DER of 11.18% under a widely used training data setting.
【2】 Speech Emotion Recognition using Supervised Deep Recurrent System for Mental Health Monitoring
标题:用于心理健康监测的有监督深度递归系统语音情感识别
链接:https://arxiv.org/abs/2208.12812
* 与cs.SD语音【10】为同一篇
作者:Nelly Elsayed,Zag ElSayed,Navid Asadizanjani,Murat Ozer,Ahmed Abdelgawad,Magdy Bayoumi
机构:School of Information Technology, University of Cincinnati, Ohio, United States, Dep. of Electrical & Computer Engineering, University of Florida, Florida, United States, School of Engineering and Technology, Central Michigan University, Michigan, United States备注:6 pages, 5 figures, 3 tables, under reviewing process in the IEEE WFIoT2022摘要:了解人类行为和监测心理健康对维护社区和社会安全至关重要。由于在COVID-19大流行期间,由于不受控制的心理健康,心理健康问题有所增加,因此早期发现心理问题至关重要。如今,智能虚拟个人助理(IVA)的使用在全球范围内增加。个人使用他们的语音来控制这些设备以满足请求并获取不同的服务。提出了一种基于门控递归神经网络和卷积神经网络的深度学习模型,用于从语音中理解人类情感,以改善其IVA服务和监测其心理健康。摘要:Understanding human behavior and monitoring mental health are essential to maintaining the community and society's safety. As there has been an increase in mental health problems during the COVID-19 pandemic due to uncontrolled mental health, early detection of mental issues is crucial. Nowadays, the usage of Intelligent Virtual Personal Assistants (IVA) has increased worldwide. Individuals use their voices to control these devices to fulfill requests and acquire different services. This paper proposes a novel deep learning model based on the gated recurrent neural network and convolution neural network to understand human emotion from speech to improve their IVA services and monitor their mental health.
【3】 Streaming Intended Query Detection using E2E Modeling for Continued Conversation
标题:基于E2E建模的连续会话流目标查询检测
链接:https://arxiv.org/abs/2208.13322
* 与cs.SD语音【1】为同一篇
作者:Shuo-yiin Chang,Guru Prakash,Zelin Wu,Qiao Liang,Tara N. Sainath,Bo Li,Adam Stambler,Shyam Upadhyay,Manaal Faruqui,Trevor Strohman备注:5 pages, Interspeech 2022摘要:在支持语音的应用中,为了关注www.example.com,通常使用预先确定的热门词汇来激活设备,query.However每次说出查询后都跟随一个热门词汇,这会给后续会话带来认知负担。为了避免重复热点词,我们提出了一种流媒体端到端(E2E)意向查询检测器,它识别指向设备的话语,并过滤掉其他非指向设备的话语。该方法将意图查询检测器引入到E2E模型中,将语音识别流水线的不同部分合并为一个神经网络www.example.comE2Enetwork.The模型中,并基于早期的部分识别结果快速地进行意图查询检测,这对于减少系统的延迟和提高系统的响应速度是非常重要的。实验结果表明,与独立的意图查询检测器相比,E2E方法在等错误率(EER)上的检测准确率提高了22%,延迟时间提高了600ms。在我们的实验中,所提出的模型在用户开始说话后的1.4秒的中位延迟内检测用户是否以8.7%的EER与设备说话。摘要:In voice-enabled applications, a predetermined hotword isusually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each timeintroduces a cognitive burden in continued conversations. Toavoid repeating a hotword, we propose a streaming end-to-end(E2E) intended query detector that identifies the utterancesdirected towards the device and filters out other utterancesnot directed towards device. The proposed approach incor-porates the intended query detector into the E2E model thatalready folds different components of the speech recognitionpipeline into one neural network.The E2E modeling onspeech decoding and intended query detection also allows us todeclare a quick intended query detection based on early partialrecognition result, which is important to decrease latencyand make the system responsive. We demonstrate that theproposed E2E approach yields a 22% relative improvement onequal error rate (EER) for the detection accuracy and 600 mslatency improvement compared with an independent intendedquery detector. In our experiment, the proposed model detectswhether the user is talking to the device with a 8.7% EERwithin 1.4 seconds of median latency after user starts speaking.
【4】 Turn-Taking Prediction for Natural Conversational Speech
标题:自然会话语音的话轮转换预测
链接:https://arxiv.org/abs/2208.13321
* 与cs.SD语音【2】为同一篇
作者:Shuo-yiin Chang,Bo Li,Tara N. Sainath,Chao Zhang,Trevor Strohman,Qiao Liang,Yanzhang He机构:Google Inc., U.S.A备注:5 pages, Interspeech 2022摘要:虽然流式语音助理系统已经在许多应用中使用,但是该系统典型地集中于假定来自单个语音查询的输入没有犹豫或不流畅的不自然的一次性交互。然而,一个普通的会话话语除了不流利之外,还经常涉及带有话轮转换的多个询问。这些不流利包括停顿思考、犹豫、单词加长、填充停顿和重复短语。这使得对会话语音(包括具有多个查询的会话语音)进行语音识别成为一项具有挑战性的任务。为了更好地对会话交互进行建模,关键的是区分不流利和查询结束,以便允许用户针对不流利保持发言权,同时在用户已经结束讲话时使系统尽可能快地做出响应。本文提出了一种建立在端到端语音识别器之上的话轮转换预测器。我们的最佳系统是通过联合优化ASR任务和检测用户何时暂停思考或结束讲话而获得的。实验结果表明,该方法在预测真实话轮转换时的召回率和准确率分别达到97%和85%以上,预测延迟仅为100 ms。摘要:While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice query without hesitation or disfluency. However, a common conversational utterance often involves multiple queries with turn-taking, in addition to disfluencies. These disfluencies include pausing to think, hesitations, word lengthening, filled pauses and repeated phrases. This makes doing speech recognition with conversational speech, including one with multiple queries, a challenging task. To better model the conversational interaction, it is critical to discriminate disfluencies and end of query in order to allow the user to hold the floor for disfluencies while having the system respond as quickly as possible when the user has finished speaking. In this paper, we present a turntaking predictor built on top of the end-to-end (E2E) speech recognizer. Our best system is obtained by jointly optimizing for ASR task and detecting when the user is paused to think or finished speaking. The proposed approach demonstrates over 97% recall rate and 85% precision rate on predicting true turn-taking with only 100 ms latency on a test set designed with 4 types of disfluencies inserted in conversational utterances.
【5】 Computing with Hypervectors for Efficient Speaker Identification
标题:高效说话人辨认的超向量计算
链接:https://arxiv.org/abs/2208.13285
* 与cs.SD语音【3】为同一篇
作者:Ping-Chen Huang,Denis Kleyko,Jan M. Rabaey,Bruno A. Olshausen,Pentti Kanerva机构:Redwood Center of Theoretical Neuroscience, University of California, Berkeley, USA, Intelligent Systems Lab, Research Institutes of Sweden, Sweden, Berkeley Wireless Research Center, University of California, Berkeley, USA摘要:本文提出了一种基于高维随机向量计算的说话人识别方法。它的优点是简单和快速。仅使用www.example.com的1.02k活动参数和128分钟的训练数据,我们在1,251个说话者的VoxCeleb1数据集上获得了31%和52%的前1和前5得分。这与CNN模型形成对比,CNN模型需要数百万个参数和数量级更高的计算复杂性,而在互信息中测量的区分能力中仅获得2 × $的增益。使用广义学习矢量量化(GLVQ)的额外92秒训练将分数提高到48%和67%。经过训练的分类器在5.7ms内对1秒的语音进行分类。所有处理都在基于标准CPU的机器上完成。摘要:We introduce a method to identify speakers by computing with high-dimensional random vectors. Its strengths are simplicity and speed. With only 1.02k active parameters and a 128-minute pass through the training data we achieve Top-1 and Top-5 scores of 31% and 52% on the VoxCeleb1 dataset of 1,251 speakers. This is in contrast to CNN models requiring several million parameters and orders of magnitude higher computational complexity for only a 2$\times$ gain in discriminative power as measured in mutual information. An additional 92 seconds of training with Generalized Learning Vector Quantization (GLVQ) raises the scores to 48% and 67%. A trained classifier classifies 1 second of speech in 5.7 ms. All processing was done on standard CPU-based machines.
【6】 Towards Disentangled Speech Representations
标题:走向无纠缠的言语表征
链接:https://arxiv.org/abs/2208.13191
* 与cs.SD语音【4】为同一篇
作者:Cal Peyser,Ronny Huang Andrew Rosenberg Tara N. Sainath,Michael Picheny,Kyunghyun Cho机构:Cho, Center for Data Science, New York University, New York City, USA, Google Inc., U.S.A摘要:在许多言语任务的方法设计中,音频表征的仔细构建已经成为一个主要特征。这种方法越来越强调“解纠缠,”其中表示只包含语音信号中与转录相关的部分,而丢弃不相关的信息.本文基于ASR和TTS的联合建模,构建了一个表征学习任务,并试图学习一种将语音信号中与转录相关的部分与不相关的部分分开的音频表征。我们提出的经验证据表明,成功地找到这样一个代表是绑在训练中固有的随机性。然后,我们观察到这些期望的、解纠缠的优化问题的解具有独特的统计特性。最后,我们表明,在训练过程中实施这些属性,相对于我们的联合建模任务,平均提高了24.5%的WER。这些观察激发了学习有效音频表示的新颖方法。摘要:The careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks. Increasingly, such approaches have emphasized "disentanglement", where a representation contains only parts of the speech signal relevant to transcription while discarding irrelevant information. In this paper, we construct a representation learning task based on joint modeling of ASR and TTS, and seek to learn a representation of audio that disentangles that part of the speech signal that is relevant to transcription from that part which is not. We present empirical evidence that successfully finding such a representation is tied to the randomness inherent in training. We then make the observation that these desired, disentangled solutions to the optimization problem possess unique statistical properties. Finally, we show that enforcing these properties during training improves WER by 24.5% relative on average for our joint modeling task. These observations motivate a novel approach to learning effective audio representations.
【7】 Training Text-To-Speech Systems From Synthetic Data: A Practical Approach For Accent Transfer Tasks
标题:从合成数据训练文本到语音系统:一种实用的口音转移任务方法
链接:https://arxiv.org/abs/2208.13183
* 与cs.SD语音【5】为同一篇
作者:Lev Finkelstein,Heiga Zen,Norman Casagrande,Chun-an Chan,Ye Jia,Tom Kenter,Alexey Petelin,Jonathan Shen,Vincent Wan,Yu Zhang,Yonghui Wu,Rob Clark备注:To be published in Interspeech 2022摘要:文本到语音(TTS)合成中的传输任务——其中一组说话者的语音的一个或多个方面被传输到最初不具有这些方面的另一组说话者——仍然是一项具有挑战性的任务。其中一个挑战是,具有高质量传输能力的模型可能存在稳定性问题,这使得它们不适用于面向用户的关键任务。本文证明,通过训练鲁棒TTS系统,可以获得高质量的传输任务,数据由设计用于高质量传输任务的鲁棒性较差的TTS系统生成;特别地,在为重音转移设计的Tacotron模型的输出上训练CHiVE-BERT单语TTS系统。虽然用这种方法不可避免地会有一些质量损失,但实验结果表明,用这种方法在合成数据上训练的模型可以产生显示重音转移的高质量音频,同时保留说话者的特征,例如说话风格。摘要:Transfer tasks in text-to-speech (TTS) synthesis - where one or more aspects of the speech of one set of speakers is transferred to another set of speakers that do not feature these aspects originally - remains a challenging task. One of the challenges is that models that have high-quality transfer capabilities can have issues in stability, making them impractical for user-facing critical tasks. This paper demonstrates that transfer can be obtained by training a robust TTS system on data generated by a less robust TTS system designed for a high-quality transfer task; in particular, a CHiVE-BERT monolingual TTS system is trained on the output of a Tacotron model designed for accent transfer. While some quality loss is inevitable with this approach, experimental results show that the models trained on synthetic data this way can produce high quality audio displaying accent transfer, while preserving speaker characteristics such as speaking style.
【8】 SA: Sliding attack for synthetic speech detection with resistance to clipping and self-splicing
标题:SA:抗剪裁和自拼接的合成语音检测滑动攻击
链接:https://arxiv.org/abs/2208.13066
* 与cs.SD语音【6】为同一篇
作者:Deng JiaCheng,Dong Li,Yan Diqun,Wang Rangding,Zeng Jiaming机构:Department of Computer Science, Ningbo University, Ningbo , China, A R T I C L E I N F O备注:12 pages, Neurocomputing摘要:深度神经网络容易受到敌对例子的影响,这些例子会用不可察觉的扰动误导模型。在音频中,尽管对抗性攻击的例子在白盒设置和黑盒设置上取得了令人难以置信的攻击成功率,但大多数现有的对抗性攻击受到输入长度的限制。一个更实际的场景是对抗性的例子必须被剪切或自拼接并输入到黑盒模型中。因此,有必要探索如何在不同的输入长度设置下提高可转移性。本文以合成语音检测任务为例,考虑了两种有代表性的SOTA模型。通过对剪切或自拼接后样本输入模型得到的梯度进行分析,发现相同样本值的片段在不同模型中的梯度是相似的。受此启发,我们提出了一种新的对抗性攻击方法——滑动攻击。具体地说,我们使每个采样点知道不同位置的梯度,这可以模拟具有不同输入长度的敌对样本被输入到黑盒模型的情况。因此,代替在梯度计算的每次迭代中直接使用当前梯度,我们经历以下三个步骤。首先,我们使用滑动窗口提取不同长度的子段。然后,我们用来自相邻域的数据扩充子段。最后,我们将子片段馈入不同的模型中以获得聚合梯度来更新对抗范例。实验结果表明,该方法能显著提高剪切或自拼接后的对抗性实例的可移植性。此外,该方法还可以增强基于不同特征的模型之间的可移植性。摘要:Deep neural networks are vulnerable to adversarial examples that mislead models with imperceptible perturbations. In audio, although adversarial examples have achieved incredible attack success rates on white-box settings and black-box settings, most existing adversarial attacks are constrained by the input length. A More practical scenario is that the adversarial examples must be clipped or self-spliced and input into the black-box model. Therefore, it is necessary to explore how to improve transferability in different input length settings. In this paper, we take the synthetic speech detection task as an example and consider two representative SOTA models. We observe that the gradients of fragments with the same sample value are similar in different models via analyzing the gradients obtained by feeding samples into the model after cropping or self-splicing. Inspired by the above observation, we propose a new adversarial attack method termed sliding attack. Specifically, we make each sampling point aware of gradients at different locations, which can simulate the situation where adversarial examples are input to black-box models with varying input lengths. Therefore, instead of using the current gradient directly in each iteration of the gradient calculation, we go through the following three steps. First, we extract subsegments of different lengths using sliding windows. We then augment the subsegments with data from the adjacent domains. Finally, we feed the sub-segments into different models to obtain aggregate gradients to update adversarial examples. Empirical results demonstrate that our method could significantly improve the transferability of adversarial examples after clipping or self-splicing. Besides, our method could also enhance the transferability between models based on different features.
【9】 Sub-mW Neuromorphic SNN audio processing applications with Rockpool and Xylo
标题:采用Rockpool和Xylo的亚毫瓦神经形态SNN音频处理应用
链接:https://arxiv.org/abs/2208.12991
* 与cs.SD语音【7】为同一篇
摘要:尖峰神经网络(SNN)为时间信号处理提供了一种高效的计算机制,尤其是在与低功耗SNN推理ASIC耦合时。SNN历来难以配置,缺乏为任意任务找到解决方案的通用方法。近年来,梯度下降优化方法已经越来越容易地应用于SNNs。因此,SNN和SNN推理处理器提供了一个良好的平台,用于在没有云依赖性的能量受限环境中进行商业低功耗信号处理。然而,到目前为止,这些方法还没有被工业中的ML工程师所使用,需要研究生水平的培训才能成功地配置单个SNN应用。在这里,我们展示了一个方便的高级流水线,用于设计、训练和部署任意时间信号处理应用到亚mW SNN推理硬件。我们采用一种新的简单的SNN结构,使用突触时间常数金字塔来提取时间尺度范围内的信号特征。我们在一个环境音频分类任务上演示了该架构,该任务以流模式部署到Xylo SNN推理处理器。我们的应用在低功耗(〈4 μ W推理功率)下实现了高准确度(98%)和低延迟(100ms)。我们的方法使得具有一般神经网络背景的ML工程师可以培训和部署SNN应用,而不需要具有尖峰神经网络的特定经验。我们希望我们的方法使神经形态硬件和SNN成为商业低功耗和边缘信号处理应用的有吸引力的选择。摘要:Spiking Neural Networks (SNNs) provide an efficient computational mechanism for temporal signal processing, especially when coupled with low-power SNN inference ASICs. SNNs have been historically difficult to configure, lacking a general method for finding solutions for arbitrary tasks. In recent years, gradient-descent optimization methods have been applied to SNNs with increasing ease. SNNs and SNN inference processors therefore offer a good platform for commercial low-power signal processing in energy constrained environments without cloud dependencies. However, to date these methods have not been accessible to ML engineers in industry, requiring graduate-level training to successfully configure a single SNN application. Here we demonstrate a convenient high-level pipeline to design, train and deploy arbitrary temporal signal processing applications to sub-mW SNN inference hardware. We apply a new straightforward SNN architecture designed for temporal signal processing, using a pyramid of synaptic time constants to extract signal features at a range of temporal scales. We demonstrate this architecture on an ambient audio classification task, deployed to the Xylo SNN inference processor in streaming mode. Our application achieves high accuracy (98%) and low latency (100ms) at low power (<4muW inference power). Our approach makes training and deploying SNN applications available to ML engineers with general NN backgrounds, without requiring specific prior experience with spiking NNs. We intend for our approach to make Neuromorphic hardware and SNNs an attractive choice for commercial low-power and edge signal processing applications.
【10】 Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation
标题:跨语言低资源ASR评估的数据分割策略研究
链接:https://arxiv.org/abs/2208.12888
* 与cs.SD语音【8】为同一篇
作者:Zoey Liu,Justin Spence,Emily Prud'hommeaux机构:Boston College, University of California, Davis, Emily Prud’hommeaux摘要:许多自动语音识别(ASR)数据集包括单个预定义的测试集,该测试集由其语音从未出现在训练集中的一个或多个说话者组成。然而,这种“保留说话人”的数据划分策略对于说话人数量非常少的数据集可能并不理想。本研究以最少的ASR训练资源,对五种语言的十种不同的数据分割方法进行了研究。我们发现:(1)模型的性能变化很大,取决于选择哪个扬声器进行测试;(2)所有保持的说话者的平均字错误率(WER)不仅与多个随机分裂上的平均WER相当,而且与任何给定的单个随机分裂相当;(3)当数据被启发式地或敌对地分割时,WER通常也是可比较的;(4)无论数据分割如何,发声持续时间和强度是相对更具有预测性的可变性因素。这些结果表明,广泛使用的ASR数据划分的保持说话人方法可能产生不反映模型对未见数据或说话人的性能的结果。在数据稀疏的情况下,随机分割可以产生更可靠和可推广的估计。摘要:Many automatic speech recognition (ASR) data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. This "hold-speaker(s)-out" data partitioning strategy, however, may not be ideal for data sets in which the number of speakers is very small. This study investigates ten different data split methods for five languages with minimal ASR training resources. We find that (1) model performance varies greatly depending on which speaker is selected for testing; (2) the average word error rate (WER) across all held-out speakers is comparable not only to the average WER over multiple random splits but also to any given individual random split; (3) WER is also generally comparable when the data is split heuristically or adversarially; (4) utterance duration and intensity are comparatively more predictive factors of variability regardless of the data split. These results suggest that the widely used hold-speakers-out approach to ASR data partitioning can yield results that do not reflect model performance on unseen data or speakers. Random splits can yield more reliable and generalizable estimates when facing data sparsity.
机器翻译,仅供参考