cs.SD语音,共计7篇,eess.AS音频处理,共计11篇


1.cs.SD语音:

【1】 Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator标题:使用树约束指针生成器最小化上下文ASR的偏向单词错误

链接:https://arxiv.org/abs/2205.09058

作者:Guangzhi Sun,Chao Zhang,Philip C Woodland
备注:This work has been submitted to the IEEE Transactions on Audio, Speech, and Language Processing for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:上下文知识对于减少高值长尾词的语音识别错误至关重要。本文提出了一种新的树约束指针生成器(TCPGen)组件,该组件允许端到端ASR模型偏向使用外部上下文信息获得的长尾词列表。TCPGen只需在内存使用和计算成本方面花费很小的开销,就可以有效地将数千个偏差词构造成一个符号前缀树,并在树和最终ASR输出之间创建一个神经捷径,以便于识别偏差词。为了增强TCPGen,我们进一步提出了一种新的最小偏差词错误(MBWE)损失,该损失可以直接优化训练期间的偏差词错误,以及测试期间的偏差词驱动语言模型折扣(BLMD)方法。所有上下文ASR系统均在公共Librispeech有声读物语料库和对话状态跟踪挑战(DSTC)数据上进行评估,并从对话系统本体中提取偏差列表。TCPGen实现了一致的单词错误率(WER)降低,这在偏误单词上尤其显著,识别错误率相对降低约40%。MBWE和BLMD进一步提高了TCPGen的有效性,并在偏倚词上实现了更显著的WER降低。TCPGen还实现了对不在音频训练集中的单词的零射门学习,大大减少了偏差列表中的词汇外单词。
摘要:Contextual knowledge is essential for reducing speech recognition errors on high-valued long-tail words. This paper proposes a novel tree-constrained pointer generator (TCPGen) component that enables end-to-end ASR models to bias towards a list of long-tail words obtained using external contextual information. With only a small overhead in memory use and computation cost, TCPGen can structure thousands of biasing words efficiently into a symbolic prefix-tree and creates a neural shortcut between the tree and the final ASR output to facilitate the recognition of the biasing words. To enhance TCPGen, we further propose a novel minimum biasing word error (MBWE) loss that directly optimises biasing word errors during training, along with a biasing-word-driven language model discounting (BLMD) method during the test. All contextual ASR systems were evaluated on the public Librispeech audiobook corpus and the data from the dialogue state tracking challenges (DSTC) with the biasing lists extracted from the dialogue-system ontology. Consistent word error rate (WER) reductions were achieved with TCPGen, which were particularly significant on the biasing words with around 40\% relative reductions in the recognition error rates. MBWE and BLMD further improved the effectiveness of TCPGen and achieved more significant WER reductions on the biasing words. TCPGen also achieved zero-shot learning of words not in the audio training set with large WER reductions on the out-of-vocabulary words in the biasing list.


【2】 Seeing Sounds, Hearing Shapes: a gamified study to evaluate sound-sketches

标题:看得见的声音,听得到的形状:评估声音素描的游戏化研究

链接:https://arxiv.org/abs/2205.08866

作者:Sebastian Löbbers,György Fazekas
备注:Accepted at International Computer Music Conference (ICMC) 2022
摘要:声音-形状关联是听觉和视觉领域之间跨模态关联的一个子集,主要是在将一组精心制作的形状与声音匹配的背景下进行研究。最近的研究探索了人类如何通过自由形式的素描来表现声音,以及如何将图形素描输入用于声音制作。在本文中,通过两次在线展览活动中的82名参与者进行了一项游戏化研究,探讨了通过这些自由形式草图传达声音特征的潜力。结果表明,参与者能够以比随机基线更高的速度识别声音,但似乎很难对细微的音色差异进行视觉编码。
摘要:Sound-shape associations, a subset of cross-modal associations between the auditory and visual domain, have been studied mainly in the context of matching a set of purposefully crafted shapes to sounds. Recent studies have explored how humans represent sound through free-form sketching and how a graphical sketch input could be used for sound production. In this paper, the potential of communicating sound characteristics through these free-form sketches is investigated in a gamified study that was conducted with eighty-two participants at two online exhibition events. The results show that participants managed to recognise sounds at a higher rate than the random baseline would suggest, however it appeared difficult to visually encode nuanced timbral differences.


【3】 Deploying self-supervised learning in the wild for hybrid automatic speech recognition

标题:在野外部署自监督学习用于混合自动语音识别

链接:https://arxiv.org/abs/2205.08598

作者:Mostafa Karimi,Changliang Liu,Kenichi Kumatani,Yao Qian,Tianyu Wu,Jian Wu
摘要:自监督学习(SSL)方法已被证明在自动语音识别(ASR)中非常成功。据报道,这些巨大的改进主要基于高度管理的数据集,如用于非流式端到端ASR模型的LibriSpeech。然而,SSL的关键特性是可用于任何未翻译的音频数据。在本文中,我们将全面探讨如何在SSL中利用未固化的音频数据,从数据预处理到部署流式混合ASR模型。更具体地说,我们展示了(1)数据预处理管道中音频事件检测(AED)模型的效果(2)对优化器选择和学习率调度的分析(3)最近开发的对比损失的比较,(4)各种预训练策略的比较,如域内和域外预训练数据的利用,单语言与多语言预训练数据,多头多语言SSL与单头多语言SSL,监督预训练与SSL。实验结果表明,与所有可选的域外预训练策略相比,使用域内未固化数据的SSL预训练可以获得更好的性能。
摘要:Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly curated datasets such as LibriSpeech for non-streaming End-to-End ASR models. However, the pivotal characteristics of SSL is to be utilized for any untranscribed audio data. In this paper, we provide a full exploration on how to utilize uncurated audio data in SSL from data pre-processing to deploying an streaming hybrid ASR model. More specifically, we present (1) the effect of Audio Event Detection (AED) model in data pre-processing pipeline (2) analysis on choosing optimizer and learning rate scheduling (3) comparison of recently developed contrastive losses, (4) comparison of various pre-training strategies such as utilization of in-domain versus out-domain pre-training data, monolingual versus multilingual pre-training data, multi-head multilingual SSL versus single-head multilingual SSL and supervised pre-training versus SSL. The experimental results show that SSL pre-training with in-domain uncurated data can achieve better performance in comparison to all the alternative out-domain pre-training strategies.


【4】 The Power of Reuse: A Multi-Scale Transformer Model for Structural Dynamic Segmentation in Symbolic Music Generation

标题:重用的力量:符号音乐生成中结构动态分割的多尺度变换模型

链接:https://arxiv.org/abs/2205.08579

作者:Guowei Wu,Shipei Liu,Xiaoya Fan
摘要:符号音乐的生成依赖于生成模型的上下文表示能力,其中最流行的方法是基于转换器的模型。不仅如此,长期语境的学习还与音乐结构的动态分割有关,即介绍、韵文和合唱,这目前被研究界所忽视。在本文中,我们提出了一种多尺度变换器,它使用粗解码器和细解码器分别在全局和部分级别对上下文进行建模。具体来说,我们设计了一个片段范围定位层,将音乐切分为多个部分,这些部分随后用于预训练精细解码器。然后,我们设计了一个音乐风格规范化层,将风格信息从原始部分传输到生成的部分,以实现音乐风格的一致性。生成的部分在聚合层中进行组合,并由粗解码器进行微调。我们的模型在两个开放的MIDI数据集上进行了评估,实验表明,我们的模型优于当代最好的符号音乐生成模型。更令人兴奋的是,视觉评估表明,我们的模型在旋律重用方面具有优势,从而产生了更逼真的音乐。
摘要:Symbolic Music Generation relies on the contextual representation capabilities of the generative model, where the most prevalent approach is the Transformer-based model. Not only that, the learning of long-term context is also related to the dynamic segmentation of musical structures, i.e. intro, verse and chorus, which is currently overlooked by the research community. In this paper, we propose a multi-scale Transformer, which uses coarse-decoder and fine-decoders to model the contexts at the global and section-level, respectively. Concretely, we designed a Fragment Scope Localization layer to syncopate the music into sections, which were later used to pre-train fine-decoders. After that, we designed a Music Style Normalization layer to transfer the style information from the original sections to the generated sections to achieve consistency in music style. The generated sections are combined in the aggregation layer and fine-tuned by the coarse decoder. Our model is evaluated on two open MIDI datasets, and experiments show that our model outperforms the best contemporary symbolic music generative models. More excitingly, visual evaluation shows that our model is superior in melody reuse, resulting in more realistic music.


【5】 Dictionary-Based Fusion of Contact and Acoustic Microphones for Wind Noise Reduction

标题:基于字典的接触式麦克风和声学麦克风的融合降噪

链接:https://arxiv.org/abs/2205.09017

作者:Marvin Tammen,Xilin Li,Simon Doclo,Lalin Theverapperuma
备注:submitted to IWAENC 22
摘要:在移动语音通信应用中,风噪声会导致语音质量和清晰度的严重降低。由于使用声学麦克风的语音增强算法的性能在极具挑战性的场景中往往会大幅下降,因此可以使用接触式麦克风等辅助传感器。虽然接触式话筒提供的记录风噪级别要低得多,但它们是以语音失真和额外的噪声分量为代价的。为了利用声学麦克风和接触式麦克风在风噪声抑制方面的优势,本文提出通过同时建模声学麦克风和接触式麦克风信号来扩展传统的基于单麦克风字典的语音增强方法。我们建议训练一个语音字典和两个噪声字典,并使用相对传递函数来建模麦克风处语音成分之间的关系。仿真结果表明,与几种基线方法相比,该方法在语音质量和可懂度方面都有所提高,其中最显著的是仅使用接触式麦克风或仅使用声学麦克风的方法。
摘要:In mobile speech communication applications, wind noise can lead to a severe reduction of speech quality and intelligibility. Since the performance of speech enhancement algorithms using acoustic microphones tends to substantially degrade in extremely challenging scenarios, auxiliary sensors such as contact microphones can be used. Although contact microphones offer a much lower recorded wind noise level, they come at the cost of speech distortion and additional noise components. Aiming at exploiting the advantages of acoustic and contact microphones for wind noise reduction, in this paper we propose to extend conventional single-microphone dictionary-based speech enhancement approaches by simultaneously modeling the acoustic and contact microphone signals. We propose to train a single speech dictionary and two noise dictionaries and use a relative transfer function to model the relationship between the speech components at the microphones. Simulation results show that the proposed approach yields improvements in both speech quality and intelligibility compared to several baseline approaches, most notably approaches using only the contact microphones or only the acoustic microphone.


【6】 Deep Multi-Frame MVDR Filtering for Binaural Noise Reduction

标题:深度多帧MVDR双耳降噪方法

链接:https://arxiv.org/abs/2205.08983

作者:Marvin Tammen,Simon Doclo
备注:submitted to IWAENC 2022
摘要:为了提高噪声环境中的语音清晰度和语音质量,头戴式辅助听力设备的双耳降噪算法至关重要。提出了几种双耳降噪算法,如著名的双耳最小方差无失真响应(MVDR)波束形成器,该算法利用了目标语音和噪声分量的空间相关性。此外,对于单话筒场景,已经提出了多帧算法,如多帧MVDR(MFMVDR)滤波器,它利用时间而不是空间相关性。在本文中,我们提出了MFMVDR滤波器的双耳扩展,它利用了空间和时间相关性。双耳MFMVDR滤波器嵌入到端到端的深度学习框架中,其中所需参数,即语音时空相关向量以及(逆)噪声时空协方差矩阵,由时间卷积网络(TCN)估计,时间卷积网络通过最小化平均谱绝对误差损失函数进行训练。仿真结果包括测量的双耳房间脉冲和信噪比为-5 dB至20 dB的各种噪声源,表明使用双耳MFMVDR滤波器结构比直接使用TCN估计双耳多帧滤波器系数更具优势。
摘要:To improve speech intelligibility and speech quality in noisy environments, binaural noise reduction algorithms for head-mounted assistive listening devices are of crucial importance. Several binaural noise reduction algorithms such as the well-known binaural minimum variance distortionless response (MVDR) beamformer have been proposed, which exploit spatial correlations of both the target speech and the noise components. Furthermore, for single-microphone scenarios, multi-frame algorithms such as the multi-frame MVDR (MFMVDR) filter have been proposed, which exploit temporal instead of spatial correlations. In this contribution, we propose a binaural extension of the MFMVDR filter, which exploits both spatial and temporal correlations. The binaural MFMVDR filters are embedded in an end-to-end deep learning framework, where the required parameters, i.e., the speech spatio-temporal correlation vectors as well as the (inverse) noise spatio-temporal covariance matrix, are estimated by temporal convolutional networks (TCNs) that are trained by minimizing the mean spectral absolute error loss function. Simulation results comprising measured binaural room impulses and diverse noise sources at signal-to-noise ratios from -5 dB to 20 dB demonstrate the advantage of utilizing the binaural MFMVDR filter structure over directly estimating the binaural multi-frame filter coefficients with TCNs.


【7】 Streaming Noise Context Aware Enhancement For Automatic Speech Recognition in Multi-Talker Environments

标题:多人环境下自动语音识别的流噪声上下文感知增强

链接:https://arxiv.org/abs/2205.08555

作者:Joe Caroselli,Arun Narayanan,Yiteng Huang
备注:Submitted to IWAENC 2022
摘要:对于智能扬声器来说,最具挑战性的场景之一是多说话人,当来自所需扬声器的目标语音与来自一个或多个扬声器的干扰语音混合时。智能助理需要确定要识别和忽略哪个语音,并且需要以流式、低延迟的方式进行。本文针对这种情况提出了两种多麦克风语音增强算法。针对设备用例,我们假设该算法能够访问hotword之前的信号,即噪声上下文。首先是上下文感知波束形成器,它使用噪声上下文和检测到的热词来确定如何定位所需的说话人。第二种是一种称为语音清洁器的自适应噪声消除算法,它使用噪声上下文训练滤波器。结果表明,这两种算法在信噪比条件下是互补的,在信噪比条件下,这两种算法都能很好地工作。我们还提出了一种基于估计信噪比的算法来选择要使用的信噪比。当使用3个麦克风通道时,最终系统在-12dB时的相对字错误率降低了55%,在12dB时的相对字错误率降低了43%。
摘要:One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which voice to recognize and which to ignore and it needs to do so in a streaming, low-latency manner. This work presents two multi-microphone speech enhancement algorithms targeted at this scenario. Targeting on-device use-cases, we assume that the algorithm has access to the signal before the hotword, which is referred to as the noise context. First is the Context Aware Beamformer which uses the noise context and detected hotword to determine how to target the desired speaker. The second is an adaptive noise cancellation algorithm called Speech Cleaner which trains a filter using the noise context. It is demonstrated that the two algorithms are complementary in the signal-to-noise ratio conditions under which they work well. We also propose an algorithm to select which one to use based on estimated SNR. When using 3 microphone channels, the final system achieves a relative word error rate reduction of 55% at -12dB, and 43\% at 12dB.


2.eess.AS音频处理:

【1】 Dictionary-Based Fusion of Contact and Acoustic Microphones for Wind Noise Reduction

标题:基于字典的接触式麦克风和声学麦克风的融合降噪

链接:https://arxiv.org/abs/2205.09017

作者:Marvin Tammen,Xilin Li,Simon Doclo,Lalin Theverapperuma
备注:submitted to IWAENC 22
摘要:在移动语音通信应用中,风噪声会导致语音质量和清晰度的严重降低。由于使用声学麦克风的语音增强算法的性能在极具挑战性的场景中往往会大幅下降,因此可以使用接触式麦克风等辅助传感器。虽然接触式话筒提供的记录风噪级别要低得多,但它们是以语音失真和额外的噪声分量为代价的。为了利用声学麦克风和接触式麦克风在风噪声抑制方面的优势,本文提出通过同时建模声学麦克风和接触式麦克风信号来扩展传统的基于单麦克风字典的语音增强方法。我们建议训练一个语音字典和两个噪声字典,并使用相对传递函数来建模麦克风处语音成分之间的关系。仿真结果表明,与几种基线方法相比,该方法在语音质量和可懂度方面都有所提高,其中最显著的是仅使用接触式麦克风或仅使用声学麦克风的方法。
摘要:In mobile speech communication applications, wind noise can lead to a severe reduction of speech quality and intelligibility. Since the performance of speech enhancement algorithms using acoustic microphones tends to substantially degrade in extremely challenging scenarios, auxiliary sensors such as contact microphones can be used. Although contact microphones offer a much lower recorded wind noise level, they come at the cost of speech distortion and additional noise components. Aiming at exploiting the advantages of acoustic and contact microphones for wind noise reduction, in this paper we propose to extend conventional single-microphone dictionary-based speech enhancement approaches by simultaneously modeling the acoustic and contact microphone signals. We propose to train a single speech dictionary and two noise dictionaries and use a relative transfer function to model the relationship between the speech components at the microphones. Simulation results show that the proposed approach yields improvements in both speech quality and intelligibility compared to several baseline approaches, most notably approaches using only the contact microphones or only the acoustic microphone.


【2】 Coherence-Based Frequency Subset Selection For Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers

标题:基于相干的多说话人双耳RTF矢量波达方向估计的频率子集选择

链接:https://arxiv.org/abs/2205.08985

作者:Daniel Fejgin,Simon Doclo
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:最近,有人提出了一种通过最小化估计的相对传递函数(RTF)向量和原型消声RTF向量数据库之间的频率平均厄米角来估计单个说话人的到达方向(DOA)的方法。本文通过引入频率平均厄米角谱并选择该空间谱的峰值,将该方法推广到多说话人定位。为了构造厄米角谱,我们只考虑频率的子集,其中可能有一个说话人占主导地位。我们比较了广义幅度平方相干和两种相干扩散比(CDR)估计器作为频率选择准则的有效性。利用双耳听觉设备对混响环境中两个说话人的DOA进行估计的仿真结果表明,使用基于双耳有效相干性的CDR估计作为频率选择准则可以获得最佳性能。
摘要:Recently, a method has been proposed to estimate the direction of arrival (DOA) of a single speaker by minimizing the frequency-averaged Hermitian angle between an estimated relative transfer function (RTF) vector and a database of prototype anechoic RTF vectors. In this paper, we extend this method to multi-speaker localization by introducing the frequency-averaged Hermitian angle spectrum and selecting peaks of this spatial spectrum. To construct the Hermitian angle spectrum, we consider only a subset of frequencies, where it is likely that one speaker is dominant. We compare the effectiveness of the generalized magnitude squared coherence and two coherent-to-diffuse ratio (CDR) estimators as frequency selection criteria. Simulation results for estimating the DOAs of two speakers in a reverberant environment with diffuse-like babble noise using binaural hearing devices show that using the binaural effective-coherence-based CDR estimate as a frequency selection criterion yields the best performance.


【3】 Deep Multi-Frame MVDR Filtering for Binaural Noise Reduction

标题:深度多帧MVDR双耳降噪方法

链接:https://arxiv.org/abs/2205.08983

作者:Marvin Tammen,Simon Doclo
备注:submitted to IWAENC 2022
摘要:为了提高噪声环境中的语音清晰度和语音质量,头戴式辅助听力设备的双耳降噪算法至关重要。提出了几种双耳降噪算法,如著名的双耳最小方差无失真响应(MVDR)波束形成器,该算法利用了目标语音和噪声分量的空间相关性。此外,对于单话筒场景,已经提出了多帧算法,如多帧MVDR(MFMVDR)滤波器,它利用时间而不是空间相关性。在本文中,我们提出了MFMVDR滤波器的双耳扩展,它利用了空间和时间相关性。双耳MFMVDR滤波器嵌入到端到端的深度学习框架中,其中所需参数,即语音时空相关向量以及(逆)噪声时空协方差矩阵,由时间卷积网络(TCN)估计,时间卷积网络通过最小化平均谱绝对误差损失函数进行训练。仿真结果包括测量的双耳房间脉冲和信噪比为-5 dB至20 dB的各种噪声源,表明使用双耳MFMVDR滤波器结构比直接使用TCN估计双耳多帧滤波器系数更具优势。
摘要:To improve speech intelligibility and speech quality in noisy environments, binaural noise reduction algorithms for head-mounted assistive listening devices are of crucial importance. Several binaural noise reduction algorithms such as the well-known binaural minimum variance distortionless response (MVDR) beamformer have been proposed, which exploit spatial correlations of both the target speech and the noise components. Furthermore, for single-microphone scenarios, multi-frame algorithms such as the multi-frame MVDR (MFMVDR) filter have been proposed, which exploit temporal instead of spatial correlations. In this contribution, we propose a binaural extension of the MFMVDR filter, which exploits both spatial and temporal correlations. The binaural MFMVDR filters are embedded in an end-to-end deep learning framework, where the required parameters, i.e., the speech spatio-temporal correlation vectors as well as the (inverse) noise spatio-temporal covariance matrix, are estimated by temporal convolutional networks (TCNs) that are trained by minimizing the mean spectral absolute error loss function. Simulation results comprising measured binaural room impulses and diverse noise sources at signal-to-noise ratios from -5 dB to 20 dB demonstrate the advantage of utilizing the binaural MFMVDR filter structure over directly estimating the binaural multi-frame filter coefficients with TCNs.


【4】 3D Single Source Localization Based on Euclidean Distance Matrices

标题:基于欧氏距离矩阵的三维单源定位

链接:https://arxiv.org/abs/2205.08960

作者:Klaus Brümann,Simon Doclo
备注:5 pages (last page references), 3 figures, 1 table, submitted to "International Workshop on Acoustic Signal Enhancement (IWAENC), Bamberg, 2022"
摘要:使用多个麦克风进行三维声源定位的一种流行方法是方向响应功率法,其中通过最大化三个连续位置变量的函数直接估计声源位置。在本文中,我们提出了一种基于距离的间接三维源定位方法,而不是直接估计源位置。基于欧几里德距离矩阵(EDMs)的性质,我们将三维声源定位问题转化为单个变量的代价函数的最小化,即声源与参考麦克风之间的距离。使用已知的麦克风几何结构和麦克风之间的估计到达时间差(TDOAs),我们展示了如何基于此变量计算三维震源位置。此外,我们提出了一种扩展,可以从一组候选TDOA估计值中选择最合适的估计值,而不是使用每个麦克风对的单个TDOA估计值,这在具有强烈早期反射的混响环境中尤其相关。对不同声源和麦克风星座的实验结果表明,基于EDM的方法始终优于转向响应功率法,尤其是当声源靠近麦克风时。
摘要:A popular approach for 3D source localization using multiple microphones is the steered-response power method, where the source position is directly estimated by maximizing a function of three continuous position variables. Instead of directly estimating the source position, in this paper we propose an indirect, distance-based method for 3D source localization. Based on properties of Euclidean distance matrices (EDMs), we reformulate the 3D source localization problem as the minimization of a cost function of a single variable, namely the distance between the source and the reference microphone. Using the known microphone geometry and estimated time-differences of arrival (TDOAs) between the microphones, we show how the 3D source position can be computed based on this variable. In addition, instead of using a single TDOA estimate per microphone pair, we propose an extension that enables to select the most appropriate estimate from a set of candidate TDOA estimates, which is especially relevant in reverberant environments with strong early reflections. Experimental results for different source and microphone constellations show that the proposed EDM-based method consistently outperforms the steered-response power method, especially when the source is close to the microphones.


【5】 U-Former: Improving Monaural Speech Enhancement with Multi-head Self and Cross Attention

标题:U-Form:利用多头自我注意和交叉注意改进单声道语音增强

链接:https://arxiv.org/abs/2205.08681

作者:Xinmeng Xu,Jianjun Hao
备注:Accepted by ICPR 2022. arXiv admin note: text overlap with arXiv:2112.06052, arXiv:2103.06104, arXiv:2103.09963 by other authors
摘要:对于有监督的语音增强,上下文信息对于精确的频谱映射非常重要。然而,常用的深度神经网络(DNN)在捕获时间上下文方面受到限制。为了利用长期上下文跟踪目标说话人,本文将语音增强视为序列到序列的映射,并提出了一种新的基于变换器的单耳语音增强U网络结构,称为U-Former。其关键思想是通过多头注意机制对长期相关性和依赖性进行建模,这对于精确的噪声语音建模至关重要。为此,U-Former在两个层面上整合了多头注意机制:1)多头自我注意模块,该模块沿时间和频率轴计算注意图,以生成时间和频率子注意图,以利用编码器功能之间的全局交互,而插入跳过连接的多头交叉注意模块通过过滤掉不相关的特征,可以在解码器中进行精细恢复。实验结果表明,U-Former的性能始终优于PESQ、STOI和SSNR分数的最新模型。
摘要:For supervised speech enhancement, contextual information is important for accurate spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term contexts for tracking a target speaker, this paper treats the speech enhancement as sequence-to-sequence mapping, and propose a novel monaural speech enhancement U-net structure based on Transformer, dubbed U-Former. The key idea is to model long-term correlations and dependencies, which are crucial for accurate noisy speech modeling, through the multi-head attention mechanisms. For this purpose, U-Former incorporates multi-head attention mechanisms at two levels: 1) a multi-head self-attention module which calculate the attention map along both time- and frequency-axis to generate time and frequency sub-attention maps for leveraging global interactions between encoder features, while 2) multi-head cross-attention module which are inserted in the skip connections allows a fine recovery in the decoder by filtering out uncorrelated features. Experimental results illustrate that the U-Former obtains consistently better performance than recent models of PESQ, STOI, and SSNR scores.


【6】 Streaming Noise Context Aware Enhancement For Automatic Speech Recognition in Multi-Talker Environments

标题:多人环境下自动语音识别的流噪声上下文感知增强

链接:https://arxiv.org/abs/2205.08555

作者:Joe Caroselli,Arun Narayanan,Yiteng Huang
备注:Submitted to IWAENC 2022
摘要:对于智能扬声器来说,最具挑战性的场景之一是多说话人,当来自所需扬声器的目标语音与来自一个或多个扬声器的干扰语音混合时。智能助理需要确定要识别和忽略哪个语音,并且需要以流式、低延迟的方式进行。本文针对这种情况提出了两种多麦克风语音增强算法。针对设备用例,我们假设该算法能够访问hotword之前的信号,即噪声上下文。首先是上下文感知波束形成器,它使用噪声上下文和检测到的热词来确定如何定位所需的说话人。第二种是一种称为语音清洁器的自适应噪声消除算法,它使用噪声上下文训练滤波器。结果表明,这两种算法在信噪比条件下是互补的,在信噪比条件下,这两种算法都能很好地工作。我们还提出了一种基于估计信噪比的算法来选择要使用的信噪比。当使用3个麦克风通道时,最终系统在-12dB时的相对字错误率降低了55%,在12dB时的相对字错误率降低了43%。
摘要:One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which voice to recognize and which to ignore and it needs to do so in a streaming, low-latency manner. This work presents two multi-microphone speech enhancement algorithms targeted at this scenario. Targeting on-device use-cases, we assume that the algorithm has access to the signal before the hotword, which is referred to as the noise context. First is the Context Aware Beamformer which uses the noise context and detected hotword to determine how to target the desired speaker. The second is an adaptive noise cancellation algorithm called Speech Cleaner which trains a filter using the noise context. It is demonstrated that the two algorithms are complementary in the signal-to-noise ratio conditions under which they work well. We also propose an algorithm to select which one to use based on estimated SNR. When using 3 microphone channels, the final system achieves a relative word error rate reduction of 55% at -12dB, and 43\% at 12dB.


【7】 Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator

标题:使用树约束指针生成器最小化上下文ASR的偏向单词错误

链接:https://arxiv.org/abs/2205.09058

作者:Guangzhi Sun,Chao Zhang,Philip C Woodland
备注:This work has been submitted to the IEEE Transactions on Audio, Speech, and Language Processing for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:上下文知识对于减少高值长尾词的语音识别错误至关重要。本文提出了一种新的树约束指针生成器(TCPGen)组件,该组件允许端到端ASR模型偏向使用外部上下文信息获得的长尾词列表。TCPGen只需在内存使用和计算成本方面花费很小的开销,就可以有效地将数千个偏差词构造成一个符号前缀树,并在树和最终ASR输出之间创建一个神经捷径,以便于识别偏差词。为了增强TCPGen,我们进一步提出了一种新的最小偏差词错误(MBWE)损失,该损失可以直接优化训练期间的偏差词错误,以及测试期间的偏差词驱动语言模型折扣(BLMD)方法。所有上下文ASR系统均在公共Librispeech有声读物语料库和对话状态跟踪挑战(DSTC)数据上进行评估,并从对话系统本体中提取偏差列表。TCPGen实现了一致的单词错误率(WER)降低,这在偏误单词上尤其显著,识别错误率相对降低约40%。MBWE和BLMD进一步提高了TCPGen的有效性,并在偏倚词上实现了更显著的WER降低。TCPGen还实现了对不在音频训练集中的单词的零射门学习,大大减少了偏差列表中的词汇外单词。
摘要:Contextual knowledge is essential for reducing speech recognition errors on high-valued long-tail words. This paper proposes a novel tree-constrained pointer generator (TCPGen) component that enables end-to-end ASR models to bias towards a list of long-tail words obtained using external contextual information. With only a small overhead in memory use and computation cost, TCPGen can structure thousands of biasing words efficiently into a symbolic prefix-tree and creates a neural shortcut between the tree and the final ASR output to facilitate the recognition of the biasing words. To enhance TCPGen, we further propose a novel minimum biasing word error (MBWE) loss that directly optimises biasing word errors during training, along with a biasing-word-driven language model discounting (BLMD) method during the test. All contextual ASR systems were evaluated on the public Librispeech audiobook corpus and the data from the dialogue state tracking challenges (DSTC) with the biasing lists extracted from the dialogue-system ontology. Consistent word error rate (WER) reductions were achieved with TCPGen, which were particularly significant on the biasing words with around 40\% relative reductions in the recognition error rates. MBWE and BLMD further improved the effectiveness of TCPGen and achieved more significant WER reductions on the biasing words. TCPGen also achieved zero-shot learning of words not in the audio training set with large WER reductions on the out-of-vocabulary words in the biasing list.


【8】 Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation

标题:利用伪标签数据改进直接语音到语音翻译

链接:https://arxiv.org/abs/2205.08993

作者:Qianqian Dong,Fengpeng Yue,Tom Ko,Mingxuan Wang,Qibing Bai,Yu Zhang
备注:Submitted to INTERSPEECH 2022
摘要:近年来,直接言语翻译(S2ST)受到了越来越多的关注。由于数据匮乏和复杂的语音到语音映射,这项任务非常具有挑战性。在本文中,我们报告了我们在S2ST方面的最新成就。首先,我们构建了一个S2STTransformer基线,其性能优于原Translatotron。其次,我们通过伪标记利用外部数据,在Fisher英语到西班牙语测试集上获得了一个新的最新结果。事实上,我们利用伪数据结合流行的技术,这些技术在应用于S2ST时并不简单。此外,我们还对句法相似(西班牙语-英语)和遥远(英语-汉语)的语言对进行了评估。我们的实现可在https://github.com/fengpeng-yue/speech-to-speech-translation.
摘要:Direct Speech-to-speech translation (S2ST) has drawn more and more attention recently. The task is very challenging due to data scarcity and complex speech-to-speech mapping. In this paper, we report our recent achievements in S2ST. Firstly, we build a S2ST Transformer baseline which outperforms the original Translatotron. Secondly, we utilize the external data by pseudo-labeling and obtain a new state-of-the-art result on the Fisher English-to-Spanish test set. Indeed, we exploit the pseudo data with a combination of popular techniques which are not trivial when applied to S2ST. Moreover, we evaluate our approach on both syntactically similar (Spanish-English) and distant (English-Chinese) language pairs. Our implementation is available at https://github.com/fengpeng-yue/speech-to-speech-translation.


【9】 Seeing Sounds, Hearing Shapes: a gamified study to evaluate sound-sketches

标题:看得见的声音,听得到的形状:评估声音素描的游戏化研究

链接:https://arxiv.org/abs/2205.08866

作者:Sebastian Löbbers,György Fazekas
备注:Accepted at International Computer Music Conference (ICMC) 2022
摘要:声音-形状关联是听觉和视觉领域之间跨模态关联的一个子集,主要是在将一组精心制作的形状与声音匹配的背景下进行研究。最近的研究探索了人类如何通过自由形式的素描来表现声音,以及如何将图形素描输入用于声音制作。在本文中,通过两次在线展览活动中的82名参与者进行了一项游戏化研究,探讨了通过这些自由形式草图传达声音特征的潜力。结果表明,参与者能够以比随机基线更高的速度识别声音,但似乎很难对细微的音色差异进行视觉编码。
摘要:Sound-shape associations, a subset of cross-modal associations between the auditory and visual domain, have been studied mainly in the context of matching a set of purposefully crafted shapes to sounds. Recent studies have explored how humans represent sound through free-form sketching and how a graphical sketch input could be used for sound production. In this paper, the potential of communicating sound characteristics through these free-form sketches is investigated in a gamified study that was conducted with eighty-two participants at two online exhibition events. The results show that participants managed to recognise sounds at a higher rate than the random baseline would suggest, however it appeared difficult to visually encode nuanced timbral differences.


【10】 Deploying self-supervised learning in the wild for hybrid automatic speech recognition

标题:在野外部署自监督学习用于混合自动语音识别

链接:https://arxiv.org/abs/2205.08598

作者:Mostafa Karimi,Changliang Liu,Kenichi Kumatani,Yao Qian,Tianyu Wu,Jian Wu
摘要:自监督学习(SSL)方法已被证明在自动语音识别(ASR)中非常成功。据报道,这些巨大的改进主要基于高度管理的数据集,如用于非流式端到端ASR模型的LibriSpeech。然而,SSL的关键特性是可用于任何未翻译的音频数据。在本文中,我们将全面探讨如何在SSL中利用未固化的音频数据,从数据预处理到部署流式混合ASR模型。更具体地说,我们展示了(1)数据预处理管道中音频事件检测(AED)模型的效果(2)对优化器选择和学习率调度的分析(3)最近开发的对比损失的比较,(4)各种预训练策略的比较,如域内和域外预训练数据的利用,单语言与多语言预训练数据,多头多语言SSL与单头多语言SSL,监督预训练与SSL。实验结果表明,与所有可选的域外预训练策略相比,使用域内未固化数据的SSL预训练可以获得更好的性能。
摘要:Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly curated datasets such as LibriSpeech for non-streaming End-to-End ASR models. However, the pivotal characteristics of SSL is to be utilized for any untranscribed audio data. In this paper, we provide a full exploration on how to utilize uncurated audio data in SSL from data pre-processing to deploying an streaming hybrid ASR model. More specifically, we present (1) the effect of Audio Event Detection (AED) model in data pre-processing pipeline (2) analysis on choosing optimizer and learning rate scheduling (3) comparison of recently developed contrastive losses, (4) comparison of various pre-training strategies such as utilization of in-domain versus out-domain pre-training data, monolingual versus multilingual pre-training data, multi-head multilingual SSL versus single-head multilingual SSL and supervised pre-training versus SSL. The experimental results show that SSL pre-training with in-domain uncurated data can achieve better performance in comparison to all the alternative out-domain pre-training strategies.


【11】 The Power of Reuse: A Multi-Scale Transformer Model for Structural Dynamic Segmentation in Symbolic Music Generation

标题:重用的力量:符号音乐生成中结构动态分割的多尺度变换模型

链接:https://arxiv.org/abs/2205.08579

作者:Guowei Wu,Shipei Liu,Xiaoya Fan
摘要:符号音乐的生成依赖于生成模型的上下文表示能力,其中最流行的方法是基于转换器的模型。不仅如此,长期语境的学习还与音乐结构的动态分割有关,即介绍、韵文和合唱,这目前被研究界所忽视。在本文中,我们提出了一种多尺度变换器,它使用粗解码器和细解码器分别在全局和部分级别对上下文进行建模。具体来说,我们设计了一个片段范围定位层,将音乐切分为多个部分,这些部分随后用于预训练精细解码器。然后,我们设计了一个音乐风格规范化层,将风格信息从原始部分传输到生成的部分,以实现音乐风格的一致性。生成的部分在聚合层中进行组合,并由粗解码器进行微调。我们的模型在两个开放的MIDI数据集上进行了评估,实验表明,我们的模型优于当代最好的符号音乐生成模型。更令人兴奋的是,视觉评估表明,我们的模型在旋律重用方面具有优势,从而产生了更逼真的音乐。
摘要:Symbolic Music Generation relies on the contextual representation capabilities of the generative model, where the most prevalent approach is the Transformer-based model. Not only that, the learning of long-term context is also related to the dynamic segmentation of musical structures, i.e. intro, verse and chorus, which is currently overlooked by the research community. In this paper, we propose a multi-scale Transformer, which uses coarse-decoder and fine-decoders to model the contexts at the global and section-level, respectively. Concretely, we designed a Fragment Scope Localization layer to syncopate the music into sections, which were later used to pre-train fine-decoders. After that, we designed a Music Style Normalization layer to transfer the style information from the original sections to the generated sections to achieve consistency in music style. The generated sections are combined in the aggregation layer and fine-tuned by the coarse decoder. Our model is evaluated on two open MIDI datasets, and experiments show that our model outperforms the best contemporary symbolic music generative models. More excitingly, visual evaluation shows that our model is superior in melody reuse, resulting in more realistic music.


机器翻译,仅供参考