今天跟大家分享一篇语音相关的论文合集:cs.SD语音14篇,eess.AS音频处理19篇。本文经arXiv每日学术速递授权转载
【1】 AVATAR: Unconstrained Audiovisual Speech Recognition
标题:阿凡达:不受限制的视听语音识别
链接:https://arxiv.org/abs/2206.07684
作者:Valentin Gabeur,Paul Hongsuck Seo,Arsha Nagrani,Chen Sun,Karteek Alahari,Cordelia Schmid机构:Inria†, Google Research摘要:视听自动语音识别(AV-ASR)是ASR的一个扩展,它包含视觉线索,通常来自说话者的嘴的运动。与只关注嘴唇运动的作品不同,我们研究了整个视觉框架(视觉动作、物体、背景等)的贡献。这对于无约束视频尤其有用,因为在这些视频中,说话者不一定可见。为了解决这一问题,我们提出了一种新的序列到序列的视听ASR转换器(AVATAR),该转换器通过频谱图和全帧RGB进行端到端的训练。为了防止音频流主导训练,我们提出了不同的单词掩蔽策略,从而鼓励我们的模型关注视频流。我们展示了视觉模式对How2 AV-ASR基准测试的贡献,尤其是在存在模拟噪声的情况下,并表明我们的模型大大优于所有其他先前的工作。最后,我们还为AV-ASR创建了一个名为VisSpeech的新的、真实的测试平台,该平台展示了视觉模态在具有挑战性的音频条件下的贡献。摘要:Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of entire visual frames (visual actions, objects, background etc.). This is particularly useful for unconstrained videos, where the speaker is not necessarily visible. To solve this task, we propose a new sequence-to-sequence AudioVisual ASR TrAnsformeR (AVATAR) which is trained end-to-end from spectrograms and full-frame RGB. To prevent the audio stream from dominating training, we propose different word-masking strategies, thereby encouraging our model to pay attention to the visual stream. We demonstrate the contribution of the visual modality on the How2 AV-ASR benchmark, especially in the presence of simulated noise, and show that our model outperforms all other prior work by a large margin. Finally, we also create a new, real-world test bed for AV-ASR called VisSpeech, which demonstrates the contribution of the visual modality under challenging audio conditions.
【2】 Exploring Capabilities of Monolingual Audio Transformers using Large Datasets in Automatic Speech Recognition of Czech标题:利用大数据集探索单语音频转换器在捷克语自动语音识别中的能力作者:Jan Lehečka,Jan Švec,Aleš Pražák,Josef V. Psutka机构:Department of Cybernetics, University of West Bohemia Pilsen, Czech Republic备注:to be published in Proceedings of INTERSPEECH 2022摘要:在本文中,我们介绍了我们在从一个包含8万多小时未标记语音的大型数据集预训练捷克单语音频转换器方面的进展,并随后结合域内数据和近6000小时的域外转录语音对自动语音识别任务的模型进行微调。我们正在展示一系列实验,其中包括在两个公共数据集(CommonVoice和VoxPopuli)和MALACH项目中一个极具挑战性的数据集上评估的各种微调设置。我们的结果表明,单语Wav2Vec 2.0模型是健壮的ASR系统,它可以利用大型标记和未标记数据集,并成功地与最先进的LVCSR系统竞争。此外,当目标ASR任务没有可用的训练数据时,Wav2Vec模型被证明是很好的零炮学习者。摘要:In this paper, we present our progress in pretraining Czech monolingual audio transformers from a large dataset containing more than 80 thousand hours of unlabeled speech, and subsequently fine-tuning the model on automatic speech recognition tasks using a combination of in-domain data and almost 6 thousand hours of out-of-domain transcribed speech. We are presenting a large palette of experiments with various fine-tuning setups evaluated on two public datasets (CommonVoice and VoxPopuli) and one extremely challenging dataset from the MALACH project. Our results show that monolingual Wav2Vec 2.0 models are robust ASR systems, which can take advantage of large labeled and unlabeled datasets and successfully compete with state-of-the-art LVCSR systems. Moreover, Wav2Vec models proved to be good zero-shot learners when no training data are available for the target ASR task.
【3】 Investigating Multi-Feature Selection and Ensembling for Audio Classification
标题:音频分类中的多特征选择与集成方法研究
链接:https://arxiv.org/abs/2206.07511
作者:Muhammad Turab,Teerath Kumar,Malika Bendechache,Takfarinas Saber机构:Mehran University of Engineering and Technology, Jamshoro, Pakistan., ADAPT – Science Foundation Ireland Research Centre, CRT AI, School of Computing, Dublin City University, Dublin, Ireland, Lero – the Irish Software Research Centre摘要:深度学习(DL)算法在不同领域表现出令人印象深刻的性能。其中,由于一些有趣的模式,尤其是在音频数据的分类方面,音频在过去几十年中吸引了许多研究人员。为了获得更好的音频分类性能,特征选择和组合起着关键作用,因为它们有可能影响任何DL模型的性能。要调查此角色,我们对具有各种最先进音频特征(即,Mel频谱图、Mel频率倒谱系数和过零率)的多个尖端DL模型(即卷积神经网络、EfficientNet、MobileNet、超级向量机和多感知器)的性能进行了广泛的评估,这些模型可以是独立的,也可以是组合的(即通过融合)三种不同的数据集(即自由语音数字数据集、音频乌尔都语数字数据集和音频古吉拉特语数字数据集)。总体而言,结果表明特征选择取决于数据集和模型。然而,特征组合应仅限于单独使用时已达到良好性能的特征(即,主要是Mel频谱图、Mel频率倒谱系数)。无论我们选择何种DL模型,这种特性组合/融合都使我们的表现优于之前最先进的结果。摘要:Deep Learning (DL) algorithms have shown impressive performance in diverse domains. Among them, audio has attracted many researchers over the last couple of decades due to some interesting patterns--particularly in classification of audio data. For better performance of audio classification, feature selection and combination play a key role as they have the potential to make or break the performance of any DL model. To investigate this role, we conduct an extensive evaluation of the performance of several cutting-edge DL models (i.e., Convolutional Neural Network, EfficientNet, MobileNet, Supper Vector Machine and Multi-Perceptron) with various state-of-the-art audio features (i.e., Mel Spectrogram, Mel Frequency Cepstral Coefficients, and Zero Crossing Rate) either independently or as a combination (i.e., through ensembling) on three different datasets (i.e., Free Spoken Digits Dataset, Audio Urdu Digits Dataset, and Audio Gujarati Digits Dataset). Overall, results suggest feature selection depends on both the dataset and the model. However, feature combinations should be restricted to the only features that already achieve good performances when used individually (i.e., mostly Mel Spectrogram, Mel Frequency Cepstral Coefficients). Such feature combination/ensembling enabled us to outperform the previous state-of-the-art results irrespective of our choice of DL model.
【4】 VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection
标题:VisageSynTalk:基于语音人脸特征选择的看不见说话人视频到语音合成
链接:https://arxiv.org/abs/2206.07458
作者:Joanna Hong,Minsu Kim,Yong Man Ro机构:Image and Video Systems Lab, School of Electrical Engineering, KAIST, South Korea备注:Submitted to ECCV 2022摘要:这项工作的目标是从无声的人脸视频中重建语音。最近的研究表明,从无声交谈的人脸视频中合成语音的性能令人印象深刻。然而,他们没有明确考虑不同说话人的不同身份特征,这给视频到语音合成带来了挑战,这在看不见的说话人设置中变得更加重要。与之前的方法不同,我们的方法是从给定的无声对话人脸视频中分离语音内容和面部风格。通过引导模型独立地关注这两种表示的建模,我们可以从模型中获得高可懂度的语音,即使给定了一个看不见的对象的输入视频。为此,我们引入了语音视觉选择模块,该模块将语音内容和说话人身份与输入视频的视觉特征分离。将分离的表示合并到一起,通过基于外观样式的合成器合成语音,该合成器通过在保持语音内容的同时涂抹外观样式来生成语音。因此,所提出的框架带来了合成包含正确内容的语音的优势,即使给出了一个看不见的主题的无声对话人脸视频。我们在网格、TCD-TIMIT志愿者和LRW数据集上验证了所提框架的有效性。合成语音可以在补充材料中听到。摘要:The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on varying identity characteristics of different speakers, which place a challenge in the video-to-speech synthesis, and this becomes more critical in unseen-speaker settings. Distinct from the previous methods, our approach is to separate the speech content and the visage-style from a given silent talking face video. By guiding the model to independently focus on modeling the two representations, we can obtain the speech of high intelligibility from the model even when the input video of an unseen subject is given. To this end, we introduce speech-visage selection module that separates the speech content and the speaker identity from the visual features of the input video. The disentangled representations are jointly incorporated to synthesize speech through visage-style based synthesizer which generates speech by coating the visage-styles while maintaining the speech content. Thus, the proposed framework brings the advantage of synthesizing the speech containing the right content even when the silent talking face video of an unseen subject is given. We validate the effectiveness of the proposed framework on the GRID, TCD-TIMIT volunteer, and LRW datasets. The synthesized speech can be heard in supplementary materials.
【5】 NatiQ: An End-to-end Text-to-Speech System for Arabic
标题:NATIQ:一个面向阿拉伯语的端到端文语转换系统
链接:https://arxiv.org/abs/2206.07373
作者:Ahmed Abdelali,Nadir Durrani,Cenk Demiroglu,Fahim Dalvi,Hamdy Mubarak,Kareem Darwish机构:Qatar Computing Research Institute - Hamad Bin Khalifa University, Doha, Qatar, ¨Ozye˘gin University, Istanbul, T¨urkiye摘要:NatiQ是阿拉伯语的端到端文本到语音系统。我们的语音合成器使用编码器-解码器架构。我们使用基于tacotron的模型(tacotron-1和tacotron-2)和快速变换器模型从字符生成mel光谱图。我们将Tacotron1与WaveRNN声码器连接,Tacotron2与WaveGlow声码器连接,ESPnet transformer与并行wavegan声码器连接,以从频谱图合成波形。我们使用了两种声音的内部语音数据:1)中性男性“Hamza”(讲述一般内容和新闻),2)富有表现力的女性“Amina”(讲述儿童故事书)来训练我们的模型。我们的最佳系统对Amina和Hamza的平均意见得分(MOS)分别为4.21和4.40。使用单词和字符错误率(WER和CER)以及实时因素测量的响应时间对系统进行客观评估,有利于端到端架构ESPnet。NatiQ演示可在线访问https://tts.qcri.org摘要:NatiQ is end-to-end text-to-speech system for Arabic. Our speech synthesizer uses an encoder-decoder architecture with attention. We used both tacotron-based models (tacotron-1 and tacotron-2) and the faster transformer model for generating mel-spectrograms from characters. We concatenated Tacotron1 with the WaveRNN vocoder, Tacotron2 with the WaveGlow vocoder and ESPnet transformer with the parallel wavegan vocoder to synthesize waveforms from the spectrograms. We used in-house speech data for two voices: 1) neutral male "Hamza"- narrating general content and news, and 2) expressive female "Amina"- narrating children story books to train our models. Our best systems achieve an average Mean Opinion Score (MOS) of 4.21 and 4.40 for Amina and Hamza respectively. The objective evaluation of the systems using word and character error rate (WER and CER) as well as the response time measured by real-time factor favored the end-to-end architecture ESPnet. NatiQ demo is available on-line at https://tts.qcri.org
【6】 On the Use of Deep Mask Estimation Module for Neural Source Separation Systems
标题:深度掩码估计模块在神经源分离系统中的应用
链接:https://arxiv.org/abs/2206.07347
作者:Kai Li,Xiaolin Hu,Yi Luo机构:†Department of Computer Science and Technology, BNRist, Tsinghua University, China, ‡Tencent AI Lab, Shenzhen, China备注:Accepted by Interspeech 2022摘要:大多数最近的神经源分离系统依赖于基于掩蔽的管道,其中一组乘法掩蔽是从输入混合的信号表示中估计并应用于该信号表示的。在几乎所有的网络体系结构中,这种掩码的估计都是由一个单层和一个可选的非线性激活函数完成的。然而,最近的文献研究了深掩模估计模块的使用,并观察到与浅掩模估计模块相比的性能改进。在本文中,我们通过将深度掩码估计模块与最近提出的无监督源分离方法相连接,分析了这种深度掩码估计模块的作用,并从经验上表明,深度掩码估计模块是所谓的过分离分组范式与传统浅层掩码估计层的有效近似。摘要:Most of the recent neural source separation systems rely on a masking-based pipeline where a set of multiplicative masks are estimated from and applied to a signal representation of the input mixture. The estimation of such masks, in almost all network architectures, is done by a single layer followed by an optional nonlinear activation function. However, recent literatures have investigated the use of a deep mask estimation module and observed performance improvement compared to a shallow mask estimation module. In this paper, we analyze the role of such deeper mask estimation module by connecting it to a recently proposed unsupervised source separation method, and empirically show that the deep mask estimation module is an efficient approximation of the so-called overseparation-grouping paradigm with the conventional shallow mask estimation layers.
【7】 On the Design and Training Strategies for RNN-based Online Neural Speech Separation Systems
标题:基于RNN的在线神经语音分离系统设计与训练策略研究
链接:https://arxiv.org/abs/2206.07340
机构:†Department of Computer Science and Technology, BNRist, Tsinghua University, China, ‡Tencent AI Lab, Shenzhen, China摘要:虽然离线神经语音分离系统的性能因新型神经网络体系结构的发展而得到了极大的提高,但系统与其在线变体之间通常存在不可避免的性能差距。在本文中,我们研究了如何将基于RNN的离线神经语音分离系统转变为在线神经语音分离系统,同时缓解性能下降。我们在双向RNN层中分解或重组前向和后向RNN层,以形成在线路径和离线路径,从而使模型能够使用相同的模型参数集执行在线和离线处理。我们进一步介绍了两种通过预训练离线模型或多任务训练目标改进在线模型的训练策略。实验结果表明,与从头开始训练的在线模型相比,所提出的层分解重组方案和训练策略可以有效地缓解两种基于RNN的离线分离模型及其在线变体之间的性能差距。摘要:While the performance of offline neural speech separation systems has been greatly advanced by the recent development of novel neural network architectures, there is typically an inevitable performance gap between the systems and their online variants. In this paper, we investigate how RNN-based offline neural speech separation systems can be changed into their online counterparts while mitigating the performance degradation. We decompose or reorganize the forward and backward RNN layers in a bidirectional RNN layer to form an online path and an offline path, which enables the model to perform both online and offline processing with a same set of model parameters. We further introduce two training strategies for improving the online model via either a pretrained offline model or a multitask training objective. Experiment results show that compared to the online models that are trained from scratch, the proposed layer decomposition and reorganization schemes and training strategies can effectively mitigate the performance gap between two RNN-based offline separation models and their online variants.
【8】 FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
标题:FRCRN:基于频率递归的改进特征表示用于单声道语音增强
链接:https://arxiv.org/abs/2206.07293
作者:Shengkui Zhao,Bin Ma,Karn N. Watcharasupat,Woon-Seng Gan机构:Alibaba Group, School of Electrical and Electronic Engineering, Nanyang Technological University (NTU), Singapore备注:The paper has been accepted by ICASSP 2022. 5 pages, 2 figures, 5 tables摘要:卷积递归网络(CRN)集成了卷积编解码(CED)结构和递归结构,在单耳语音增强方面取得了良好的性能。然而,由于CED卷积中的感受野有限,跨频率上下文的特征表示受到高度限制。在本文中,我们提出了一种卷积循环编码器-解码器(CRED)结构来增强沿频率轴的特征表示。CRED在每次卷积后沿频率轴对三维卷积特征映射应用频率递归,因此,它能够捕捉长距离的频率相关性并增强语音输入的特征表示。所提出的频率递归是利用前馈顺序存储网络(FSMN)有效实现的。除了CRED之外,我们在编码器和解码器之间插入两个堆叠的FSMN层,以模拟进一步的时间动态。我们将该框架命名为频率递归CRN(FRCRN)。我们设计FRCRN在复值域预测复理想比掩模(cIRM),并利用时频域和时域损耗优化FRCRN。我们提出的方法在宽带基准数据集上取得了最先进的性能,并在ICASSP 2022深度噪声抑制(DNS)挑战赛中,在平均意见得分(MOS)和单词准确性(WAcc)方面取得了实时全频段跟踪的第二名。摘要:Convolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across frequency context is highly constrained due to limited receptive fields in the convolutions of CED. In this paper, we propose a convolutional recurrent encoder-decoder (CRED) structure to boost feature representation along the frequency axis. The CRED applies frequency recurrence on 3D convolutional feature maps along the frequency axis following each convolution, therefore, it is capable of catching long-range frequency correlations and enhancing feature representations of speech inputs. The proposed frequency recurrence is realized efficiently using a feedforward sequential memory network (FSMN). Besides the CRED, we insert two stacked FSMN layers between the encoder and the decoder to model further temporal dynamics. We name the proposed framework as Frequency Recurrent CRN (FRCRN). We design FRCRN to predict complex Ideal Ratio Mask (cIRM) in complex-valued domain and optimize FRCRN using both time-frequency-domain and time-domain losses. Our proposed approach achieved state-of-the-art performance on wideband benchmark datasets and achieved 2nd place for the real-time fullband track in terms of Mean Opinion Score (MOS) and Word Accuracy (WAcc) in the ICASSP 2022 Deep Noise Suppression (DNS) challenge.
【9】 Text-Aware End-to-end Mispronunciation Detection and Diagnosis
标题:文本感知的端到端发音错误检测与诊断
链接:https://arxiv.org/abs/2206.07289
作者:Linkai Peng,Yingming Gao,Binghuai Lin,Dengfeng Ke,Yanlu Xie,Jinsong Zhang机构:School of Information Sciences, Beijing Language and Culture University, Beijing, China, Smart Platform Product Department,Tencent Technology Co., Ltd, Beijing, China备注:Rejected by Interspeech2022摘要:发音错误检测与诊断(MDD)技术是计算机辅助发音训练系统(CAPT)的关键组成部分。在评估受限语音的发音质量方面,给定的抄本可以起到教师的作用。传统的方法已经充分利用了先前的文本来构建模型或提高系统性能,例如强制对齐和扩展识别网络。最近,一些基于端到端的方法试图将先前的文本纳入到模型训练中,并初步证明了其有效性。然而,以往的研究大多考虑将原始注意机制应用于音频表征与文本表征的融合,而没有考虑可能的文本发音失配。在本文中,我们提出了一种门控策略,该策略在抑制无关文本信息的同时,更加重视相关音频特征。此外,在给定转录本的情况下,我们设计了一个额外的对比损失,以缩小音位识别的学习目标与MDD之间的差距。我们使用两个公开的数据集(TIMIT和L2 Arctic)进行了实验,与基线相比,我们的最佳模型将F1得分从57.51美元提高到61.75美元。此外,我们还对门控机制和对比学习对MDD的有效性进行了详细的分析。摘要:Mispronunciation detection and diagnosis (MDD) technology is a key component of computer-assisted pronunciation training system (CAPT). In the field of assessing the pronunciation quality of constrained speech, the given transcriptions can play the role of a teacher. Conventional methods have fully utilized the prior texts for the model construction or improving the system performance, e.g. forced-alignment and extended recognition networks. Recently, some end-to-end based methods attempt to incorporate the prior texts into model training and preliminarily show the effectiveness. However, previous studies mostly consider applying raw attention mechanism to fuse audio representations with text representations, without taking possible text-pronunciation mismatch into account. In this paper, we present a gating strategy that assigns more importance to the relevant audio features while suppressing irrelevant text information. Moreover, given the transcriptions, we design an extra contrastive loss to reduce the gap between the learning objective of phoneme recognition and MDD. We conducted experiments using two publicly available datasets (TIMIT and L2-Arctic) and our best model improved the F1 score from $57.51\%$ to $61.75\%$ compared to the baselines. Besides, we provide a detailed analysis to shed light on the effectiveness of gating mechanism and contrastive learning on MDD.
【10】 Streaming non-autoregressive model for any-to-many voice conversion
标题:用于任意对多语音转换的流非自回归模型
链接:https://arxiv.org/abs/2206.07288
作者:Ziyi Chen,Haoran Miao,Pengyuan Zhang机构:Key Laboratory of Speech Acoustics & Content Understanding, Institute of Acoustics, Chinese, University of Chinese Academy of Sciences, Beijing, China摘要:语音转换模型已经发展了几十年,当前的主流研究集中在非流式语音转换上。然而,流式语音转换比非流式语音转换更适合实际应用场景。在本文中,我们提出了一种基于完全非自回归模型的流式任意多语音转换,包括基于流式变换器的声学模型和流式声码器。基于流Transformer的声学模型由基于流端到端自动语音识别模型的预训练编码器和基于FastSpeech块的解码器组成。流式声码器采用伪正交镜像滤波器组和因果卷积设计,用于流式任务。实验结果表明,该方法在延迟和转换质量方面都取得了显著的性能,并且可以在CPU和GPU上实现实时性。摘要:Voice conversion models have developed for decades, and current mainstream research focuses on non-streaming voice conversion. However, streaming voice conversion is more suitable for practical application scenarios than non-streaming voice conversion. In this paper, we propose a streaming any-to-many voice conversion based on fully non-autoregressive model, which includes a streaming transformer based acoustic model and a streaming vocoder. Streaming transformer based acoustic model is composed of a pre-trained encoder from streaming end-to-end based automatic speech recognition model and a decoder modified on FastSpeech blocks. Streaming vocoder is designed for streaming task with pseudo quadrature mirror filter bank and causal convolution. Experimental results show that the proposed method achieves significant performance both in latency and conversion quality and can be real-time on CPU and GPU.
【11】 Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning
标题:基于数据驱动深度学习的可见语音和未见语音情感强度精确评估
链接:https://arxiv.org/abs/2206.07229
作者:Rui Liu,Berrak Sisman,Björn Schuller,Guanglai Gao,Haizhou Li机构:Inner Mongolia University, China , The Chinese University of Hong Kong, Shenzhen, China, National University of Singapore , Singapore University of Technology and Design, Imperial College London, United Kingdom备注:To appear in INTERSPEECH 2022. 5 pages, 4 figures. Substantial text overlap with arXiv:2110.03156摘要:在情感文本到语音和语音转换等应用中,需要对语音进行情感分类并评估情感强度。提出了基于支持向量机(SVM)的情感属性排序函数来预测情感语音语料库的情感强度。然而,经过训练的排序函数并没有推广到新的领域,这限制了它的应用范围,尤其是对于域外或看不见的语音。在本文中,我们提出了一种数据驱动的深度学习模型,即StrengthNet,以改进对可见和不可见语音的情感强度评估的泛化。这是通过融合来自不同领域的情感数据来实现的。我们遵循一个多任务学习网络架构,该架构包括一个声学编码器、一个强度预测器和一个辅助情绪预测器。实验表明,所提出的StrengthNet的预测情感强度与可见和不可见语音的地面真实度得分高度相关。我们在以下位置发布源代码:https://github.com/ttslr/StrengthNet.摘要:Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet.
【12】 Frequency-centroid features for word recognition of non-native English speakers
标题:非英语母语者单词识别的频率中心特征
链接:https://arxiv.org/abs/2206.07176
作者:Pierre Berjon,Rajib Sharma,Avishek Nag,Soumyabrata Dev机构:Department de Sciences du Numérique, INP-ENSEEIHT, Toulouse, France, Department of DSIS, Indian Institute of Information Technology Dharwad, India, School of Electrical and Electronic Engineering, University College Dublin, Ireland备注:Published in IEEE Irish Signals & Systems Conference (ISSC), 2022摘要:这项工作的目的是研究互补特征,以帮助不同母语的非英语母语者在封闭、有限集合词识别任务中使用典型的Mel频率倒谱系数(MFCC)。与MFCC不同的是,MFCC源自语音信号的频谱能量,建议的频率质心(FCs)封装了语音频谱不同频带的频谱中心,频带由Mel滤波器组定义。这些特征与MFCC相结合,可以提高英语单词识别的相对性能,尤其是在各种噪声条件下。采用两级卷积神经网络(CNN)对英语中带有阿拉伯语、法语和西班牙语口音的单词的特征进行建模。摘要:The objective of this work is to investigate complementary features which can aid the quintessential Mel frequency cepstral coefficients (MFCCs) in the task of closed, limited set word recognition for non-native English speakers of different mother-tongues. Unlike the MFCCs, which are derived from the spectral energy of the speech signal, the proposed frequency-centroids (FCs) encapsulate the spectral centres of the different bands of the speech spectrum, with the bands defined by the Mel filterbank. These features, in combination with the MFCCs, are observed to provide relative performance improvement in English word recognition, particularly under varied noisy conditions. A two-stage Convolution Neural Network (CNN) is used to model the features of the English words uttered with Arabic, French and Spanish accents.
【13】 End-to-End Voice Conversion with Information Perturbation
标题:带信息扰动的端到端语音转换
链接:https://arxiv.org/abs/2206.07569
作者:Qicong Xie,Shan Yang,Yi Lei,Lei Xie,Dan Su机构:Northwestern Polytechnical University, Xi’an, China, Tencent AI Lab, China摘要:语音转换的理想目标是将源说话人的语音转换为与目标说话人一样自然的声音,同时保持源语音的语言内容和韵律。然而,现有的方法不足以在转换语音中实现全面的源韵律转换和目标说话人音色保持,并且由于声学模型和声码器之间的不匹配,转换语音的质量也不令人满意。在本文中,我们利用信息扰动方面的最新进展,提出了一种完全端到端的方法来进行高质量的语音转换。我们首先采用信息扰动去除源语音中与说话人相关的信息,以分离说话人音色和语言内容,然后通过内容编码器对语言信息进行建模。为了更好地将源语音的韵律传递到目标语音,我们特别介绍了一种与说话人相关的基音编码器,它可以保持源说话人的一般基音模式,同时灵活地修改生成语音的基音强度。最后,通过连续的说话人空间建模建立一次语音转换。实验结果表明,所提出的端到端方法在可懂度、自然度和说话人相似度方面明显优于现有的模型。摘要:The ideal goal of voice conversion is to convert the source speaker's speech to sound naturally like the target speaker while maintaining the linguistic content and the prosody of the source speech. However, current approaches are insufficient to achieve comprehensive source prosody transfer and target speaker timbre preservation in the converted speech, and the quality of the converted speech is also unsatisfied due to the mismatch between the acoustic model and the vocoder. In this paper, we leverage the recent advances in information perturbation and propose a fully end-to-end approach to conduct high-quality voice conversion. We first adopt information perturbation to remove speaker-related information in the source speech to disentangle speaker timbre and linguistic content and thus the linguistic information is subsequently modeled by a content encoder. To better transfer the prosody of the source speech to the target, we particularly introduce a speaker-related pitch encoder which can maintain the general pitch pattern of the source speaker while flexibly modifying the pitch intensity of the generated speech. Finally, one-shot voice conversion is set up through continuous speaker space modeling. Experimental results indicate that the proposed end-to-end approach significantly outperforms the state-of-the-art models in terms of intelligibility, naturalness, and speaker similarity.
【14】 Residual Language Model for End-to-end Speech Recognition
标题:端到端语音识别的残差语言模型
链接:https://arxiv.org/abs/2206.07430
作者:Emiru Tsunoo,Yosuke Kashiwagi,Chaitanya Narisetty,Shinji Watanabe机构:Sony Group Corporation, Japan, Carnegie Mellon University, USA备注:Accepted for Interspeech2022摘要:端到端自动语音识别尽管使用了大量成对的音频-文本数据进行训练,但仍无法适应未知的目标域语音。最近的研究估计,该模型的语言偏向为内部语言模型(LM)。为了有效地适应目标域,在推理过程中从后验域中减去内部LM,并与外部目标域LM融合。然而,这种融合使推理复杂化,内部LM的估计可能并不总是准确的。在本文中,我们提出了一种简单的外部LM融合域自适应方法,该方法在训练过程中考虑了内部LM估计。我们直接对外部和内部LMs的残差因子建模,即残差LM。为了稳定地训练残差LM,我们提出对估计的内部LM进行平滑处理,并结合交叉熵和均方误差损失对其进行优化,这考虑了内部LM在目标域数据中的统计行为。我们通过实验证实,在大多数跨域和域内场景中,所提出的残差LM的性能优于内部LM估计。摘要:End-to-end automatic speech recognition suffers from adaptation to unknown target domain speech despite being trained with a large amount of paired audio--text data. Recent studies estimate a linguistic bias of the model as the internal language model (LM). To effectively adapt to the target domain, the internal LM is subtracted from the posterior during inference and fused with an external target-domain LM. However, this fusion complicates the inference and the estimation of the internal LM may not always be accurate. In this paper, we propose a simple external LM fusion method for domain adaptation, which considers the internal LM estimation in its training. We directly model the residual factor of the external and internal LMs, namely the residual LM. To stably train the residual LM, we propose smoothing the estimated internal LM and optimizing it with a combination of cross-entropy and mean-squared-error losses, which consider the statistical behaviors of the internal LM in the target domain data. We experimentally confirmed that the proposed residual LM performs better than the internal LM estimation in most of the cross-domain and intra-domain scenarios.
【1】 End-to-End Voice Conversion with Information Perturbation标题:带信息扰动的端到端语音转换
链接:https://arxiv.org/abs/2206.07569
作者:Qicong Xie,Shan Yang,Yi Lei,Lei Xie,Dan Su机构:Northwestern Polytechnical University, Xi’an, China, Tencent AI Lab, China摘要:语音转换的理想目标是将源说话人的语音转换为与目标说话人一样自然的声音,同时保持源语音的语言内容和韵律。然而,现有的方法不足以在转换语音中实现全面的源韵律转换和目标说话人音色保持,并且由于声学模型和声码器之间的不匹配,转换语音的质量也不令人满意。在本文中,我们利用信息扰动方面的最新进展,提出了一种完全端到端的方法来进行高质量的语音转换。我们首先采用信息扰动去除源语音中与说话人相关的信息,以分离说话人音色和语言内容,然后通过内容编码器对语言信息进行建模。为了更好地将源语音的韵律传递到目标语音,我们特别介绍了一种与说话人相关的基音编码器,它可以保持源说话人的一般基音模式,同时灵活地修改生成语音的基音强度。最后,通过连续的说话人空间建模建立一次语音转换。实验结果表明,所提出的端到端方法在可懂度、自然度和说话人相似度方面明显优于现有的模型。摘要:The ideal goal of voice conversion is to convert the source speaker's speech to sound naturally like the target speaker while maintaining the linguistic content and the prosody of the source speech. However, current approaches are insufficient to achieve comprehensive source prosody transfer and target speaker timbre preservation in the converted speech, and the quality of the converted speech is also unsatisfied due to the mismatch between the acoustic model and the vocoder. In this paper, we leverage the recent advances in information perturbation and propose a fully end-to-end approach to conduct high-quality voice conversion. We first adopt information perturbation to remove speaker-related information in the source speech to disentangle speaker timbre and linguistic content and thus the linguistic information is subsequently modeled by a content encoder. To better transfer the prosody of the source speech to the target, we particularly introduce a speaker-related pitch encoder which can maintain the general pitch pattern of the source speaker while flexibly modifying the pitch intensity of the generated speech. Finally, one-shot voice conversion is set up through continuous speaker space modeling. Experimental results indicate that the proposed end-to-end approach significantly outperforms the state-of-the-art models in terms of intelligibility, naturalness, and speaker similarity.
【2】 Learnable Frequency Filters for Speech Feature Extraction in Speaker Verification
标题:用于说话人确认中语音特征提取的可学习频率滤波器
链接:https://arxiv.org/abs/2206.07563
作者:Jingyu Li,Yusheng Tian,Tan Lee机构:Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong摘要:Mel尺度谱特征用于语音信号的各种识别和分类任务。没有理由期望这些功能对于所有不同的任务都是最佳的,包括说话人验证(SV)。本文描述了一种可学习的前端特征提取模型。该模型包括一组用于变换傅里叶光谱的滤波器。定义这些过滤器的模型参数经过端到端的训练,并专门针对说话人验证任务进行优化。与标准Mel尺度滤波器组相比,滤波器的带宽和中心频率可调。实验结果表明,与传统的Mel尺度谱特征相比,可学习声学前端的应用提高了说话人验证性能。对学习滤波器参数的分析表明,窄带信息有利于SV系统的性能。该模型在性能和计算成本之间取得了很好的平衡。在资源受限的计算环境中,该模型明显优于基于CNN的可学习前端。在不同的嵌入提取模型和数据集上,验证了该模型的泛化能力。摘要:Mel-scale spectrum features are used in various recognition and classification tasks on speech signals. There is no reason to expect that these features are optimal for all different tasks, including speaker verification (SV). This paper describes a learnable front-end feature extraction model. The model comprises a group of filters to transform the Fourier spectrum. Model parameters that define these filters are trained end-to-end and optimized specifically for the task of speaker verification. Compared to the standard Mel-scale filter-bank, the filters' bandwidths and center frequencies are adjustable. Experimental results show that applying the learnable acoustic front-end improves speaker verification performance over conventional Mel-scale spectrum features. Analysis on the learned filter parameters suggests that narrow-band information benefits the SV system performance. The proposed model achieves a good balance between performance and computation cost. In resource-constrained computation settings, the model significantly outperforms CNN-based learnable front-ends. The generalization ability of the proposed model is also demonstrated on different embedding extraction models and datasets.
【3】 EDITnet: A Lightweight Network for Unsupervised Domain Adaptation in Speaker Verification
标题:EDITnet:说话人确认中的轻量级无监督域自适应网络
链接:https://arxiv.org/abs/2206.07548
作者:Jingyu Li,Wei Liu,Tan Lee机构:Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong备注:Accepted by Interspeech2022摘要:在对不同语言的语音数据应用说话人验证系统时,语言不匹配导致的性能下降是一个常见的问题。本文提出了一种域传输网络EDITnet,以缓解说话人嵌入中的语言不匹配问题,而不需要说话人标签。网络利用条件变分自动编码器将嵌入从目标域传输到源域。为了增加不同说话人嵌入信息之间的余弦距离,对所转移的嵌入信息采用了自监督学习策略。在EDITnet的训练过程中,嵌入抽取模型是固定的,无需微调,这使得训练效率高,成本低。在Voxceleb和CN-Celeb上的实验表明,使用ECPA-TDNN512,EDITnet传输的嵌入比未传输的嵌入性能高出约30%。使用其他嵌入提取模型(例如TDNN、SE-ResNet34)也可以实现性能改进。摘要:Performance degradation caused by language mismatch is a common problem when applying a speaker verification system on speech data in different languages. This paper proposes a domain transfer network, named EDITnet, to alleviate the language-mismatch problem on speaker embeddings without requiring speaker labels. The network leverages a conditional variational auto-encoder to transfer embeddings from the target domain into the source domain. A self-supervised learning strategy is imposed on the transferred embeddings so as to increase the cosine distance between embeddings from different speakers. In the training process of the EDITnet, the embedding extraction model is fixed without fine-tuning, which renders the training efficient and low-cost. Experiments on Voxceleb and CN-Celeb show that the embeddings transferred by EDITnet outperform the un-transferred ones by around 30% with the ECAPA-TDNN512. Performance improvement can also be achieved with other embedding extraction models, e.g., TDNN, SE-ResNet34.
【4】 The ZevoMOS entry to VoiceMOS Challenge 2022
标题:参加2022年语音MOS挑战赛的ZevoMOS
链接:https://arxiv.org/abs/2206.07448
机构:Communications Department, Technical University of Cluj-Napoca, Romania备注:Accepted at Interspeech 2022 - VoiceMOS Challenge; 5 pages, 2 figures, 2 tables摘要:本文介绍了ZevoMOS进入2022年VoiceMOS挑战赛的主要赛道。ZevoMOS的提交基于预训练自监督学习(SSL)语音模型的两步微调。第一步是对自然语音和合成语音进行分类,而第二步是预测与每个训练样本相关的MOS分数。然后,将微调过程的结果与从自动语音识别模型中提取的置信度得分以及从wav2vec SSL语音模型中获取的训练样本的原始嵌入相结合。VoiceMOS挑战中分配给ZevoMOS系统的团队id为T01。系统级SRCC排名第14位,话语级MSE排名第9位。本文还介绍了对中间结果的附加评估。摘要:This paper introduces the ZevoMOS entry to the main track of the VoiceMOS Challenge 2022. The ZevoMOS submission is based on a two-step finetuning of pretrained self-supervised learning (SSL) speech models. The first step uses a task of classifying natural versus synthetic speech, while the second step's task is to predict the MOS scores associated with each training sample. The results of the finetuning process are then combined with the confidence scores extracted from an automatic speech recognition model, as well as the raw embeddings of the training samples obtained from a wav2vec SSL speech model. The team id assigned to the ZevoMOS system within the VoiceMOS Challenge is T01. The submission was placed on the 14th place with respect to the system-level SRCC, and on the 9th place with respect to the utterance-level MSE. The paper also introduces additional evaluations of the intermediate results.
【5】 Residual Language Model for End-to-end Speech Recognition
标题:端到端语音识别的残差语言模型
链接:https://arxiv.org/abs/2206.07430
作者:Emiru Tsunoo,Yosuke Kashiwagi,Chaitanya Narisetty,Shinji Watanabe机构:Sony Group Corporation, Japan, Carnegie Mellon University, USA备注:Accepted for Interspeech2022摘要:端到端自动语音识别尽管使用了大量成对的音频-文本数据进行训练,但仍无法适应未知的目标域语音。最近的研究估计,该模型的语言偏向为内部语言模型(LM)。为了有效地适应目标域,在推理过程中从后验域中减去内部LM,并与外部目标域LM融合。然而,这种融合使推理复杂化,内部LM的估计可能并不总是准确的。在本文中,我们提出了一种简单的外部LM融合域自适应方法,该方法在训练过程中考虑了内部LM估计。我们直接对外部和内部LMs的残差因子建模,即残差LM。为了稳定地训练残差LM,我们提出对估计的内部LM进行平滑处理,并结合交叉熵和均方误差损失对其进行优化,这考虑了内部LM在目标域数据中的统计行为。我们通过实验证实,在大多数跨域和域内场景中,所提出的残差LM的性能优于内部LM估计。摘要:End-to-end automatic speech recognition suffers from adaptation to unknown target domain speech despite being trained with a large amount of paired audio--text data. Recent studies estimate a linguistic bias of the model as the internal language model (LM). To effectively adapt to the target domain, the internal LM is subtracted from the posterior during inference and fused with an external target-domain LM. However, this fusion complicates the inference and the estimation of the internal LM may not always be accurate. In this paper, we propose a simple external LM fusion method for domain adaptation, which considers the internal LM estimation in its training. We directly model the residual factor of the external and internal LMs, namely the residual LM. To stably train the residual LM, we propose smoothing the estimated internal LM and optimizing it with a combination of cross-entropy and mean-squared-error losses, which consider the statistical behaviors of the internal LM in the target domain data. We experimentally confirmed that the proposed residual LM performs better than the internal LM estimation in most of the cross-domain and intra-domain scenarios.
【6】 Exploiting Cross-domain And Cross-Lingual Ultrasound Tongue Imaging Features For Elderly And Dysarthric Speech Recognition
标题:利用跨域和跨语言的超声舌象特征进行老年人和节律障碍语音识别
链接:https://arxiv.org/abs/2206.07327
作者:Shujie Hu,Xurong Xie,Mengzhe Geng,Mingyu Cui,Jiajun Deng,Tianzi Wang,Xunying Liu,Helen Meng机构:The Chinese University of Hong Kong, Hong Kong SAR, China, Institute of Software, Chinese Academy of Sciences, China备注:arXiv admin note: text overlap with arXiv:2203.10274摘要:发音特征对声音信号失真具有固有的不变性,并已成功地融入为正常语音设计的自动语音识别(ASR)系统中。他们在非典型任务领域的实际应用,如老年人和跨语言的语音障碍,往往受到从目标说话者收集此类专家数据的困难的限制。本文提出了一种跨域、跨语言的A2A反转方法,该方法在A2A模型预训练中利用24小时TaL语料库的并行音频、视频和超声舌成像(UTI)数据,然后跨域、跨语言地适应两种语言的三个数据集:英语DementiaBank Pitt和粤语JCocc MoCA老年语音语料库;和英语TORGO构音障碍语音数据,以生成基于UTI的发音特征。在三项任务上进行的实验表明,通过统计显著的单词错误率或字符错误率降低达2.64%,合并生成的发音特征始终优于使用声学特征构建的基线混合TDNN和基于构象的端到端系统,数据增强和说话人自适应后,绝对值为1.92%,相对值为1.21%(8.17%,7.89%,13.28%)。摘要:Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems designed for normal speech. Their practical application to atypical task domains such as elderly and disordered speech across languages is often limited by the difficulty in collecting such specialist data from target speakers. This paper presents a cross-domain and cross-lingual A2A inversion approach that utilizes the parallel audio, visual and ultrasound tongue imaging (UTI) data of the 24-hour TaL corpus in A2A model pre-training before being cross-domain and cross-lingual adapted to three datasets across two languages: the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech corpora; and the English TORGO dysarthric speech data, to produce UTI based articulatory features. Experiments conducted on three tasks suggested incorporating the generated articulatory features consistently outperformed the baseline hybrid TDNN and Conformer based end-to-end systems constructed using acoustic features only by statistically significant word error rate or character error rate reductions up to 2.64%, 1.92% and 1.21% absolute (8.17%, 7.89% and 13.28% relative) after data augmentation and speaker adaptation were applied.
【7】 Latency Control for Keyword Spotting
标题:关键词识别的时延控制
链接:https://arxiv.org/abs/2206.07261
作者:Christin Jose,Joseph Wang,Grant P. Strimel,Mohammad Omar Khursheed,Yuriy Mishchenko,Brian Kulis机构:Amazon Science, United States备注:Proceedings of INTERSPEECH摘要:会话代理通常利用关键字定位(KWS)来启动与用户的语音交互。出于用户体验和隐私考虑,现有的KWS方法主要关注准确性,这通常会以引入延迟为代价。为了解决这个折衷问题,我们提出了一种控制KWS模型延迟的新方法,并将其推广到任何损失函数,而无需明确了解关键字端点。通过一个可调的超参数,我们的方法可以平衡目标应用程序的检测延迟和准确性。从经验上看,与现有方法相比,我们的方法在延迟约束下提供了优异的性能。也就是说,与最先进的基线状态相比,我们对固定延迟目标做出了25%的相对错误接受改善。我们还表明,当我们的方法与最大池损失结合使用时,与交叉熵损失相比,我们能够在固定延迟下将相对错误接受提高25%。摘要:Conversational agents commonly utilize keyword spotting (KWS) to initiate voice interaction with the user. For user experience and privacy considerations, existing approaches to KWS largely focus on accuracy, which can often come at the expense of introduced latency. To address this tradeoff, we propose a novel approach to control KWS model latency and which generalizes to any loss function without explicit knowledge of the keyword endpoint. Through a single, tunable hyperparameter, our approach enables one to balance detection latency and accuracy for the targeted application. Empirically, we show that our approach gives superior performance under latency constraints when compared to existing methods. Namely, we make a substantial 25\% relative false accepts improvement for a fixed latency target when compared to the baseline state-of-the-art. We also show that when our approach is used in conjunction with a max-pooling loss, we are able to improve relative false accepts by 25 % at a fixed latency when compared to cross entropy loss.
【8】 AVATAR: Unconstrained Audiovisual Speech Recognition
标题:阿凡达:不受限制的视听语音识别
链接:https://arxiv.org/abs/2206.07684
作者:Valentin Gabeur,Paul Hongsuck Seo,Arsha Nagrani,Chen Sun,Karteek Alahari,Cordelia Schmid机构:Inria†, Google Research摘要:视听自动语音识别(AV-ASR)是ASR的一个扩展,它包含视觉线索,通常来自说话者的嘴的运动。与只关注嘴唇运动的作品不同,我们研究了整个视觉框架(视觉动作、物体、背景等)的贡献。这对于无约束视频尤其有用,因为在这些视频中,说话者不一定可见。为了解决这一问题,我们提出了一种新的序列到序列的视听ASR转换器(AVATAR),该转换器通过频谱图和全帧RGB进行端到端的训练。为了防止音频流主导训练,我们提出了不同的单词掩蔽策略,从而鼓励我们的模型关注视频流。我们展示了视觉模式对How2 AV-ASR基准测试的贡献,尤其是在存在模拟噪声的情况下,并表明我们的模型大大优于所有其他先前的工作。最后,我们还为AV-ASR创建了一个名为VisSpeech的新的、真实的测试平台,该平台展示了视觉模态在具有挑战性的音频条件下的贡献。摘要:Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of entire visual frames (visual actions, objects, background etc.). This is particularly useful for unconstrained videos, where the speaker is not necessarily visible. To solve this task, we propose a new sequence-to-sequence AudioVisual ASR TrAnsformeR (AVATAR) which is trained end-to-end from spectrograms and full-frame RGB. To prevent the audio stream from dominating training, we propose different word-masking strategies, thereby encouraging our model to pay attention to the visual stream. We demonstrate the contribution of the visual modality on the How2 AV-ASR benchmark, especially in the presence of simulated noise, and show that our model outperforms all other prior work by a large margin. Finally, we also create a new, real-world test bed for AV-ASR called VisSpeech, which demonstrates the contribution of the visual modality under challenging audio conditions.
【9】 Exploring Capabilities of Monolingual Audio Transformers using Large Datasets in Automatic Speech Recognition of Czech
标题:利用大数据集探索单语音频转换器在捷克语自动语音识别中的能力
链接:https://arxiv.org/abs/2206.07627
作者:Jan Lehečka,Jan Švec,Aleš Pražák,Josef V. Psutka机构:Department of Cybernetics, University of West Bohemia Pilsen, Czech Republic备注:to be published in Proceedings of INTERSPEECH 2022摘要:在本文中,我们介绍了我们在从一个包含8万多小时未标记语音的大型数据集预训练捷克单语音频转换器方面的进展,并随后结合域内数据和近6000小时的域外转录语音对自动语音识别任务的模型进行微调。我们正在展示一系列实验,其中包括在两个公共数据集(CommonVoice和VoxPopuli)和MALACH项目中一个极具挑战性的数据集上评估的各种微调设置。我们的结果表明,单语Wav2Vec 2.0模型是健壮的ASR系统,它可以利用大型标记和未标记数据集,并成功地与最先进的LVCSR系统竞争。此外,当目标ASR任务没有可用的训练数据时,Wav2Vec模型被证明是很好的零炮学习者。摘要:In this paper, we present our progress in pretraining Czech monolingual audio transformers from a large dataset containing more than 80 thousand hours of unlabeled speech, and subsequently fine-tuning the model on automatic speech recognition tasks using a combination of in-domain data and almost 6 thousand hours of out-of-domain transcribed speech. We are presenting a large palette of experiments with various fine-tuning setups evaluated on two public datasets (CommonVoice and VoxPopuli) and one extremely challenging dataset from the MALACH project. Our results show that monolingual Wav2Vec 2.0 models are robust ASR systems, which can take advantage of large labeled and unlabeled datasets and successfully compete with state-of-the-art LVCSR systems. Moreover, Wav2Vec models proved to be good zero-shot learners when no training data are available for the target ASR task.
【10】 Investigating Multi-Feature Selection and Ensembling for Audio Classification
标题:音频分类中的多特征选择与集成方法研究
链接:https://arxiv.org/abs/2206.07511
作者:Muhammad Turab,Teerath Kumar,Malika Bendechache,Takfarinas Saber机构:Mehran University of Engineering and Technology, Jamshoro, Pakistan., ADAPT – Science Foundation Ireland Research Centre, CRT AI, School of Computing, Dublin City University, Dublin, Ireland, Lero – the Irish Software Research Centre摘要:深度学习(DL)算法在不同领域表现出令人印象深刻的性能。其中,由于一些有趣的模式,尤其是在音频数据的分类方面,音频在过去几十年中吸引了许多研究人员。为了获得更好的音频分类性能,特征选择和组合起着关键作用,因为它们有可能影响任何DL模型的性能。要调查此角色,我们对具有各种最先进音频特征(即,Mel频谱图、Mel频率倒谱系数和过零率)的多个尖端DL模型(即卷积神经网络、EfficientNet、MobileNet、超级向量机和多感知器)的性能进行了广泛的评估,这些模型可以是独立的,也可以是组合的(即通过融合)三种不同的数据集(即自由语音数字数据集、音频乌尔都语数字数据集和音频古吉拉特语数字数据集)。总体而言,结果表明特征选择取决于数据集和模型。然而,特征组合应仅限于单独使用时已达到良好性能的特征(即,主要是Mel频谱图、Mel频率倒谱系数)。无论我们选择何种DL模型,这种特性组合/融合都使我们的表现优于之前最先进的结果。摘要:Deep Learning (DL) algorithms have shown impressive performance in diverse domains. Among them, audio has attracted many researchers over the last couple of decades due to some interesting patterns--particularly in classification of audio data. For better performance of audio classification, feature selection and combination play a key role as they have the potential to make or break the performance of any DL model. To investigate this role, we conduct an extensive evaluation of the performance of several cutting-edge DL models (i.e., Convolutional Neural Network, EfficientNet, MobileNet, Supper Vector Machine and Multi-Perceptron) with various state-of-the-art audio features (i.e., Mel Spectrogram, Mel Frequency Cepstral Coefficients, and Zero Crossing Rate) either independently or as a combination (i.e., through ensembling) on three different datasets (i.e., Free Spoken Digits Dataset, Audio Urdu Digits Dataset, and Audio Gujarati Digits Dataset). Overall, results suggest feature selection depends on both the dataset and the model. However, feature combinations should be restricted to the only features that already achieve good performances when used individually (i.e., mostly Mel Spectrogram, Mel Frequency Cepstral Coefficients). Such feature combination/ensembling enabled us to outperform the previous state-of-the-art results irrespective of our choice of DL model.
【11】 VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection
标题:VisageSynTalk:基于语音人脸特征选择的看不见说话人视频到语音合成
链接:https://arxiv.org/abs/2206.07458
作者:Joanna Hong,Minsu Kim,Yong Man Ro机构:Image and Video Systems Lab, School of Electrical Engineering, KAIST, South Korea备注:Submitted to ECCV 2022摘要:这项工作的目标是从无声的人脸视频中重建语音。最近的研究表明,从无声交谈的人脸视频中合成语音的性能令人印象深刻。然而,他们没有明确考虑不同说话人的不同身份特征,这给视频到语音合成带来了挑战,这在看不见的说话人设置中变得更加重要。与之前的方法不同,我们的方法是从给定的无声对话人脸视频中分离语音内容和面部风格。通过引导模型独立地关注这两种表示的建模,我们可以从模型中获得高可懂度的语音,即使给定了一个看不见的对象的输入视频。为此,我们引入了语音视觉选择模块,该模块将语音内容和说话人身份与输入视频的视觉特征分离。将分离的表示合并到一起,通过基于外观样式的合成器合成语音,该合成器通过在保持语音内容的同时涂抹外观样式来生成语音。因此,所提出的框架带来了合成包含正确内容的语音的优势,即使给出了一个看不见的主题的无声对话人脸视频。我们在网格、TCD-TIMIT志愿者和LRW数据集上验证了所提框架的有效性。合成语音可以在补充材料中听到。摘要:The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on varying identity characteristics of different speakers, which place a challenge in the video-to-speech synthesis, and this becomes more critical in unseen-speaker settings. Distinct from the previous methods, our approach is to separate the speech content and the visage-style from a given silent talking face video. By guiding the model to independently focus on modeling the two representations, we can obtain the speech of high intelligibility from the model even when the input video of an unseen subject is given. To this end, we introduce speech-visage selection module that separates the speech content and the speaker identity from the visual features of the input video. The disentangled representations are jointly incorporated to synthesize speech through visage-style based synthesizer which generates speech by coating the visage-styles while maintaining the speech content. Thus, the proposed framework brings the advantage of synthesizing the speech containing the right content even when the silent talking face video of an unseen subject is given. We validate the effectiveness of the proposed framework on the GRID, TCD-TIMIT volunteer, and LRW datasets. The synthesized speech can be heard in supplementary materials.
【12】 NatiQ: An End-to-end Text-to-Speech System for Arabic
标题:NATIQ:一个面向阿拉伯语的端到端文语转换系统
链接:https://arxiv.org/abs/2206.07373
作者:Ahmed Abdelali,Nadir Durrani,Cenk Demiroglu,Fahim Dalvi,Hamdy Mubarak,Kareem Darwish机构:Qatar Computing Research Institute - Hamad Bin Khalifa University, Doha, Qatar, ¨Ozye˘gin University, Istanbul, T¨urkiye摘要:NatiQ是阿拉伯语的端到端文本到语音系统。我们的语音合成器使用编码器-解码器架构。我们使用基于tacotron的模型(tacotron-1和tacotron-2)和快速变换器模型从字符生成mel光谱图。我们将Tacotron1与WaveRNN声码器连接,Tacotron2与WaveGlow声码器连接,ESPnet transformer与并行wavegan声码器连接,以从频谱图合成波形。我们使用了两种声音的内部语音数据:1)中性男性“Hamza”(讲述一般内容和新闻),2)富有表现力的女性“Amina”(讲述儿童故事书)来训练我们的模型。我们的最佳系统对Amina和Hamza的平均意见得分(MOS)分别为4.21和4.40。使用单词和字符错误率(WER和CER)以及实时因素测量的响应时间对系统进行客观评估,有利于端到端架构ESPnet。NatiQ演示可在线访问https://tts.qcri.org摘要:NatiQ is end-to-end text-to-speech system for Arabic. Our speech synthesizer uses an encoder-decoder architecture with attention. We used both tacotron-based models (tacotron-1 and tacotron-2) and the faster transformer model for generating mel-spectrograms from characters. We concatenated Tacotron1 with the WaveRNN vocoder, Tacotron2 with the WaveGlow vocoder and ESPnet transformer with the parallel wavegan vocoder to synthesize waveforms from the spectrograms. We used in-house speech data for two voices: 1) neutral male "Hamza"- narrating general content and news, and 2) expressive female "Amina"- narrating children story books to train our models. Our best systems achieve an average Mean Opinion Score (MOS) of 4.21 and 4.40 for Amina and Hamza respectively. The objective evaluation of the systems using word and character error rate (WER and CER) as well as the response time measured by real-time factor favored the end-to-end architecture ESPnet. NatiQ demo is available on-line at https://tts.qcri.org
【13】 On the Use of Deep Mask Estimation Module for Neural Source Separation Systems
标题:深度掩码估计模块在神经源分离系统中的应用
链接:https://arxiv.org/abs/2206.07347
作者:Kai Li,Xiaolin Hu,Yi Luo机构:†Department of Computer Science and Technology, BNRist, Tsinghua University, China, ‡Tencent AI Lab, Shenzhen, China备注:Accepted by Interspeech 2022摘要:大多数最近的神经源分离系统依赖于基于掩蔽的管道,其中一组乘法掩蔽是从输入混合的信号表示中估计并应用于该信号表示的。在几乎所有的网络体系结构中,这种掩码的估计都是由一个单层和一个可选的非线性激活函数完成的。然而,最近的文献研究了深掩模估计模块的使用,并观察到与浅掩模估计模块相比的性能改进。在本文中,我们通过将深度掩码估计模块与最近提出的无监督源分离方法相连接,分析了这种深度掩码估计模块的作用,并从经验上表明,深度掩码估计模块是所谓的过分离分组范式与传统浅层掩码估计层的有效近似。摘要:Most of the recent neural source separation systems rely on a masking-based pipeline where a set of multiplicative masks are estimated from and applied to a signal representation of the input mixture. The estimation of such masks, in almost all network architectures, is done by a single layer followed by an optional nonlinear activation function. However, recent literatures have investigated the use of a deep mask estimation module and observed performance improvement compared to a shallow mask estimation module. In this paper, we analyze the role of such deeper mask estimation module by connecting it to a recently proposed unsupervised source separation method, and empirically show that the deep mask estimation module is an efficient approximation of the so-called overseparation-grouping paradigm with the conventional shallow mask estimation layers.
【14】 On the Design and Training Strategies for RNN-based Online Neural Speech Separation Systems
标题:基于RNN的在线神经语音分离系统设计与训练策略研究
链接:https://arxiv.org/abs/2206.07340
机构:†Department of Computer Science and Technology, BNRist, Tsinghua University, China, ‡Tencent AI Lab, Shenzhen, China摘要:虽然离线神经语音分离系统的性能因新型神经网络体系结构的发展而得到了极大的提高,但系统与其在线变体之间通常存在不可避免的性能差距。在本文中,我们研究了如何将基于RNN的离线神经语音分离系统转变为在线神经语音分离系统,同时缓解性能下降。我们在双向RNN层中分解或重组前向和后向RNN层,以形成在线路径和离线路径,从而使模型能够使用相同的模型参数集执行在线和离线处理。我们进一步介绍了两种通过预训练离线模型或多任务训练目标改进在线模型的训练策略。实验结果表明,与从头开始训练的在线模型相比,所提出的层分解重组方案和训练策略可以有效地缓解两种基于RNN的离线分离模型及其在线变体之间的性能差距。摘要:While the performance of offline neural speech separation systems has been greatly advanced by the recent development of novel neural network architectures, there is typically an inevitable performance gap between the systems and their online variants. In this paper, we investigate how RNN-based offline neural speech separation systems can be changed into their online counterparts while mitigating the performance degradation. We decompose or reorganize the forward and backward RNN layers in a bidirectional RNN layer to form an online path and an offline path, which enables the model to perform both online and offline processing with a same set of model parameters. We further introduce two training strategies for improving the online model via either a pretrained offline model or a multitask training objective. Experiment results show that compared to the online models that are trained from scratch, the proposed layer decomposition and reorganization schemes and training strategies can effectively mitigate the performance gap between two RNN-based offline separation models and their online variants.
【15】 FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
标题:FRCRN:基于频率递归的改进特征表示用于单声道语音增强
链接:https://arxiv.org/abs/2206.07293
作者:Shengkui Zhao,Bin Ma,Karn N. Watcharasupat,Woon-Seng Gan机构:Alibaba Group, School of Electrical and Electronic Engineering, Nanyang Technological University (NTU), Singapore备注:The paper has been accepted by ICASSP 2022. 5 pages, 2 figures, 5 tables摘要:卷积递归网络(CRN)集成了卷积编解码(CED)结构和递归结构,在单耳语音增强方面取得了良好的性能。然而,由于CED卷积中的感受野有限,跨频率上下文的特征表示受到高度限制。在本文中,我们提出了一种卷积循环编码器-解码器(CRED)结构来增强沿频率轴的特征表示。CRED在每次卷积后沿频率轴对三维卷积特征映射应用频率递归,因此,它能够捕捉长距离的频率相关性并增强语音输入的特征表示。所提出的频率递归是利用前馈顺序存储网络(FSMN)有效实现的。除了CRED之外,我们在编码器和解码器之间插入两个堆叠的FSMN层,以模拟进一步的时间动态。我们将该框架命名为频率递归CRN(FRCRN)。我们设计FRCRN在复值域预测复理想比掩模(cIRM),并利用时频域和时域损耗优化FRCRN。我们提出的方法在宽带基准数据集上取得了最先进的性能,并在ICASSP 2022深度噪声抑制(DNS)挑战赛中,在平均意见得分(MOS)和单词准确性(WAcc)方面取得了实时全频段跟踪的第二名。摘要:Convolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across frequency context is highly constrained due to limited receptive fields in the convolutions of CED. In this paper, we propose a convolutional recurrent encoder-decoder (CRED) structure to boost feature representation along the frequency axis. The CRED applies frequency recurrence on 3D convolutional feature maps along the frequency axis following each convolution, therefore, it is capable of catching long-range frequency correlations and enhancing feature representations of speech inputs. The proposed frequency recurrence is realized efficiently using a feedforward sequential memory network (FSMN). Besides the CRED, we insert two stacked FSMN layers between the encoder and the decoder to model further temporal dynamics. We name the proposed framework as Frequency Recurrent CRN (FRCRN). We design FRCRN to predict complex Ideal Ratio Mask (cIRM) in complex-valued domain and optimize FRCRN using both time-frequency-domain and time-domain losses. Our proposed approach achieved state-of-the-art performance on wideband benchmark datasets and achieved 2nd place for the real-time fullband track in terms of Mean Opinion Score (MOS) and Word Accuracy (WAcc) in the ICASSP 2022 Deep Noise Suppression (DNS) challenge.
【16】 Text-Aware End-to-end Mispronunciation Detection and Diagnosis
标题:文本感知的端到端发音错误检测与诊断
链接:https://arxiv.org/abs/2206.07289
作者:Linkai Peng,Yingming Gao,Binghuai Lin,Dengfeng Ke,Yanlu Xie,Jinsong Zhang机构:School of Information Sciences, Beijing Language and Culture University, Beijing, China, Smart Platform Product Department,Tencent Technology Co., Ltd, Beijing, China备注:Rejected by Interspeech2022摘要:发音错误检测与诊断(MDD)技术是计算机辅助发音训练系统(CAPT)的关键组成部分。在评估受限语音的发音质量方面,给定的抄本可以起到教师的作用。传统的方法已经充分利用了先前的文本来构建模型或提高系统性能,例如强制对齐和扩展识别网络。最近,一些基于端到端的方法试图将先前的文本纳入到模型训练中,并初步证明了其有效性。然而,以往的研究大多考虑将原始注意机制应用于音频表征与文本表征的融合,而没有考虑可能的文本发音失配。在本文中,我们提出了一种门控策略,该策略在抑制无关文本信息的同时,更加重视相关音频特征。此外,在给定转录本的情况下,我们设计了一个额外的对比损失,以缩小音位识别的学习目标与MDD之间的差距。我们使用两个公开的数据集(TIMIT和L2 Arctic)进行了实验,与基线相比,我们的最佳模型将F1得分从57.51美元提高到61.75美元。此外,我们还对门控机制和对比学习对MDD的有效性进行了详细的分析。摘要:Mispronunciation detection and diagnosis (MDD) technology is a key component of computer-assisted pronunciation training system (CAPT). In the field of assessing the pronunciation quality of constrained speech, the given transcriptions can play the role of a teacher. Conventional methods have fully utilized the prior texts for the model construction or improving the system performance, e.g. forced-alignment and extended recognition networks. Recently, some end-to-end based methods attempt to incorporate the prior texts into model training and preliminarily show the effectiveness. However, previous studies mostly consider applying raw attention mechanism to fuse audio representations with text representations, without taking possible text-pronunciation mismatch into account. In this paper, we present a gating strategy that assigns more importance to the relevant audio features while suppressing irrelevant text information. Moreover, given the transcriptions, we design an extra contrastive loss to reduce the gap between the learning objective of phoneme recognition and MDD. We conducted experiments using two publicly available datasets (TIMIT and L2-Arctic) and our best model improved the F1 score from $57.51\%$ to $61.75\%$ compared to the baselines. Besides, we provide a detailed analysis to shed light on the effectiveness of gating mechanism and contrastive learning on MDD.
【17】 Streaming non-autoregressive model for any-to-many voice conversion
标题:用于任意对多语音转换的流非自回归模型
链接:https://arxiv.org/abs/2206.07288
作者:Ziyi Chen,Haoran Miao,Pengyuan Zhang机构:Key Laboratory of Speech Acoustics & Content Understanding, Institute of Acoustics, Chinese, University of Chinese Academy of Sciences, Beijing, China摘要:语音转换模型已经发展了几十年,当前的主流研究集中在非流式语音转换上。然而,流式语音转换比非流式语音转换更适合实际应用场景。在本文中,我们提出了一种基于完全非自回归模型的流式任意多语音转换,包括基于流式变换器的声学模型和流式声码器。基于流Transformer的声学模型由基于流端到端自动语音识别模型的预训练编码器和基于FastSpeech块的解码器组成。流式声码器采用伪正交镜像滤波器组和因果卷积设计,用于流式任务。实验结果表明,该方法在延迟和转换质量方面都取得了显著的性能,并且可以在CPU和GPU上实现实时性。摘要:Voice conversion models have developed for decades, and current mainstream research focuses on non-streaming voice conversion. However, streaming voice conversion is more suitable for practical application scenarios than non-streaming voice conversion. In this paper, we propose a streaming any-to-many voice conversion based on fully non-autoregressive model, which includes a streaming transformer based acoustic model and a streaming vocoder. Streaming transformer based acoustic model is composed of a pre-trained encoder from streaming end-to-end based automatic speech recognition model and a decoder modified on FastSpeech blocks. Streaming vocoder is designed for streaming task with pseudo quadrature mirror filter bank and causal convolution. Experimental results show that the proposed method achieves significant performance both in latency and conversion quality and can be real-time on CPU and GPU.
【18】 Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning
标题:基于数据驱动深度学习的可见语音和未见语音情感强度精确评估
链接:https://arxiv.org/abs/2206.07229
作者:Rui Liu,Berrak Sisman,Björn Schuller,Guanglai Gao,Haizhou Li机构:Inner Mongolia University, China , The Chinese University of Hong Kong, Shenzhen, China, National University of Singapore , Singapore University of Technology and Design, Imperial College London, United Kingdom备注:To appear in INTERSPEECH 2022. 5 pages, 4 figures. Substantial text overlap with arXiv:2110.03156摘要:在情感文本到语音和语音转换等应用中,需要对语音进行情感分类并评估情感强度。提出了基于支持向量机(SVM)的情感属性排序函数来预测情感语音语料库的情感强度。然而,经过训练的排序函数并没有推广到新的领域,这限制了它的应用范围,尤其是对于域外或看不见的语音。在本文中,我们提出了一种数据驱动的深度学习模型,即StrengthNet,以改进对可见和不可见语音的情感强度评估的泛化。这是通过融合来自不同领域的情感数据来实现的。我们遵循一个多任务学习网络架构,该架构包括一个声学编码器、一个强度预测器和一个辅助情绪预测器。实验表明,所提出的StrengthNet的预测情感强度与可见和不可见语音的地面真实度得分高度相关。我们在以下位置发布源代码:https://github.com/ttslr/StrengthNet.摘要:Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet.
【19】 Frequency-centroid features for word recognition of non-native English speakers
标题:非英语母语者单词识别的频率中心特征
链接:https://arxiv.org/abs/2206.07176
作者:Pierre Berjon,Rajib Sharma,Avishek Nag,Soumyabrata Dev机构:Department de Sciences du Numérique, INP-ENSEEIHT, Toulouse, France, Department of DSIS, Indian Institute of Information Technology Dharwad, India, School of Electrical and Electronic Engineering, University College Dublin, Ireland备注:Published in IEEE Irish Signals & Systems Conference (ISSC), 2022摘要:这项工作的目的是研究互补特征,以帮助不同母语的非英语母语者在封闭、有限集合词识别任务中使用典型的Mel频率倒谱系数(MFCC)。与MFCC不同的是,MFCC源自语音信号的频谱能量,建议的频率质心(FCs)封装了语音频谱不同频带的频谱中心,频带由Mel滤波器组定义。这些特征与MFCC相结合,可以提高英语单词识别的相对性能,尤其是在各种噪声条件下。采用两级卷积神经网络(CNN)对英语中带有阿拉伯语、法语和西班牙语口音的单词的特征进行建模。摘要:The objective of this work is to investigate complementary features which can aid the quintessential Mel frequency cepstral coefficients (MFCCs) in the task of closed, limited set word recognition for non-native English speakers of different mother-tongues. Unlike the MFCCs, which are derived from the spectral energy of the speech signal, the proposed frequency-centroids (FCs) encapsulate the spectral centres of the different bands of the speech spectrum, with the bands defined by the Mel filterbank. These features, in combination with the MFCCs, are observed to provide relative performance improvement in English word recognition, particularly under varied noisy conditions. A two-stage Convolution Neural Network (CNN) is used to model the features of the English words uttered with Arabic, French and Spanish accents.
机器翻译,仅供参考