【1】 Masked Autoencoders that Listen
标题:可监听的蒙面自动编码器
链接:https://arxiv.org/abs/2207.06405
作者:Po-Yao,Huang,Hu Xu,Juncheng Li,Alexei Baevski,Michael Auli,Wojciech Galuba,Florian Metze,Christoph Feichtenhofer机构:FAIR, Meta AI, Carnegie Mellon University摘要:本文研究了基于图像的掩码自动编码器(MAE)的简单扩展,以从音频频谱图中进行自监督表示学习。根据MAE中的Transformer编码器-解码器设计,我们的音频MAE首先以高掩蔽率对音频频谱图块进行编码,通过编码器层仅向非掩蔽令牌馈电。然后,解码器对填充有掩码令牌的编码上下文进行重新排序和解码,以重建输入频谱图。我们发现,在解码器中加入局部窗口注意是有益的,因为音频频谱在局部时间和频带上高度相关。然后,我们在目标数据集上以较低的掩蔽率微调编码器。根据经验,音频MAE在六个音频和语音分类任务上设定了最新的性能,优于其他使用外部监督预训练的最新模型。代码和模型将在https://github.com/facebookresearch/AudioMAE.摘要:This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only the non-masked tokens through encoder layers. The decoder then re-orders and decodes the encoded context padded with mask tokens, in order to reconstruct the input spectrogram. We find it beneficial to incorporate local window attention in the decoder, as audio spectrograms are highly correlated in local time and frequency bands. We then fine-tune the encoder with a lower masking ratio on target datasets. Empirically, Audio-MAE sets new state-of-the-art performance on six audio and speech classification tasks, outperforming other recent models that use external supervised pre-training. The code and models will be at https://github.com/facebookresearch/AudioMAE.
【2】 Polyphonic sound event detection for highly dense birdsong scenes
标题:用于高密度鸟鸣场景的复音事件检测
链接:https://arxiv.org/abs/2207.06349
作者:Alberto García Arroba Parrilla,Dan Stowell机构:. Tilburg University, the Netherlands, . Naturalis, Biodiversity, Center, Leiden摘要:日出前一个小时,你可以体验黎明合唱,来自不同物种的鸟类一起唱歌。在这种情况下,与重叠声源的数量一样,很容易出现高水平的复调,从而导致复杂的声学结果。声音事件检测任务分析声音场景,以识别发生的事件及其各自的时间信息。然而,高密度场景可能难以处理,并且尚未进行深入研究。在这里,我们使用卷积递归神经网络(CRNN)展示了在处理更高的复调时如何检测鸟鸣复调场景,以及这种模型如何有效地面对多达10只重叠鸟类的非常密集的场景。我们发现,使用更密集的示例(即更高的复调)训练的模型与在训练集中使用更简单样本的模型学习速度相似。此外,使用密度最大的样本训练的模型对所有复调都保持一致的分数,而使用密度最小的样本训练的模型随着复调的增加而退化。我们的结果表明,使用CRNN可以处理高密度声学场景。我们希望这项研究可以作为研究高密度鸟类场景的起点,例如黎明合唱或其他密集声学问题。摘要:One hour before sunrise, one can experience the dawn chorus where birds from different species sing together. In this scenario, high levels of polyphony, as in the number of overlapping sound sources, are prone to happen resulting in a complex acoustic outcome. Sound Event Detection (SED) tasks analyze acoustic scenarios in order to identify the occurring events and their respective temporal information. However, highly dense scenarios can be hard to process and have not been studied in depth. Here we show, using a Convolutional Recurrent Neural Network (CRNN), how birdsong polyphonic scenarios can be detected when dealing with higher polyphony and how effectively this type of model can face a very dense scene with up to 10 overlapping birds. We found that models trained with denser examples (i.e., higher polyphony) learn at a similar rate as models that used simpler samples in their training set. Additionally, the model trained with the densest samples maintained a consistent score for all polyphonies, while the model trained with the least dense samples degraded as the polyphony increased. Our results demonstrate that highly dense acoustic scenarios can be dealt with using CRNNs. We expect that this study serves as a starting point for working on highly populated bird scenarios such as dawn chorus or other dense acoustic problems.
【3】 Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech
标题:可控无损非自回归端到端文语转换
链接:https://arxiv.org/abs/2207.06088
作者:Zhengxi Liu,Qiao Tian,Chenxu Hu,Xudong Liu,Menglin Wu,Yuping Wang,Hang Zhao,Yuxuan Wang机构:Speech, Audio and Music Intelligence (SAMI), ByteDance, IIIS, Tsinghua University摘要:最近的一些研究已经证明了单阶段神经文本到语音的可行性,它不需要生成mel频谱,而是直接从文本生成原始波形。单阶段文本到语音通常面临两个问题:a)由于多个语音变体而导致的一对多映射问题和b)由于在训练期间缺乏对地面真实声学特征的监督而导致高频重建不足。为了解决a)问题并生成更具表现力的语音,我们提出了一种新的音素级韵律建模方法,该方法基于具有归一化流的变分自动编码器来建模语音中的潜在韵律信息。我们还使用韵律预测器来支持端到端的表达性语音合成。此外,我们提出了双并行自动编码器,在训练期间引入对地面真实声学特征的监督,以解决b)问题,使我们的模型能够生成高质量的语音。我们在内部表达性英语数据集上比较了合成质量与最先进的文本语音系统。定性和定量评估都证明了我们的方法在无损语音生成方面的优越性和鲁棒性,同时也显示出强大的韵律建模能力。摘要:Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces two problems: a) the one-to-many mapping problem due to multiple speech variations and b) insufficiency of high frequency reconstruction due to the lack of supervision of ground-truth acoustic features during training. To solve the a) problem and generate more expressive speech, we propose a novel phoneme-level prosody modeling method based on a variational autoencoder with normalizing flows to model underlying prosodic information in speech. We also use the prosody predictor to support end-to-end expressive speech synthesis. Furthermore, we propose the dual parallel autoencoder to introduce supervision of the ground-truth acoustic features during training to solve the b) problem enabling our model to generate high-quality speech. We compare the synthesis quality with state-of-the-art text-to-speech systems on an internal expressive English dataset. Both qualitative and quantitative evaluations demonstrate the superiority and robustness of our method for lossless speech generation while also showing a strong capability in prosody modeling.
【4】 Subband-based Generative Adversarial Network for Non-parallel Many-to-many Voice Conversion
标题:基于子带的非并行多对多语音转换生成对抗性网络
链接:https://arxiv.org/abs/2207.06057
作者:Jian Ma,Zhedong Zheng,Hao Fei,Feng Zheng,Tat-seng Chua,Yi Yang机构:Southern University of Science and Technology摘要:语音转换是生成具有源内容和目标语音风格的新语音。在本文中,我们重点关注一种通用设置,即非并行多对多语音转换,这与真实场景非常接近。顾名思义,非并行多对多语音转换不需要成对的源语音和参考语音,可以应用于任意语音传输。近年来,生成对抗网络(GAN)和条件变分自动编码器(CVAE)等技术在这一领域取得了长足的进展。然而,由于语音转换的复杂性,转换语音的风格相似性仍然不能令人满意。受mel谱图固有结构的启发,我们提出了一种新的语音转换框架,即基于子带的生成式语音转换对抗网络(SGAN-VC)。SGAN-VC通过明确利用不同子带之间的空间特征,分别转换源语音的每个子带内容。SGAN-VC包含一个样式编码器、一个内容编码器和一个解码器。特别是,风格编码器网络设计用于学习目标说话人不同子带的风格代码。内容编码器网络可以捕获源语音上的内容信息。最后,解码器生成特定子带内容。此外,我们提出了一个基音偏移模块来微调源说话人的基音,使转换后的音调更准确和易于解释。大量实验表明,无论是在可见数据还是不可见数据上,该方法在VCTK语料库和Aisell3数据集上都取得了最先进的性能。此外,SGAN-VC在不可见数据上的内容可懂度甚至超过了有ASR网络辅助的StarGANv2-VC。摘要:Voice conversion is to generate a new speech with the source content and a target voice style. In this paper, we focus on one general setting, i.e., non-parallel many-to-many voice conversion, which is close to the real-world scenario. As the name implies, non-parallel many-to-many voice conversion does not require the paired source and reference speeches and can be applied to arbitrary voice transfer. In recent years, Generative Adversarial Networks (GANs) and other techniques such as Conditional Variational Autoencoders (CVAEs) have made considerable progress in this field. However, due to the sophistication of voice conversion, the style similarity of the converted speech is still unsatisfactory. Inspired by the inherent structure of mel-spectrogram, we propose a new voice conversion framework, i.e., Subband-based Generative Adversarial Network for Voice Conversion (SGAN-VC). SGAN-VC converts each subband content of the source speech separately by explicitly utilizing the spatial characteristics between different subbands. SGAN-VC contains one style encoder, one content encoder, and one decoder. In particular, the style encoder network is designed to learn style codes for different subbands of the target speaker. The content encoder network can capture the content information on the source speech. Finally, the decoder generates particular subband content. In addition, we propose a pitch-shift module to fine-tune the pitch of the source speaker, making the converted tone more accurate and explainable. Extensive experiments demonstrate that the proposed approach achieves state-of-the-art performance on VCTK Corpus and AISHELL3 datasets both qualitatively and quantitatively, whether on seen or unseen data. Furthermore, the content intelligibility of SGAN-VC on unseen data even exceeds that of StarGANv2-VC with ASR network assistance.
【5】 Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech Recognition
标题:面向端到端语音识别的视觉上下文驱动的音频特征增强
链接:https://arxiv.org/abs/2207.06020
作者:Joanna Hong,Minsu Kim,Daehun Yoo,Yong Man Ro机构:KAIST, Daejeon, South Korea, Genesis Lab Inc., Seoul, South Korea备注:Accepted at Interspeech 2022摘要:本文主要研究设计一种抗噪声的端到端视听语音识别系统。为此,我们提出了视觉上下文驱动的音频特征增强模块(V-CAFE),以借助视听对应来增强输入的噪声音频语音。提出的V-CAFE旨在捕捉嘴唇运动的过渡,即视觉语境,并通过考虑获得的视觉语境生成降噪掩码。通过上下文相关建模,可以细化视位到音素映射中的歧义,以生成掩码。噪声表示用降噪掩码掩盖,从而增强音频特征。将增强的音频特征与视觉特征融合,并将其用于语音识别的编码器-解码器模型,该模型由构象器和变换器组成。我们表明,使用V-CAFE的端到端AVSR可以进一步提高AVSR的噪声鲁棒性。使用两个最大的视听数据集LRS2和LRS3,在噪声语音识别和重叠语音识别实验中评估了该方法的有效性。摘要:This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a help of audio-visual correspondence. The proposed V-CAFE is designed to capture the transition of lip movements, namely visual context and to generate a noise reduction mask by considering the obtained visual context. Through context-dependent modeling, the ambiguity in viseme-to-phoneme mapping can be refined for mask generation. The noisy representations are masked out with the noise reduction mask resulting in enhanced audio features. The enhanced audio features are fused with the visual features and taken to an encoder-decoder model composed of Conformer and Transformer for speech recognition. We show the proposed end-to-end AVSR with the V-CAFE can further improve the noise-robustness of AVSR. The effectiveness of the proposed method is evaluated in noisy speech recognition and overlapped speech recognition experiments using the two largest audio-visual datasets, LRS2 and LRS3.
【6】 Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
标题:神经文语转换中文本驱动的情绪风格控制与交叉语者风格迁移
链接:https://arxiv.org/abs/2207.06000
作者:Yookyung Shin,Younggun Lee,Suhee Jo,Yeongtae Hwang,Taesu Kim备注:Accepted to Interspeech 2022摘要:近年来,具有表现力的文本到语音的性能有所提高。然而,合成语音的风格控制通常局限于离散的情感类别,需要目标说话人以目标风格记录训练数据。在许多实际情况下,用户可能没有在目标情感中记录参考语音,但仍然有兴趣通过键入所需情感风格的文本描述来控制语音风格。在本文中,我们提出了一种基于文本的界面,用于多说话人TTS中的情感风格控制和跨说话人风格转换。我们提出了双模风格编码器,该编码器使用预训练的语言模型来模拟文本描述嵌入和语音风格嵌入之间的语义关系。为了进一步改善不相交、多风格数据集上的跨说话人风格转换,我们提出了新的风格损失。实验结果表明,我们的模型可以生成高质量的表达性语音,即使是在看不见的风格。摘要:Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the target style. In many practical situations, users may not have reference speech recorded in target emotion but still be interested in controlling speech style just by typing text description of desired emotional style. In this paper, we propose a text-based interface for emotional style control and cross-speaker style transfer in multi-speaker TTS. We propose the bi-modal style encoder which models the semantic relationship between text description embedding and speech style embedding with a pretrained language model. To further improve cross-speaker style transfer on disjoint, multi-style datasets, we propose the novel style loss. The experimental results show that our model can generate high-quality expressive speech even in unseen style.
【7】 NEC: Speaker Selective Cancellation via Neural Enhanced Ultrasound Shadowing
标题:NEC:基于神经增强超声跟踪的说话人选择性对消
链接:https://arxiv.org/abs/2207.05848
作者:Hanqing Guo,Chenning Li,Lingkun Li,Zhichao Cao,Qiben Yan,Li Xiao机构:Department of Computer Science and Engineering, Michigan State University, ∗These authors contributed equally摘要:在本文中,我们提出了NEC(神经增强消除),这是一种防御机制,可防止未经授权的麦克风捕捉目标说话人的声音。与现有的基于置乱的音频抵消方法相比,NEC可以有选择地从混合语音中删除目标说话人的语音,而不会对其他语音造成干扰。具体来说,对于目标说话人,我们设计了一个深度神经网络(DNN)模型,从其参考音频中提取特定于说话人但与话语无关的高级语音特征。当麦克风录制时,DNN生成阴影声音以实时取消目标声音。此外,我们将可听的阴影声音调制成超声频率,使人类听不到。通过利用麦克风电路的非线性,麦克风可以准确解码阴影声音以消除目标语音。我们在不同设置下使用8个智能手机麦克风全面实施和评估NEC。结果表明,NEC在不干扰其他用户正常对话的情况下,有效地使目标扬声器在麦克风前静音。摘要:In this paper, we propose NEC (Neural Enhanced Cancellation), a defense mechanism, which prevents unauthorized microphones from capturing a target speaker's voice. Compared with the existing scrambling-based audio cancellation approaches, NEC can selectively remove a target speaker's voice from a mixed speech without causing interference to others. Specifically, for a target speaker, we design a Deep Neural Network (DNN) model to extract high-level speaker-specific but utterance-independent vocal features from his/her reference audios. When the microphone is recording, the DNN generates a shadow sound to cancel the target voice in real-time. Moreover, we modulate the audible shadow sound onto an ultrasound frequency, making it inaudible for humans. By leveraging the non-linearity of the microphone circuit, the microphone can accurately decode the shadow sound for target voice cancellation. We implement and evaluate NEC comprehensively with 8 smartphone microphones in different settings. The results show that NEC effectively mutes the target speaker at a microphone without interfering with other users' normal conversations.
【8】 Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices
标题:适用于低资源设备的二进制神经网络提取非语义语音嵌入
链接:https://arxiv.org/abs/2207.05784
作者:Harlin Lee,Aaqib Saeed机构:edu) is with the Department of Mathematics, University of California摘要:本文介绍了BRILLsson,一种新的基于二进制神经网络的表示学习模型,用于广泛的非语义语音任务。我们通过从一个大型实值TRILLsson模型中提取知识来训练该模型,其中只有一小部分数据集用于训练TRILLsson。由此产生的BRILLsson模型大小仅为2MB,延迟小于8ms,适合部署在可穿戴设备等低资源设备中。我们在八个基准任务(包括但不限于口语识别、情绪识别、健康状况诊断和关键词识别)上评估了BRILLsson,并证明了我们提出的超轻和低延迟模型与大规模模型的性能相同。摘要:This work introduces BRILLsson, a novel binary neural network-based representation learning model for a broad range of non-semantic speech tasks. We train the model with knowledge distillation from a large and real-valued TRILLsson model with only a fraction of the dataset used to train TRILLsson. The resulting BRILLsson models are only 2MB in size with a latency less than 8ms, making them suitable for deployment in low-resource devices such as wearables. We evaluate BRILLsson on eight benchmark tasks (including but not limited to spoken language identification, emotion recognition, heath condition diagnosis, and keyword spotting), and demonstrate that our proposed ultra-light and low-latency models perform as well as large-scale models.
【9】 ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
标题:ProDiff:高质量文语转换的渐进式快速扩散模型
链接:https://arxiv.org/abs/2207.06389
作者:Rongjie Huang,Zhou Zhao,Huadai Liu,Jinglin Liu,Chenye Cui,Yi Ren备注:Accepted by ACM Multimedia 2022摘要:去噪扩散概率模型(DDPM)最近在许多生成任务中取得了领先的性能。然而,继承的迭代采样过程成本阻碍了其在文本到语音部署中的应用。通过对扩散模型参数化的初步研究,我们发现以前基于梯度的TTS模型需要数百或数千次迭代才能保证高样本质量,这对加速采样提出了挑战。在这项工作中,我们提出了ProDiff,关于高质量文本到语音的渐进快速扩散模型。与以前估计数据密度梯度的工作不同,ProDiff通过直接预测干净数据来参数化去噪模型,以避免加速采样时出现明显的质量下降。为了通过减少扩散迭代来应对模型收敛挑战,ProDiff通过知识提取来减少目标站点中的数据方差。具体来说,去噪模型使用从N步DDIM教师生成的mel谱图作为训练目标,并将行为提取到具有N/2步的新模型中。因此,它允许TTS模型进行精确预测,并进一步将采样时间减少了几个数量级。我们的评估表明,ProDiff只需2次迭代即可合成高保真mel谱图,同时它使用数百个步骤保持了与最先进模型相比的样本质量和多样性。ProDiff使采样速度比单个NVIDIA 2080Ti GPU上的实时速度快24倍,使扩散模型首次实际适用于文本语音合成部署。我们广泛的烧蚀研究表明,ProDiff中的每种设计都是有效的,并且我们进一步表明,ProDiff可以很容易地扩展到多扬声器设置。音频样本位于{https://ProDiff.github.io/.}摘要:Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinder their applications to text-to-speech deployment. Through the preliminary study on diffusion model parameterization, we find that previous gradient-based TTS models require hundreds or thousands of iterations to guarantee high sample quality, which poses a challenge for accelerating sampling. In this work, we propose ProDiff, on progressive fast diffusion model for high-quality text-to-speech. Unlike previous work estimating the gradient for data density, ProDiff parameterizes the denoising model by directly predicting clean data to avoid distinct quality degradation in accelerating sampling. To tackle the model convergence challenge with decreased diffusion iterations, ProDiff reduces the data variance in the target site via knowledge distillation. Specifically, the denoising model uses the generated mel-spectrogram from an N-step DDIM teacher as the training target and distills the behavior into a new model with N/2 steps. As such, it allows the TTS model to make sharp predictions and further reduces the sampling time by orders of magnitude. Our evaluation demonstrates that ProDiff needs only 2 iterations to synthesize high-fidelity mel-spectrograms, while it maintains sample quality and diversity competitive with state-of-the-art models using hundreds of steps. ProDiff enables a sampling speed of 24x faster than real-time on a single NVIDIA 2080Ti GPU, making diffusion models practically applicable to text-to-speech synthesis deployment for the first time. Our extensive ablation studies demonstrate that each design in ProDiff is effective, and we further show that ProDiff can be easily extended to the multi-speaker setting. Audio samples are available at \url{https://ProDiff.github.io/.}
【10】 MM-ALT: A Multimodal Automatic Lyric Transcription System
标题:MM-ALT:一个多通道的歌词自动转录系统
链接:https://arxiv.org/abs/2207.06127
作者:Xiangming Gu,Longshen Ou,Danielle Ong,Ye Wang机构:Integrative Sciences and Engineering Programme, NUS, Graduate School, National University of Singapore, School of Computing, National University of Singapore备注:Camera ready version. Accepted by ACM Multimedia 2022摘要:自动歌词转录是一个新兴的研究领域,由于其巨大的应用潜力,吸引了语音和音乐信息检索界越来越多的兴趣。然而,由于乐器伴奏和音乐限制,仅使用音频数据的ALT是一项众所周知的困难任务,这会导致语音提示和歌词清晰度的下降。为了应对这一挑战,我们提出了多模式自动歌词转录系统(MM-ALT),以及一个新的数据集N20EM,该数据集包括录音、嘴唇运动视频和表演歌手佩戴的耳塞的惯性测量单元(IMU)数据。我们首先将wav2vec 2.0框架从自动语音识别(ASR)调整到ALT任务。然后,我们提出了一种基于视频的ALT方法和一种基于IMU的语音活动检测(VAD)方法。此外,我们提出了剩余交叉注意力(RCA)机制来融合来自三种模式(即音频、视频和IMU)的数据。实验证明了我们提出的MM-ALT系统的有效性,尤其是在噪声鲁棒性方面。摘要:Automatic lyric transcription (ALT) is a nascent field of study attracting increasing interest from both the speech and music information retrieval communities, given its significant application potential. However, ALT with audio data alone is a notoriously difficult task due to instrumental accompaniment and musical constraints resulting in degradation of both the phonetic cues and the intelligibility of sung lyrics. To tackle this challenge, we propose the MultiModal Automatic Lyric Transcription system (MM-ALT), together with a new dataset, N20EM, which consists of audio recordings, videos of lip movements, and inertial measurement unit (IMU) data of an earbud worn by the performing singer. We first adapt the wav2vec 2.0 framework from automatic speech recognition (ASR) to the ALT task. We then propose a video-based ALT method and an IMU-based voice activity detection (VAD) method. In addition, we put forward the Residual Cross Attention (RCA) mechanism to fuse data from the three modalities (i.e., audio, video, and IMU). Experiments show the effectiveness of our proposed MM-ALT system, especially in terms of noise robustness.
【11】 SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate
标题:萨茨:演讲者文本到语音的吸引人,通过学会分离来学习说话
链接:https://arxiv.org/abs/2207.06011
作者:Nabarun Goswami,Tatsuya Harada机构:The University of Tokyo, Japan, RIKEN, Japan备注:Accepted to Interspeech 2022. Visit this https URL for a demo摘要:文本到语音(TTS)的映射是不确定的,字母的发音可能会根据语境而不同,或者音素可能会因性别、年龄、口音、情绪等各种生理和文体因素而不同。神经说话人嵌入,经过训练以识别或验证说话人通常用于表示这些特征,并将其从参考语音转换为合成语音。另一方面,语音分离是一项具有挑战性的任务,将单个说话人从不同说话人的重叠混合信号中分离出来。说话人吸引子是高维嵌入向量,将每个说话人语音的时频箱拉向自己,同时排斥其他说话人的时频箱。在这项工作中,我们探索了在多说话人TTS合成中使用这些强大的说话人吸引子进行零拍说话人自适应的可能性,并提出了说话人吸引子文本到语音(SATT)。通过各种实验,我们表明,SATT可以从未知目标说话人的参考信号中从文本合成自然语音,该参考信号可能具有低于理想的记录条件,即混响或与其他说话人混合。摘要:The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, emotions, etc. Neural speaker embeddings, trained to identify or verify speakers are typically used to represent and transfer such characteristics from reference speech to synthesized speech. Speech separation on the other hand is the challenging task of separating individual speakers from an overlapping mixed signal of various speakers. Speaker attractors are high-dimensional embedding vectors that pull the time-frequency bins of each speaker's speech towards themselves while repelling those belonging to other speakers. In this work, we explore the possibility of using these powerful speaker attractors for zero-shot speaker adaptation in multi-speaker TTS synthesis and propose speaker attractor text to speech (SATTS). Through various experiments, we show that SATTS can synthesize natural speech from text from an unseen target speaker's reference signal which might have less than ideal recording conditions, i.e. reverberations or mixed with other speakers.
【12】 Cross-Age Speaker Verification: Learning Age-Invariant Speaker Embeddings
标题:跨年龄说话人验证:学习年龄不变的说话人嵌入
链接:https://arxiv.org/abs/2207.05929
作者:Xiaoyi Qin,Na Li,Chao Weng,Dan Su,Ming Li机构:School of Computer Science, Wuhan University, Wuhan, China, Data Science Research Center, Duke Kunshan University, Kunshan, China, Tencent AI Lab, Shenzhen, China备注:Accepted by Interspeech2022摘要:近年来,说话人自动识别技术取得了显著的进展。然而,由于相关数据不足,跨年龄说话人验证的研究很少。在本文中,我们基于VoxCeleb数据集挖掘跨年龄测试集,并提出了年龄不变说话人表示(AISR)学习方法。由于VoxCeleb是从YouTube平台收集的,因此数据集本质上由跨年龄数据组成。然而,元数据不包含说话人年龄标签。因此,我们采用人脸年龄估计方法从相关的视觉数据中预测说话人的年龄值,然后用估计的年龄标记音频记录。我们在VoxCeleb(Vox-CA)上构建了多个交叉年龄测试集,故意选择年龄差距较大的阳性试验。此外,在选择与Vox-H病例对应的阴性配对时,还考虑了国籍和性别的影响。基线系统性能从Vox-H测试集的1.939\%EER下降到Vox-CA20测试集的10.419\%,这表明跨年龄场景有多困难。因此,我们提出了一种年龄解耦对抗学习(ADAL)方法,以缓解年龄差距的负面影响,减少类内方差。我们的方法优于基线系统,在Vox-CA20测试集上相关EER减少了10%以上。源代码和试用资源可在https://github.com/qinxiaoyi/Cross-Age_Speaker_Verification摘要:Automatic speaker verification has achieved remarkable progress in recent years. However, there is little research on cross-age speaker verification (CASV) due to insufficient relevant data. In this paper, we mine cross-age test sets based on the VoxCeleb dataset and propose our age-invariant speaker representation(AISR) learning method. Since the VoxCeleb is collected from the YouTube platform, the dataset consists of cross-age data inherently. However, the meta-data does not contain the speaker age label. Therefore, we adopt the face age estimation method to predict the speaker age value from the associated visual data, then label the audio recording with the estimated age. We construct multiple Cross-Age test sets on VoxCeleb (Vox-CA), which deliberately select the positive trials with large age-gap. Also, the effect of nationality and gender is considered in selecting negative pairs to align with Vox-H cases. The baseline system performance drops from 1.939\% EER on the Vox-H test set to 10.419\% on the Vox-CA20 test set, which indicates how difficult the cross-age scenario is. Consequently, we propose an age-decoupling adversarial learning (ADAL) method to alleviate the negative effect of the age gap and reduce intra-class variance. Our method outperforms the baseline system by over 10\% related EER reduction on the Vox-CA20 test set. The source code and trial resources are available on https://github.com/qinxiaoyi/Cross-Age_Speaker_Verification
【13】 Online Target Speaker Voice Activity Detection for Speaker Diarization
标题:用于说话人二值化的在线目标说话人语音活动检测
链接:https://arxiv.org/abs/2207.05920
作者:Weiqing Wang,Qingjian Lin,Ming Li机构:Department of Electrical & Computer Engineering, Duke University, Durham, NC , USA, Data Science Research Center, Duke Kunshan University, Kunshan , PR China, AI Lab, Lenovo Research, Beijing , PR China备注:Accepted by Interspeech 2022摘要:本文提出了一种用于说话人日记任务的在线目标说话人语音活动检测系统,该系统不需要来自基于聚类的日记系统的先验知识来获得目标说话人嵌入。首先,我们采用基于ResNet的前端模型来提取每个信号块的帧级扬声器嵌入。接下来,我们基于这些帧级说话人嵌入和先前估计的目标说话人嵌入来预测每个说话人的检测状态。然后,通过根据当前块中的预测聚合这些帧级说话人嵌入来更新目标说话人嵌入。我们迭代地提取每个块的结果,并更新目标说话人嵌入,直到到达信号的末尾。实验结果表明,该方法优于AliMeeting数据集上基于离线聚类的日记系统。摘要:This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings. First, we employ a ResNet-based front-end model to extract the frame-level speaker embeddings for each coming block of a signal. Next, we predict the detection state of each speaker based on these frame-level speaker embeddings and the previously estimated target speaker embedding. Then, the target speaker embeddings are updated by aggregating these frame-level speaker embeddings according to the predictions in the current block. We iteratively extract the results for each block and update the target speaker embedding until reaching the end of the signal. Experimental results show that the proposed method is better than the offline clustering-based diarization system on the AliMeeting dataset.
【14】 A Cyclical Approach to Synthetic and Natural Speech Mismatch Refinement of Neural Post-filter for Low-cost Text-to-speech System
标题:低成本文语转换系统中人工语音与自然语音失配的循环优化神经后置滤波器
链接:https://arxiv.org/abs/2207.05913
作者:Yi-Chiao Wu,Patrick Lumban Tobing,Kazuki Yasuhara,Noriyuki Matsunaga,Yamato Ohtani,Tomoki Toda机构:The written permission of Cambridge University Press must be obtained for commercial re-use, Nagoya University备注:15 pages, 7 figures, 10 tables摘要:由于神经网络的快速发展,基于神经的文本到语音(TTS)系统实现了非常高保真的语音生成。然而,巨大的标记语料库和高计算成本要求限制了小公司或个人开发高保真TTS系统的可能性。另一方面,在基于神经的TTS系统中广泛用于语音生成的神经声码器可以用相对较小的未标记语料库进行训练。因此,在本文中,我们探索了一个通用框架,用于开发使用神经声码器的低成本TTS系统的神经后滤波器(NPF)。提出了一种循环方法来解决开发NPF时的声学和时间不匹配(AM和TM)。进行了客观和主观评估,以证明AM和TM问题以及拟议框架的有效性。摘要:Neural-based text-to-speech (TTS) systems achieve very high-fidelity speech generation because of the rapid neural network developments. However, the huge labeled corpus and high computation cost requirements limit the possibility of developing a high-fidelity TTS system by small companies or individuals. On the other hand, a neural vocoder, which has been widely adopted for the speech generation in neural-based TTS systems, can be trained with a relatively small unlabeled corpus. Therefore, in this paper, we explore a general framework to develop a neural post-filter (NPF) for low-cost TTS systems using neural vocoders. A cyclical approach is proposed to tackle the acoustic and temporal mismatches (AM and TM) of developing an NPF. Both objective and subjective evaluations have been conducted to demonstrate the AM and TM problems and the effectiveness of the proposed framework.【1】 ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
标题:ProDiff:高质量文语转换的渐进式快速扩散模型
链接:https://arxiv.org/abs/2207.06389
作者:Rongjie Huang,Zhou Zhao,Huadai Liu,Jinglin Liu,Chenye Cui,Yi Ren备注:Accepted by ACM Multimedia 2022摘要:去噪扩散概率模型(DDPM)最近在许多生成任务中取得了领先的性能。然而,继承的迭代采样过程成本阻碍了其在文本到语音部署中的应用。通过对扩散模型参数化的初步研究,我们发现以前基于梯度的TTS模型需要数百或数千次迭代才能保证高样本质量,这对加速采样提出了挑战。在这项工作中,我们提出了ProDiff,关于高质量文本到语音的渐进快速扩散模型。与以前估计数据密度梯度的工作不同,ProDiff通过直接预测干净数据来参数化去噪模型,以避免加速采样时出现明显的质量下降。为了通过减少扩散迭代来应对模型收敛挑战,ProDiff通过知识提取来减少目标站点中的数据方差。具体来说,去噪模型使用从N步DDIM教师生成的mel谱图作为训练目标,并将行为提取到具有N/2步的新模型中。因此,它允许TTS模型进行精确预测,并进一步将采样时间减少了几个数量级。我们的评估表明,ProDiff只需2次迭代即可合成高保真mel谱图,同时它使用数百个步骤保持了与最先进模型相比的样本质量和多样性。ProDiff使采样速度比单个NVIDIA 2080Ti GPU上的实时速度快24倍,使扩散模型首次实际适用于文本语音合成部署。我们广泛的烧蚀研究表明,ProDiff中的每种设计都是有效的,并且我们进一步表明,ProDiff可以很容易地扩展到多扬声器设置。音频样本位于{https://ProDiff.github.io/.}摘要:Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinder their applications to text-to-speech deployment. Through the preliminary study on diffusion model parameterization, we find that previous gradient-based TTS models require hundreds or thousands of iterations to guarantee high sample quality, which poses a challenge for accelerating sampling. In this work, we propose ProDiff, on progressive fast diffusion model for high-quality text-to-speech. Unlike previous work estimating the gradient for data density, ProDiff parameterizes the denoising model by directly predicting clean data to avoid distinct quality degradation in accelerating sampling. To tackle the model convergence challenge with decreased diffusion iterations, ProDiff reduces the data variance in the target site via knowledge distillation. Specifically, the denoising model uses the generated mel-spectrogram from an N-step DDIM teacher as the training target and distills the behavior into a new model with N/2 steps. As such, it allows the TTS model to make sharp predictions and further reduces the sampling time by orders of magnitude. Our evaluation demonstrates that ProDiff needs only 2 iterations to synthesize high-fidelity mel-spectrograms, while it maintains sample quality and diversity competitive with state-of-the-art models using hundreds of steps. ProDiff enables a sampling speed of 24x faster than real-time on a single NVIDIA 2080Ti GPU, making diffusion models practically applicable to text-to-speech synthesis deployment for the first time. Our extensive ablation studies demonstrate that each design in ProDiff is effective, and we further show that ProDiff can be easily extended to the multi-speaker setting. Audio samples are available at \url{https://ProDiff.github.io/.}
【2】 MM-ALT: A Multimodal Automatic Lyric Transcription System
标题:MM-ALT:一个多通道的歌词自动转录系统
链接:https://arxiv.org/abs/2207.06127
作者:Xiangming Gu,Longshen Ou,Danielle Ong,Ye Wang机构:Integrative Sciences and Engineering Programme, NUS, Graduate School, National University of Singapore, School of Computing, National University of Singapore备注:Camera ready version. Accepted by ACM Multimedia 2022摘要:自动歌词转录是一个新兴的研究领域,由于其巨大的应用潜力,吸引了语音和音乐信息检索界越来越多的兴趣。然而,由于乐器伴奏和音乐限制,仅使用音频数据的ALT是一项众所周知的困难任务,这会导致语音提示和歌词清晰度的下降。为了应对这一挑战,我们提出了多模式自动歌词转录系统(MM-ALT),以及一个新的数据集N20EM,该数据集包括录音、嘴唇运动视频和表演歌手佩戴的耳塞的惯性测量单元(IMU)数据。我们首先将wav2vec 2.0框架从自动语音识别(ASR)调整到ALT任务。然后,我们提出了一种基于视频的ALT方法和一种基于IMU的语音活动检测(VAD)方法。此外,我们提出了剩余交叉注意力(RCA)机制来融合来自三种模式(即音频、视频和IMU)的数据。实验证明了我们提出的MM-ALT系统的有效性,尤其是在噪声鲁棒性方面。摘要:Automatic lyric transcription (ALT) is a nascent field of study attracting increasing interest from both the speech and music information retrieval communities, given its significant application potential. However, ALT with audio data alone is a notoriously difficult task due to instrumental accompaniment and musical constraints resulting in degradation of both the phonetic cues and the intelligibility of sung lyrics. To tackle this challenge, we propose the MultiModal Automatic Lyric Transcription system (MM-ALT), together with a new dataset, N20EM, which consists of audio recordings, videos of lip movements, and inertial measurement unit (IMU) data of an earbud worn by the performing singer. We first adapt the wav2vec 2.0 framework from automatic speech recognition (ASR) to the ALT task. We then propose a video-based ALT method and an IMU-based voice activity detection (VAD) method. In addition, we put forward the Residual Cross Attention (RCA) mechanism to fuse data from the three modalities (i.e., audio, video, and IMU). Experiments show the effectiveness of our proposed MM-ALT system, especially in terms of noise robustness.
【3】 SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate
标题:萨茨:演讲者文本到语音的吸引人,通过学会分离来学习说话
链接:https://arxiv.org/abs/2207.06011
作者:Nabarun Goswami,Tatsuya Harada机构:The University of Tokyo, Japan, RIKEN, Japan备注:Accepted to Interspeech 2022. Visit this https URL for a demo摘要:文本到语音(TTS)的映射是不确定的,字母的发音可能会根据语境而不同,或者音素可能会因性别、年龄、口音、情绪等各种生理和文体因素而不同。神经说话人嵌入,经过训练以识别或验证说话人通常用于表示这些特征,并将其从参考语音转换为合成语音。另一方面,语音分离是一项具有挑战性的任务,将单个说话人从不同说话人的重叠混合信号中分离出来。说话人吸引子是高维嵌入向量,将每个说话人语音的时频箱拉向自己,同时排斥其他说话人的时频箱。在这项工作中,我们探索了在多说话人TTS合成中使用这些强大的说话人吸引子进行零拍说话人自适应的可能性,并提出了说话人吸引子文本到语音(SATT)。通过各种实验,我们表明,SATT可以从未知目标说话人的参考信号中从文本合成自然语音,该参考信号可能具有低于理想的记录条件,即混响或与其他说话人混合。摘要:The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, emotions, etc. Neural speaker embeddings, trained to identify or verify speakers are typically used to represent and transfer such characteristics from reference speech to synthesized speech. Speech separation on the other hand is the challenging task of separating individual speakers from an overlapping mixed signal of various speakers. Speaker attractors are high-dimensional embedding vectors that pull the time-frequency bins of each speaker's speech towards themselves while repelling those belonging to other speakers. In this work, we explore the possibility of using these powerful speaker attractors for zero-shot speaker adaptation in multi-speaker TTS synthesis and propose speaker attractor text to speech (SATTS). Through various experiments, we show that SATTS can synthesize natural speech from text from an unseen target speaker's reference signal which might have less than ideal recording conditions, i.e. reverberations or mixed with other speakers.
【4】 Cross-Age Speaker Verification: Learning Age-Invariant Speaker Embeddings
标题:跨年龄说话人验证:学习年龄不变的说话人嵌入
链接:https://arxiv.org/abs/2207.05929
作者:Xiaoyi Qin,Na Li,Chao Weng,Dan Su,Ming Li机构:School of Computer Science, Wuhan University, Wuhan, China, Data Science Research Center, Duke Kunshan University, Kunshan, China, Tencent AI Lab, Shenzhen, China备注:Accepted by Interspeech2022摘要:近年来,说话人自动识别技术取得了显著的进展。然而,由于相关数据不足,跨年龄说话人验证的研究很少。在本文中,我们基于VoxCeleb数据集挖掘跨年龄测试集,并提出了年龄不变说话人表示(AISR)学习方法。由于VoxCeleb是从YouTube平台收集的,因此数据集本质上由跨年龄数据组成。然而,元数据不包含说话人年龄标签。因此,我们采用人脸年龄估计方法从相关的视觉数据中预测说话人的年龄值,然后用估计的年龄标记音频记录。我们在VoxCeleb(Vox-CA)上构建了多个交叉年龄测试集,故意选择年龄差距较大的阳性试验。此外,在选择与Vox-H病例对应的阴性配对时,还考虑了国籍和性别的影响。基线系统性能从Vox-H测试集的1.939\%EER下降到Vox-CA20测试集的10.419\%,这表明跨年龄场景有多困难。因此,我们提出了一种年龄解耦对抗学习(ADAL)方法,以缓解年龄差距的负面影响,减少类内方差。我们的方法优于基线系统,在Vox-CA20测试集上相关EER减少了10%以上。源代码和试用资源可在https://github.com/qinxiaoyi/Cross-Age_Speaker_Verification摘要:Automatic speaker verification has achieved remarkable progress in recent years. However, there is little research on cross-age speaker verification (CASV) due to insufficient relevant data. In this paper, we mine cross-age test sets based on the VoxCeleb dataset and propose our age-invariant speaker representation(AISR) learning method. Since the VoxCeleb is collected from the YouTube platform, the dataset consists of cross-age data inherently. However, the meta-data does not contain the speaker age label. Therefore, we adopt the face age estimation method to predict the speaker age value from the associated visual data, then label the audio recording with the estimated age. We construct multiple Cross-Age test sets on VoxCeleb (Vox-CA), which deliberately select the positive trials with large age-gap. Also, the effect of nationality and gender is considered in selecting negative pairs to align with Vox-H cases. The baseline system performance drops from 1.939\% EER on the Vox-H test set to 10.419\% on the Vox-CA20 test set, which indicates how difficult the cross-age scenario is. Consequently, we propose an age-decoupling adversarial learning (ADAL) method to alleviate the negative effect of the age gap and reduce intra-class variance. Our method outperforms the baseline system by over 10\% related EER reduction on the Vox-CA20 test set. The source code and trial resources are available on https://github.com/qinxiaoyi/Cross-Age_Speaker_Verification
【5】 Online Target Speaker Voice Activity Detection for Speaker Diarization
标题:用于说话人二值化的在线目标说话人语音活动检测
链接:https://arxiv.org/abs/2207.05920
作者:Weiqing Wang,Qingjian Lin,Ming Li机构:Department of Electrical & Computer Engineering, Duke University, Durham, NC , USA, Data Science Research Center, Duke Kunshan University, Kunshan , PR China, AI Lab, Lenovo Research, Beijing , PR China备注:Accepted by Interspeech 2022摘要:本文提出了一种用于说话人日记任务的在线目标说话人语音活动检测系统,该系统不需要来自基于聚类的日记系统的先验知识来获得目标说话人嵌入。首先,我们采用基于ResNet的前端模型来提取每个信号块的帧级扬声器嵌入。接下来,我们基于这些帧级说话人嵌入和先前估计的目标说话人嵌入来预测每个说话人的检测状态。然后,通过根据当前块中的预测聚合这些帧级说话人嵌入来更新目标说话人嵌入。我们迭代地提取每个块的结果,并更新目标说话人嵌入,直到到达信号的末尾。实验结果表明,该方法优于AliMeeting数据集上基于离线聚类的日记系统。摘要:This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target speaker embeddings. First, we employ a ResNet-based front-end model to extract the frame-level speaker embeddings for each coming block of a signal. Next, we predict the detection state of each speaker based on these frame-level speaker embeddings and the previously estimated target speaker embedding. Then, the target speaker embeddings are updated by aggregating these frame-level speaker embeddings according to the predictions in the current block. We iteratively extract the results for each block and update the target speaker embedding until reaching the end of the signal. Experimental results show that the proposed method is better than the offline clustering-based diarization system on the AliMeeting dataset.
【6】 A Cyclical Approach to Synthetic and Natural Speech Mismatch Refinement of Neural Post-filter for Low-cost Text-to-speech System
标题:低成本文语转换系统中人工语音与自然语音失配的循环优化神经后置滤波器
链接:https://arxiv.org/abs/2207.05913
作者:Yi-Chiao Wu,Patrick Lumban Tobing,Kazuki Yasuhara,Noriyuki Matsunaga,Yamato Ohtani,Tomoki Toda机构:The written permission of Cambridge University Press must be obtained for commercial re-use, Nagoya University备注:15 pages, 7 figures, 10 tables摘要:由于神经网络的快速发展,基于神经的文本到语音(TTS)系统实现了非常高保真的语音生成。然而,巨大的标记语料库和高计算成本要求限制了小公司或个人开发高保真TTS系统的可能性。另一方面,在基于神经的TTS系统中广泛用于语音生成的神经声码器可以用相对较小的未标记语料库进行训练。因此,在本文中,我们探索了一个通用框架,用于开发使用神经声码器的低成本TTS系统的神经后滤波器(NPF)。提出了一种循环方法来解决开发NPF时的声学和时间不匹配(AM和TM)。进行了客观和主观评估,以证明AM和TM问题以及拟议框架的有效性。摘要:Neural-based text-to-speech (TTS) systems achieve very high-fidelity speech generation because of the rapid neural network developments. However, the huge labeled corpus and high computation cost requirements limit the possibility of developing a high-fidelity TTS system by small companies or individuals. On the other hand, a neural vocoder, which has been widely adopted for the speech generation in neural-based TTS systems, can be trained with a relatively small unlabeled corpus. Therefore, in this paper, we explore a general framework to develop a neural post-filter (NPF) for low-cost TTS systems using neural vocoders. A cyclical approach is proposed to tackle the acoustic and temporal mismatches (AM and TM) of developing an NPF. Both objective and subjective evaluations have been conducted to demonstrate the AM and TM problems and the effectiveness of the proposed framework.
【7】 Masked Autoencoders that Listen
标题:可监听的蒙面自动编码器
链接:https://arxiv.org/abs/2207.06405
作者:Po-Yao,Huang,Hu Xu,Juncheng Li,Alexei Baevski,Michael Auli,Wojciech Galuba,Florian Metze,Christoph Feichtenhofer机构:FAIR, Meta AI, Carnegie Mellon University摘要:本文研究了基于图像的掩码自动编码器(MAE)的简单扩展,以从音频频谱图中进行自监督表示学习。根据MAE中的Transformer编码器-解码器设计,我们的音频MAE首先以高掩蔽率对音频频谱图块进行编码,通过编码器层仅向非掩蔽令牌馈电。然后,解码器对填充有掩码令牌的编码上下文进行重新排序和解码,以重建输入频谱图。我们发现,在解码器中加入局部窗口注意是有益的,因为音频频谱在局部时间和频带上高度相关。然后,我们在目标数据集上以较低的掩蔽率微调编码器。根据经验,音频MAE在六个音频和语音分类任务上设定了最新的性能,优于其他使用外部监督预训练的最新模型。代码和模型将在https://github.com/facebookresearch/AudioMAE.摘要:This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only the non-masked tokens through encoder layers. The decoder then re-orders and decodes the encoded context padded with mask tokens, in order to reconstruct the input spectrogram. We find it beneficial to incorporate local window attention in the decoder, as audio spectrograms are highly correlated in local time and frequency bands. We then fine-tune the encoder with a lower masking ratio on target datasets. Empirically, Audio-MAE sets new state-of-the-art performance on six audio and speech classification tasks, outperforming other recent models that use external supervised pre-training. The code and models will be at https://github.com/facebookresearch/AudioMAE.
【8】 Polyphonic sound event detection for highly dense birdsong scenes
标题:用于高密度鸟鸣场景的复音事件检测
链接:https://arxiv.org/abs/2207.06349
作者:Alberto García Arroba Parrilla,Dan Stowell机构:. Tilburg University, the Netherlands, . Naturalis, Biodiversity, Center, Leiden摘要:日出前一个小时,你可以体验黎明合唱,来自不同物种的鸟类一起唱歌。在这种情况下,与重叠声源的数量一样,很容易出现高水平的复调,从而导致复杂的声学结果。声音事件检测任务分析声音场景,以识别发生的事件及其各自的时间信息。然而,高密度场景可能难以处理,并且尚未进行深入研究。在这里,我们使用卷积递归神经网络(CRNN)展示了在处理更高的复调时如何检测鸟鸣复调场景,以及这种模型如何有效地面对多达10只重叠鸟类的非常密集的场景。我们发现,使用更密集的示例(即更高的复调)训练的模型与在训练集中使用更简单样本的模型学习速度相似。此外,使用密度最大的样本训练的模型对所有复调都保持一致的分数,而使用密度最小的样本训练的模型随着复调的增加而退化。我们的结果表明,使用CRNN可以处理高密度声学场景。我们希望这项研究可以作为研究高密度鸟类场景的起点,例如黎明合唱或其他密集声学问题。摘要:One hour before sunrise, one can experience the dawn chorus where birds from different species sing together. In this scenario, high levels of polyphony, as in the number of overlapping sound sources, are prone to happen resulting in a complex acoustic outcome. Sound Event Detection (SED) tasks analyze acoustic scenarios in order to identify the occurring events and their respective temporal information. However, highly dense scenarios can be hard to process and have not been studied in depth. Here we show, using a Convolutional Recurrent Neural Network (CRNN), how birdsong polyphonic scenarios can be detected when dealing with higher polyphony and how effectively this type of model can face a very dense scene with up to 10 overlapping birds. We found that models trained with denser examples (i.e., higher polyphony) learn at a similar rate as models that used simpler samples in their training set. Additionally, the model trained with the densest samples maintained a consistent score for all polyphonies, while the model trained with the least dense samples degraded as the polyphony increased. Our results demonstrate that highly dense acoustic scenarios can be dealt with using CRNNs. We expect that this study serves as a starting point for working on highly populated bird scenarios such as dawn chorus or other dense acoustic problems.
【9】 Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech
标题:可控无损非自回归端到端文语转换
链接:https://arxiv.org/abs/2207.06088
作者:Zhengxi Liu,Qiao Tian,Chenxu Hu,Xudong Liu,Menglin Wu,Yuping Wang,Hang Zhao,Yuxuan Wang机构:Speech, Audio and Music Intelligence (SAMI), ByteDance, IIIS, Tsinghua University摘要:最近的一些研究已经证明了单阶段神经文本到语音的可行性,它不需要生成mel频谱,而是直接从文本生成原始波形。单阶段文本到语音通常面临两个问题:a)由于多个语音变体而导致的一对多映射问题和b)由于在训练期间缺乏对地面真实声学特征的监督而导致高频重建不足。为了解决a)问题并生成更具表现力的语音,我们提出了一种新的音素级韵律建模方法,该方法基于具有归一化流的变分自动编码器来建模语音中的潜在韵律信息。我们还使用韵律预测器来支持端到端的表达性语音合成。此外,我们提出了双并行自动编码器,在训练期间引入对地面真实声学特征的监督,以解决b)问题,使我们的模型能够生成高质量的语音。我们在内部表达性英语数据集上比较了合成质量与最先进的文本语音系统。定性和定量评估都证明了我们的方法在无损语音生成方面的优越性和鲁棒性,同时也显示出强大的韵律建模能力。摘要:Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces two problems: a) the one-to-many mapping problem due to multiple speech variations and b) insufficiency of high frequency reconstruction due to the lack of supervision of ground-truth acoustic features during training. To solve the a) problem and generate more expressive speech, we propose a novel phoneme-level prosody modeling method based on a variational autoencoder with normalizing flows to model underlying prosodic information in speech. We also use the prosody predictor to support end-to-end expressive speech synthesis. Furthermore, we propose the dual parallel autoencoder to introduce supervision of the ground-truth acoustic features during training to solve the b) problem enabling our model to generate high-quality speech. We compare the synthesis quality with state-of-the-art text-to-speech systems on an internal expressive English dataset. Both qualitative and quantitative evaluations demonstrate the superiority and robustness of our method for lossless speech generation while also showing a strong capability in prosody modeling.
【10】 Subband-based Generative Adversarial Network for Non-parallel Many-to-many Voice Conversion
标题:基于子带的非并行多对多语音转换生成对抗性网络
链接:https://arxiv.org/abs/2207.06057
作者:Jian Ma,Zhedong Zheng,Hao Fei,Feng Zheng,Tat-seng Chua,Yi Yang机构:Southern University of Science and Technology摘要:语音转换是生成具有源内容和目标语音风格的新语音。在本文中,我们重点关注一种通用设置,即非并行多对多语音转换,这与真实场景非常接近。顾名思义,非并行多对多语音转换不需要成对的源语音和参考语音,可以应用于任意语音传输。近年来,生成对抗网络(GAN)和条件变分自动编码器(CVAE)等技术在这一领域取得了长足的进展。然而,由于语音转换的复杂性,转换语音的风格相似性仍然不能令人满意。受mel谱图固有结构的启发,我们提出了一种新的语音转换框架,即基于子带的生成式语音转换对抗网络(SGAN-VC)。SGAN-VC通过明确利用不同子带之间的空间特征,分别转换源语音的每个子带内容。SGAN-VC包含一个样式编码器、一个内容编码器和一个解码器。特别是,风格编码器网络设计用于学习目标说话人不同子带的风格代码。内容编码器网络可以捕获源语音上的内容信息。最后,解码器生成特定子带内容。此外,我们提出了一个基音偏移模块来微调源说话人的基音,使转换后的音调更准确和易于解释。大量实验表明,无论是在可见数据还是不可见数据上,该方法在VCTK语料库和Aisell3数据集上都取得了最先进的性能。此外,SGAN-VC在不可见数据上的内容可懂度甚至超过了有ASR网络辅助的StarGANv2-VC。摘要:Voice conversion is to generate a new speech with the source content and a target voice style. In this paper, we focus on one general setting, i.e., non-parallel many-to-many voice conversion, which is close to the real-world scenario. As the name implies, non-parallel many-to-many voice conversion does not require the paired source and reference speeches and can be applied to arbitrary voice transfer. In recent years, Generative Adversarial Networks (GANs) and other techniques such as Conditional Variational Autoencoders (CVAEs) have made considerable progress in this field. However, due to the sophistication of voice conversion, the style similarity of the converted speech is still unsatisfactory. Inspired by the inherent structure of mel-spectrogram, we propose a new voice conversion framework, i.e., Subband-based Generative Adversarial Network for Voice Conversion (SGAN-VC). SGAN-VC converts each subband content of the source speech separately by explicitly utilizing the spatial characteristics between different subbands. SGAN-VC contains one style encoder, one content encoder, and one decoder. In particular, the style encoder network is designed to learn style codes for different subbands of the target speaker. The content encoder network can capture the content information on the source speech. Finally, the decoder generates particular subband content. In addition, we propose a pitch-shift module to fine-tune the pitch of the source speaker, making the converted tone more accurate and explainable. Extensive experiments demonstrate that the proposed approach achieves state-of-the-art performance on VCTK Corpus and AISHELL3 datasets both qualitatively and quantitatively, whether on seen or unseen data. Furthermore, the content intelligibility of SGAN-VC on unseen data even exceeds that of StarGANv2-VC with ASR network assistance.
【11】 Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech Recognition
标题:面向端到端语音识别的视觉上下文驱动的音频特征增强
链接:https://arxiv.org/abs/2207.06020
作者:Joanna Hong,Minsu Kim,Daehun Yoo,Yong Man Ro机构:KAIST, Daejeon, South Korea, Genesis Lab Inc., Seoul, South Korea备注:Accepted at Interspeech 2022摘要:本文主要研究设计一种抗噪声的端到端视听语音识别系统。为此,我们提出了视觉上下文驱动的音频特征增强模块(V-CAFE),以借助视听对应来增强输入的噪声音频语音。提出的V-CAFE旨在捕捉嘴唇运动的过渡,即视觉语境,并通过考虑获得的视觉语境生成降噪掩码。通过上下文相关建模,可以细化视位到音素映射中的歧义,以生成掩码。噪声表示用降噪掩码掩盖,从而增强音频特征。将增强的音频特征与视觉特征融合,并将其用于语音识别的编码器-解码器模型,该模型由构象器和变换器组成。我们表明,使用V-CAFE的端到端AVSR可以进一步提高AVSR的噪声鲁棒性。使用两个最大的视听数据集LRS2和LRS3,在噪声语音识别和重叠语音识别实验中评估了该方法的有效性。摘要:This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a help of audio-visual correspondence. The proposed V-CAFE is designed to capture the transition of lip movements, namely visual context and to generate a noise reduction mask by considering the obtained visual context. Through context-dependent modeling, the ambiguity in viseme-to-phoneme mapping can be refined for mask generation. The noisy representations are masked out with the noise reduction mask resulting in enhanced audio features. The enhanced audio features are fused with the visual features and taken to an encoder-decoder model composed of Conformer and Transformer for speech recognition. We show the proposed end-to-end AVSR with the V-CAFE can further improve the noise-robustness of AVSR. The effectiveness of the proposed method is evaluated in noisy speech recognition and overlapped speech recognition experiments using the two largest audio-visual datasets, LRS2 and LRS3.
【12】 Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
标题:神经文语转换中文本驱动的情绪风格控制与交叉语者风格迁移
链接:https://arxiv.org/abs/2207.06000
作者:Yookyung Shin,Younggun Lee,Suhee Jo,Yeongtae Hwang,Taesu Kim备注:Accepted to Interspeech 2022摘要:近年来,具有表现力的文本到语音的性能有所提高。然而,合成语音的风格控制通常局限于离散的情感类别,需要目标说话人以目标风格记录训练数据。在许多实际情况下,用户可能没有在目标情感中记录参考语音,但仍然有兴趣通过键入所需情感风格的文本描述来控制语音风格。在本文中,我们提出了一种基于文本的界面,用于多说话人TTS中的情感风格控制和跨说话人风格转换。我们提出了双模风格编码器,该编码器使用预训练的语言模型来模拟文本描述嵌入和语音风格嵌入之间的语义关系。为了进一步改善不相交、多风格数据集上的跨说话人风格转换,我们提出了新的风格损失。实验结果表明,我们的模型可以生成高质量的表达性语音,即使是在看不见的风格。摘要:Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the target style. In many practical situations, users may not have reference speech recorded in target emotion but still be interested in controlling speech style just by typing text description of desired emotional style. In this paper, we propose a text-based interface for emotional style control and cross-speaker style transfer in multi-speaker TTS. We propose the bi-modal style encoder which models the semantic relationship between text description embedding and speech style embedding with a pretrained language model. To further improve cross-speaker style transfer on disjoint, multi-style datasets, we propose the novel style loss. The experimental results show that our model can generate high-quality expressive speech even in unseen style.
【13】 NEC: Speaker Selective Cancellation via Neural Enhanced Ultrasound Shadowing
标题:NEC:基于神经增强超声跟踪的说话人选择性对消
链接:https://arxiv.org/abs/2207.05848
作者:Hanqing Guo,Chenning Li,Lingkun Li,Zhichao Cao,Qiben Yan,Li Xiao机构:Department of Computer Science and Engineering, Michigan State University, ∗These authors contributed equally摘要:在本文中,我们提出了NEC(神经增强消除),这是一种防御机制,可防止未经授权的麦克风捕捉目标说话人的声音。与现有的基于置乱的音频抵消方法相比,NEC可以有选择地从混合语音中删除目标说话人的语音,而不会对其他语音造成干扰。具体来说,对于目标说话人,我们设计了一个深度神经网络(DNN)模型,从其参考音频中提取特定于说话人但与话语无关的高级语音特征。当麦克风录制时,DNN生成阴影声音以实时取消目标声音。此外,我们将可听的阴影声音调制成超声频率,使人类听不到。通过利用麦克风电路的非线性,麦克风可以准确解码阴影声音以消除目标语音。我们在不同设置下使用8个智能手机麦克风全面实施和评估NEC。结果表明,NEC在不干扰其他用户正常对话的情况下,有效地使目标扬声器在麦克风前静音。摘要:In this paper, we propose NEC (Neural Enhanced Cancellation), a defense mechanism, which prevents unauthorized microphones from capturing a target speaker's voice. Compared with the existing scrambling-based audio cancellation approaches, NEC can selectively remove a target speaker's voice from a mixed speech without causing interference to others. Specifically, for a target speaker, we design a Deep Neural Network (DNN) model to extract high-level speaker-specific but utterance-independent vocal features from his/her reference audios. When the microphone is recording, the DNN generates a shadow sound to cancel the target voice in real-time. Moreover, we modulate the audible shadow sound onto an ultrasound frequency, making it inaudible for humans. By leveraging the non-linearity of the microphone circuit, the microphone can accurately decode the shadow sound for target voice cancellation. We implement and evaluate NEC comprehensively with 8 smartphone microphones in different settings. The results show that NEC effectively mutes the target speaker at a microphone without interfering with other users' normal conversations.
【14】 Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices
标题:适用于低资源设备的二进制神经网络提取非语义语音嵌入
链接:https://arxiv.org/abs/2207.05784
作者:Harlin Lee,Aaqib Saeed机构:edu) is with the Department of Mathematics, University of California摘要:本文介绍了BRILLsson,一种新的基于二进制神经网络的表示学习模型,用于广泛的非语义语音任务。我们通过从一个大型实值TRILLsson模型中提取知识来训练该模型,其中只有一小部分数据集用于训练TRILLsson。由此产生的BRILLsson模型大小仅为2MB,延迟小于8ms,适合部署在可穿戴设备等低资源设备中。我们在八个基准任务(包括但不限于口语识别、情绪识别、健康状况诊断和关键词识别)上评估了BRILLsson,并证明了我们提出的超轻和低延迟模型与大规模模型的性能相同。摘要:This work introduces BRILLsson, a novel binary neural network-based representation learning model for a broad range of non-semantic speech tasks. We train the model with knowledge distillation from a large and real-valued TRILLsson model with only a fraction of the dataset used to train TRILLsson. The resulting BRILLsson models are only 2MB in size with a latency less than 8ms, making them suitable for deployment in low-resource devices such as wearables. We evaluate BRILLsson on eight benchmark tasks (including but not limited to spoken language identification, emotion recognition, heath condition diagnosis, and keyword spotting), and demonstrate that our proposed ultra-light and low-latency models perform as well as large-scale models.
机器翻译,仅供参考