【1】 A cross-corpus study on speech emotion recognition
标题:语音情感识别的跨语料库研究
链接:https://arxiv.org/abs/2207.02104
作者:Rosanna Milner,Md Asif Jalal,Raymond W. M. Ng,Thomas Hain摘要:对于语音情感数据集,很难获得大量可靠的数据,与日常生活中表现较少的情感相比,行为情感可能过高。最近,创建了具有自然情感的更大数据集。本研究不是忽略较小的行为数据集,而是调查从行为情绪中学习的信息是否有助于检测自然情绪。跨语料库研究大多考虑跨语言甚至跨年龄的数据集,不同的情感注释方法会导致性能下降,这会带来困难。为了保持一致,我们考虑了四个成人英语数据集,包括行为情感、诱发情感和自然情感。提出了一种最先进的模型来准确地研究性能下降。该系统包括一个双向LSTM,该LSTM具有一个注意力机制,用于跨数据集对情绪进行分类。实验以跨语料库和多领域的方式研究了训练模型的效果,结果表明信息传递并不成功。域外模型,然后适应缺失的数据集,域对抗训练(DAT)被证明更适合推广到跨数据集的情绪。这表明了从行为数据集到那些具有更自然情感的数据集的积极信息传递,以及在不同语料库上训练的益处。摘要:For speech emotion datasets, it has been difficult to acquire large quantities of reliable data and acted emotions may be over the top compared to less expressive emotions displayed in everyday life. Lately, larger datasets with natural emotions have been created. Instead of ignoring smaller, acted datasets, this study investigates whether information learnt from acted emotions is useful for detecting natural emotions. Cross-corpus research has mostly considered cross-lingual and even cross-age datasets, and difficulties arise from different methods of annotating emotions causing a drop in performance. To be consistent, four adult English datasets covering acted, elicited and natural emotions are considered. A state-of-the-art model is proposed to accurately investigate the degradation of performance. The system involves a bi-directional LSTM with an attention mechanism to classify emotions across datasets. Experiments study the effects of training models in a cross-corpus and multi-domain fashion and results show the transfer of information is not successful. Out-of-domain models, followed by adapting to the missing dataset, and domain adversarial training (DAT) are shown to be more suitable to generalising to emotions across datasets. This shows positive information transfer from acted datasets to those with more natural emotions and the benefits from training on different corpora.
【2】 WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training
标题:WeSinger 2:基于多歌手条件对抗训练的完全并行歌声合成
链接:https://arxiv.org/abs/2207.01886
作者:Zewang Zhang,Yibin Zheng,Xinhui Li,Li Lu备注:5 pages, 2 figures, 4 tables摘要:本文旨在介绍一种鲁棒的歌声合成(SVS)系统,利用对抗训练策略有效地生成高质量的歌声。一方面,我们设计了简单但通用的随机区域条件鉴别器来帮助监督声学模型,它可以有效地避免基于持续时间分配Transformer的声学模型的过度平滑频谱预测。另一方面,我们巧妙地将频谱图与帧级线性插值F0序列相结合,作为神经声码器的输入,然后借助波形域中的多个对抗性鉴别器和频域中的多尺度距离函数对神经声码器进行优化。实验结果和烧蚀研究表明,与我们以前的自回归工作相比,我们的新系统可以通过对不同的歌唱数据集进行微调,有效地产生高质量的歌唱声音,这些数据集覆盖几分钟到几小时。网上有一些合成的歌唱样本[https://zzw922cn.github.io/wesinger2 ].摘要:This paper aims to introduce a robust singing voice synthesis (SVS) system to produce high-quality singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random area conditional discriminators to help supervise the acoustic model, which can effectively avoid the over-smoothed spectrogram prediction by the duration-allocated Transformer-based acoustic model. On the other hand, we subtly combined the spectrogram with the frame-level linearly-interpolated F0 sequence as the input for the neural vocoder, which is then optimized with the help of multiple adversarial discriminators in the waveform domain and multi-scale distance functions in the frequency domain. The experimental results and ablation studies concluded that, compared with our previous auto-regressive work, our new system can produce high-quality singing voices efficiently by fine-tuning on different singing datasets covering several minutes to a few hours. Some synthesized singing samples are available online [https://zzw922cn.github.io/wesinger2 ].
【3】 Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice Conversion
标题:Glow-WaveGan 2:高质量零激发文本到语音合成和任意到任意语音转换
链接:https://arxiv.org/abs/2207.01832
作者:Yi Lei,Shan Yang,Jian Cong,Lei Xie,Dan Su摘要:语音生成的零拍场景旨在仅用目标说话人的一句话合成一种新的看不见的语音。虽然在零炮场景中,在声学建模和声码器这两个阶段都存在适应新语音的挑战,但以前的工作通常只从一个阶段考虑这个问题。在本文中,我们将之前的Glow-WaveGAN扩展到Glow-WaveGAN 2,旨在从两个阶段解决高质量零炮文本到语音和任意到任意语音转换的问题。我们首先建立了一个通用的WaveGAN模型,用于提取语音的潜在分布$p(z)$并从中重建波形。然后,基于流的声学模型只需要从文本中学习相同的$p(z)$,这自然避免了声学模型和声码器之间的不匹配,从而在不进行模型微调的情况下生成高质量的语音。基于连续说话人空间和流的可逆性,可以获得任何说话人的条件分布,从而可以进一步为新说话人生成高质量的零拍语音。我们特别研究了两种构建说话人空间的方法,即预训练说话人编码器和联合训练说话人编码器。在LibriTTS语料库和VTCK语料库上进行的TTS和VC实验证明了Glow-WaveGAN-2的优越性。摘要:The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker. Although the challenges of adapting new voices in zero-shot scenario exist in both stages -- acoustic modeling and vocoder, previous works usually consider the problem from only one stage. In this paper, we extend our previous Glow-WaveGAN to Glow-WaveGAN 2, aiming to solve the problem from both stages for high-quality zero-shot text-to-speech and any-to-any voice conversion. We first build a universal WaveGAN model for extracting latent distribution $p(z)$ of speech and reconstructing waveform from it. Then a flow-based acoustic model only needs to learn the same $p(z)$ from texts, which naturally avoids the mismatch between the acoustic model and the vocoder, resulting in high-quality generated speech without model fine-tuning. Based on a continuous speaker space and the reversible property of flows, the conditional distribution can be obtained for any speaker, and thus we can further conduct high-quality zero-shot speech generation for new speakers. We particularly investigate two methods to construct the speaker space, namely pre-trained speaker encoder and jointly-trained speaker encoder. The superiority of Glow-WaveGAN 2 has been proved through TTS and VC experiments conducted on LibriTTS corpus and VTCK corpus.
【4】 Backend Ensemble for Speaker Verification and Spoofing Countermeasure
标题:用于说话人验证的后端集成及欺骗对策
链接:https://arxiv.org/abs/2207.01802
作者:Li Zhang,Yue Li,Huan Zhao,Qing Wang,Lei Xie摘要:本文描述了提交给2022年防欺骗说话人验证挑战的NPU系统。我们从三个方面特别关注用于说话人验证和欺骗对策的\textit{backend-embly}。首先,除了简单的级联之外,我们还提出了循环矩阵变换和堆叠用于说话人嵌入和对策嵌入。通过对新定义的循环嵌入的叠加操作,我们几乎探索了说话人嵌入和对策嵌入之间的所有可能交互。其次,我们尝试使用不同的卷积神经网络,将嵌入的显著区域选择性地融合到具有卷积核的通道中。最后,我们设计了一维卷积神经网络中的并行注意,以学习通道维度中的全局相关性以及特征维度中的重要部分。同时,我们在2D卷积神经网络中嵌入挤压和激发注意,以学习说话人嵌入和对策嵌入之间的全局依赖性。实验结果表明,上述方法都是有效的。经过上述方法增强的四个训练有素的模型融合后,我们获得的最佳SASV-EER、SPF-EER和SV-EER在评估集上分别为0.559%、0.354%和0.857%。加上上述贡献,我们的提交系统在这项挑战中获得第五名。摘要:This paper describes the NPU system submitted to Spoofing Aware Speaker Verification Challenge 2022. We particularly focus on the \textit{backend ensemble} for speaker verification and spoofing countermeasure from three aspects. Firstly, besides simple concatenation, we propose circulant matrix transformation and stacking for speaker embeddings and countermeasure embeddings. With the stacking operation of newly-defined circulant embeddings, we almost explore all the possible interactions between speaker embeddings and countermeasure embeddings. Secondly, we attempt different convolution neural networks to selectively fuse the embeddings' salient regions into channels with convolution kernels. Finally, we design parallel attention in 1D convolution neural networks to learn the global correlation in channel dimensions as well as to learn the important parts in feature dimensions. Meanwhile, we embed squeeze-and-excitation attention in 2D convolutional neural networks to learn the global dependence among speaker embeddings and countermeasure embeddings. Experimental results demonstrate that all the above methods are effective. After fusion of four well-trained models enhanced by the mentioned methods, the best SASV-EER, SPF-EER and SV-EER we achieve are 0.559\%, 0.354\% and 0.857\% on the evaluation set respectively. Together with the above contributions, our submission system achieves the fifth place in this challenge.
【5】 An adaptive music generation architecture for games based on the deep learning Transformer mode
标题:一种基于深度学习变换模式的自适应游戏音乐生成架构
链接:https://arxiv.org/abs/2207.01698
作者:Gustavo Amaral Costa dos Santos,Augusto Baffa,Jean-Pierre Briot,Bruno Feijó,Antonio Luz Furtado摘要:本文提出了一种基于Transformer深度学习模型的视频游戏音乐生成架构。该系统按照目前设计视频游戏音乐的作曲家使用的标准分层策略,生成不同层次的音乐。根据唤醒价模型,音乐适应演奏者的心理环境。我们的动机是根据玩家的口味定制音乐,玩家可以通过一组训练音乐示例来选择自己喜欢的音乐风格。我们讨论了当前的局限性和未来的前景,例如音乐组件的协作和交互控制。摘要:This paper presents an architecture for generating music for video games based on the Transformer deep learning model. The system generates music in various layers, following the standard layering strategy currently used by composers designing video game music. The music is adaptive to the psychological context of the player, according to the arousal-valence model. Our motivation is to customize music according to the player's tastes, who can select his preferred style of music through a set of training examples of music. We discuss current limitations and prospects for the future, such as collaborative and interactive control of the musical components.
【6】 Stochastic Restoration of Heavily Compressed Musical Audio using Generative Adversarial Networks
标题:基于产生式对抗网络的重压缩音乐音频随机恢复
链接:https://arxiv.org/abs/2207.01667
作者:Stefan Lattner,Javier Nistal摘要:有损音频编解码器通过删除人类感知中往往听不见的信息来压缩(和解压缩)数字音频流。在高压缩率下,此类编解码器可能会在音频信号中引入各种损伤。许多作品都使用深度学习技术解决了音频增强和压缩伪影消除的问题。然而,只有少数作品解决了音乐领域中严重压缩音频信号的恢复问题。在这种情况下,没有唯一的解决方案来恢复原始信号。因此,在本研究中,我们测试了生成对抗网络(GAN)架构的随机生成器。这种随机发生器以高度压缩的音乐音频信号为条件,有朝一日可能会产生与高质量版本无法区分的输出。因此,本研究可能有助于深入了解更高效的音乐数据存储和传输。我们在16、32和64kbit/s的MP3压缩音频信号上训练随机和确定性生成器。我们利用客观指标和听力测试对不同的实验进行了广泛的评估。我们发现,与MP3版本相比,该模型可以提高16和32kbit/s音频信号的质量,并且随机生成器能够生成比确定性生成器更接近原始信号的输出。摘要:Lossy audio codecs compress (and decompress) digital audio streams by removing information that tends to be inaudible in human perception. Under high compression rates, such codecs may introduce a variety of impairments in the audio signal. Many works have tackled the problem of audio enhancement and compression artifact removal using deep learning techniques. However, only a few works tackle the restoration of heavily compressed audio signals in the musical domain. In such a scenario, there is no unique solution for the restoration of the original signal. Therefore, in this study, we test a stochastic generator of a Generative Adversarial Network (GAN) architecture for this task. Such a stochastic generator, conditioned on highly compressed musical audio signals, could one day generate outputs indistinguishable from high-quality releases. Therefore, the present study may yield insights into more efficient musical data storage and transmission. We train stochastic and deterministic generators on MP3-compressed audio signals with 16, 32, and 64 kbit/s. We perform an extensive evaluation of the different experiments utilizing objective metrics and listening tests. We find that the models can improve the quality of the audio signals over the MP3 versions for 16 and 32 kbit/s and that the stochastic generators are capable of generating outputs that are closer to the original signals than those of the deterministic generators.
eess.AS音频处理
【1】 Relating the fundamental frequency of speech with EEG using a dilated convolutional network
标题:利用扩张卷积网络将语音基频与脑电信号联系起来
链接:https://arxiv.org/abs/2207.01963
作者:Corentin Puffay,Jana Van Canneyt,Jonas Vanthornhout,Hugo Van hamme,Tom Francart备注:Accepted for Interspeech 2022摘要:为了研究语音在大脑中的处理方式,我们可以建模自然语音信号的特征与相应记录的脑电图(EEG)之间的关系。通常,线性模型用于回归任务。可以预测脑电图,也可以重建语音,预测信号和实际信号之间的相关性用于测量大脑的解码能力。然而,考虑到大脑的非线性性质,线性模型的建模能力是有限的。最近的研究引入了非线性模型,将语音包络与脑电图联系起来。我们着手包括未在信封中编码的语音的其他特征,尤其是语音的基频(f0)。F0是一种主要在脑干至中脑水平编码的高频特征。我们提出了一个扩展卷积模型,以提供f0神经跟踪的证据。我们表明,f0和语音包络的组合提高了最先进的基于包络的模型的性能。这表明扩张卷积模型可以从f0和包络中提取非冗余信息。我们还展示了扩张卷积模型推广到训练期间未包括的受试者的能力。后一个发现将加速基于f0的听力诊断。摘要:To investigate how speech is processed in the brain, we can model the relation between features of a natural speech signal and the corresponding recorded electroencephalogram (EEG). Usually, linear models are used in regression tasks. Either EEG is predicted, or speech is reconstructed, and the correlation between predicted and actual signal is used to measure the brain's decoding ability. However, given the nonlinear nature of the brain, the modeling ability of linear models is limited. Recent studies introduced nonlinear models to relate the speech envelope to EEG. We set out to include other features of speech that are not coded in the envelope, notably the fundamental frequency of the voice (f0). F0 is a higher-frequency feature primarily coded at the brainstem to midbrain level. We present a dilated-convolutional model to provide evidence of neural tracking of the f0. We show that a combination of f0 and the speech envelope improves the performance of a state-of-the-art envelope-based model. This suggests the dilated-convolutional model can extract non-redundant information from both f0 and the envelope. We also show the ability of the dilated-convolutional model to generalize to subjects not included during training. This latter finding will accelerate f0-based hearing diagnosis.
【2】 DEFORMER: Coupling Deformed Localized Patterns with Global Context for Robust End-to-end Speech Recognition
标题:DEFORMER:将变形的局部化模式与全局上下文相结合实现健壮的端到端语音识别
链接:https://arxiv.org/abs/2207.01732
作者:Jiamin Xie,John H. L. Hansen备注:Accepted to Interspeech 2022摘要:卷积神经网络(CNN)利用局部时频模式,大大提高了语音识别性能。但是,传统的CNN操作假设这些模式出现在对称和刚性内核中。这引发了一个问题:不对称内核怎么办?在本研究中,我们说明了自适应视图可以发现局部特征,与输入的固定视图相比,局部特征与注意力更好地结合。我们用一个可变形的对应物(称为“变形器”)替换构象结构中的深度CNN。通过分析我们表现最好的模型,我们可视化了变形器学习的局部感受野和全局注意力图,并显示了在话语层面上增加的特征关联。通过对学习到的核偏移量的统计分析,可以深入了解特征信息随网络深度的变化。最后,仅替换编码器中的一半层,变形器在不使用LM的情况下提高了+5.6%的相对功率,在WSJ eval92集合上使用LM的情况下提高了+6.4%的相对功率。摘要:Convolutional neural networks (CNN) have improved speech recognition performance greatly by exploiting localized time-frequency patterns. But these patterns are assumed to appear in symmetric and rigid kernels by the conventional CNN operation. It motivates the question: What about asymmetric kernels? In this study, we illustrate adaptive views can discover local features which couple better with attention than fixed views of the input. We replace depthwise CNNs in the Conformer architecture with a deformable counterpart, dubbed this "Deformer". By analyzing our best-performing model, we visualize both local receptive fields and global attention maps learned by the Deformer and show increased feature associations on the utterance level. The statistical analysis of learned kernel offsets provides an insight into the change of information in features with the network depth. Finally, replacing only half of the layers in the encoder, the Deformer improves +5.6% relative WER without a LM and +6.4% relative WER with a LM over the Conformer baseline on the WSJ eval92 set.
【3】 Adversarial Multi-Task Deep Learning for Noise-Robust Voice Activity Detection with Low Algorithmic Delay
标题:基于对抗性多任务深度学习的低延迟抗噪语音检测
链接:https://arxiv.org/abs/2207.01691
作者:Claus Meyer Larsen,Peter Koch,Zheng-Hua Tan摘要:语音活动检测(VAD)是各种语音处理系统中一个重要的预处理步骤。VAD在实际应用中应该能够在噪声和无噪声环境中检测语音,同时不会引入显著的延迟。在这项工作中,我们提出了在训练监督VAD时使用对抗性多任务学习方法。该方法已应用于最先进的基于VAD波形的语音活动检测。此外,还研究了VADis在不同算法延迟下的性能,这是延迟的一个重要因素。观察到,在模型中引入对抗性多任务学习可以提高曲线下面积(AUC)的性能,尤其是在噪声环境中,而在较高的SNR水平下,性能不会降低。对抗式多任务学习仅适用于训练阶段,因此在测试中不会产生额外成本。此外,还研究了性能与算法延迟之间的相关性,观察到当算法延迟从398 ms降低到23 ms时,VAD性能的下降只是适度的。摘要:Voice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments, while not introducing significant latency. In this work we propose using an adversarial multi-task learning method when training a supervised VAD. The method has been applied to the state-of-the-art VAD Waveform-based Voice Activity Detection. Additionally the performance of the VADis investigated under different algorithmic delays, which is an important factor in latency. Introducing adversarial multi-task learning to the model is observed to increase performance in terms of Area Under Curve (AUC), particularly in noisy environments, while the performance is not degraded at higher SNR levels. The adversarial multi-task learning is only applied in the training phase and thus introduces no additional cost in testing. Furthermore the correlation between performance and algorithmic delays is investigated, and it is observed that the VAD performance degradation is only moderate when lowering the algorithmic delay from 398 ms to 23 ms.
【4】 A cross-corpus study on speech emotion recognition
标题:语音情感识别的跨语料库研究
链接:https://arxiv.org/abs/2207.02104
* 与cs.SD语音【1】为同一篇
作者:Rosanna Milner,Md Asif Jalal,Raymond W. M. Ng,Thomas Hain摘要:对于语音情感数据集,很难获得大量可靠的数据,与日常生活中表现较少的情感相比,行为情感可能过高。最近,创建了具有自然情感的更大数据集。本研究不是忽略较小的行为数据集,而是调查从行为情绪中学习的信息是否有助于检测自然情绪。跨语料库研究大多考虑跨语言甚至跨年龄的数据集,不同的情感注释方法会导致性能下降,这会带来困难。为了保持一致,我们考虑了四个成人英语数据集,包括行为情感、诱发情感和自然情感。提出了一种最先进的模型来准确地研究性能下降。该系统包括一个双向LSTM,该LSTM具有一个注意力机制,用于跨数据集对情绪进行分类。实验以跨语料库和多领域的方式研究了训练模型的效果,结果表明信息传递并不成功。域外模型,然后适应缺失的数据集,域对抗训练(DAT)被证明更适合推广到跨数据集的情绪。这表明了从行为数据集到那些具有更自然情感的数据集的积极信息传递,以及在不同语料库上训练的益处。摘要:For speech emotion datasets, it has been difficult to acquire large quantities of reliable data and acted emotions may be over the top compared to less expressive emotions displayed in everyday life. Lately, larger datasets with natural emotions have been created. Instead of ignoring smaller, acted datasets, this study investigates whether information learnt from acted emotions is useful for detecting natural emotions. Cross-corpus research has mostly considered cross-lingual and even cross-age datasets, and difficulties arise from different methods of annotating emotions causing a drop in performance. To be consistent, four adult English datasets covering acted, elicited and natural emotions are considered. A state-of-the-art model is proposed to accurately investigate the degradation of performance. The system involves a bi-directional LSTM with an attention mechanism to classify emotions across datasets. Experiments study the effects of training models in a cross-corpus and multi-domain fashion and results show the transfer of information is not successful. Out-of-domain models, followed by adapting to the missing dataset, and domain adversarial training (DAT) are shown to be more suitable to generalising to emotions across datasets. This shows positive information transfer from acted datasets to those with more natural emotions and the benefits from training on different corpora.
【5】 WeSinger 2: Fully Parallel Singing Voice Synthesis via Multi-Singer Conditional Adversarial Training
标题:WeSinger 2:基于多歌手条件对抗训练的完全并行歌声合成
链接:https://arxiv.org/abs/2207.01886
* 与cs.SD语音【2】为同一篇
作者:Zewang Zhang,Yibin Zheng,Xinhui Li,Li Lu备注:5 pages, 2 figures, 4 tables摘要:本文旨在介绍一种鲁棒的歌声合成(SVS)系统,利用对抗训练策略有效地生成高质量的歌声。一方面,我们设计了简单但通用的随机区域条件鉴别器来帮助监督声学模型,它可以有效地避免基于持续时间分配Transformer的声学模型的过度平滑频谱预测。另一方面,我们巧妙地将频谱图与帧级线性插值F0序列相结合,作为神经声码器的输入,然后借助波形域中的多个对抗性鉴别器和频域中的多尺度距离函数对神经声码器进行优化。实验结果和烧蚀研究表明,与我们以前的自回归工作相比,我们的新系统可以通过对不同的歌唱数据集进行微调,有效地产生高质量的歌唱声音,这些数据集覆盖几分钟到几小时。网上有一些合成的歌唱样本[https://zzw922cn.github.io/wesinger2 ].摘要:This paper aims to introduce a robust singing voice synthesis (SVS) system to produce high-quality singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random area conditional discriminators to help supervise the acoustic model, which can effectively avoid the over-smoothed spectrogram prediction by the duration-allocated Transformer-based acoustic model. On the other hand, we subtly combined the spectrogram with the frame-level linearly-interpolated F0 sequence as the input for the neural vocoder, which is then optimized with the help of multiple adversarial discriminators in the waveform domain and multi-scale distance functions in the frequency domain. The experimental results and ablation studies concluded that, compared with our previous auto-regressive work, our new system can produce high-quality singing voices efficiently by fine-tuning on different singing datasets covering several minutes to a few hours. Some synthesized singing samples are available online [https://zzw922cn.github.io/wesinger2 ].
【6】 Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice Conversion
标题:Glow-WaveGan 2:高质量零激发文本到语音合成和任意到任意语音转换
链接:https://arxiv.org/abs/2207.01832
* 与cs.SD语音【3】为同一篇
作者:Yi Lei,Shan Yang,Jian Cong,Lei Xie,Dan Su摘要:语音生成的零拍场景旨在仅用目标说话人的一句话合成一种新的看不见的语音。虽然在零炮场景中,在声学建模和声码器这两个阶段都存在适应新语音的挑战,但以前的工作通常只从一个阶段考虑这个问题。在本文中,我们将之前的Glow-WaveGAN扩展到Glow-WaveGAN 2,旨在从两个阶段解决高质量零炮文本到语音和任意到任意语音转换的问题。我们首先建立了一个通用的WaveGAN模型,用于提取语音的潜在分布$p(z)$并从中重建波形。然后,基于流的声学模型只需要从文本中学习相同的$p(z)$,这自然避免了声学模型和声码器之间的不匹配,从而在不进行模型微调的情况下生成高质量的语音。基于连续说话人空间和流的可逆性,可以获得任何说话人的条件分布,从而可以进一步为新说话人生成高质量的零拍语音。我们特别研究了两种构建说话人空间的方法,即预训练说话人编码器和联合训练说话人编码器。在LibriTTS语料库和VTCK语料库上进行的TTS和VC实验证明了Glow-WaveGAN-2的优越性。摘要:The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker. Although the challenges of adapting new voices in zero-shot scenario exist in both stages -- acoustic modeling and vocoder, previous works usually consider the problem from only one stage. In this paper, we extend our previous Glow-WaveGAN to Glow-WaveGAN 2, aiming to solve the problem from both stages for high-quality zero-shot text-to-speech and any-to-any voice conversion. We first build a universal WaveGAN model for extracting latent distribution $p(z)$ of speech and reconstructing waveform from it. Then a flow-based acoustic model only needs to learn the same $p(z)$ from texts, which naturally avoids the mismatch between the acoustic model and the vocoder, resulting in high-quality generated speech without model fine-tuning. Based on a continuous speaker space and the reversible property of flows, the conditional distribution can be obtained for any speaker, and thus we can further conduct high-quality zero-shot speech generation for new speakers. We particularly investigate two methods to construct the speaker space, namely pre-trained speaker encoder and jointly-trained speaker encoder. The superiority of Glow-WaveGAN 2 has been proved through TTS and VC experiments conducted on LibriTTS corpus and VTCK corpus.
【7】 Backend Ensemble for Speaker Verification and Spoofing Countermeasure
标题:用于说话人验证的后端集成及欺骗对策
链接:https://arxiv.org/abs/2207.01802
* 与cs.SD语音【4】为同一篇
作者:Li Zhang,Yue Li,Huan Zhao,Qing Wang,Lei Xie摘要:本文描述了提交给2022年防欺骗说话人验证挑战的NPU系统。我们从三个方面特别关注用于说话人验证和欺骗对策的\textit{backend-embly}。首先,除了简单的级联之外,我们还提出了循环矩阵变换和堆叠用于说话人嵌入和对策嵌入。通过对新定义的循环嵌入的叠加操作,我们几乎探索了说话人嵌入和对策嵌入之间的所有可能交互。其次,我们尝试使用不同的卷积神经网络,将嵌入的显著区域选择性地融合到具有卷积核的通道中。最后,我们设计了一维卷积神经网络中的并行注意,以学习通道维度中的全局相关性以及特征维度中的重要部分。同时,我们在2D卷积神经网络中嵌入挤压和激发注意,以学习说话人嵌入和对策嵌入之间的全局依赖性。实验结果表明,上述方法都是有效的。经过上述方法增强的四个训练有素的模型融合后,我们获得的最佳SASV-EER、SPF-EER和SV-EER在评估集上分别为0.559%、0.354%和0.857%。加上上述贡献,我们的提交系统在这项挑战中获得第五名。摘要:This paper describes the NPU system submitted to Spoofing Aware Speaker Verification Challenge 2022. We particularly focus on the \textit{backend ensemble} for speaker verification and spoofing countermeasure from three aspects. Firstly, besides simple concatenation, we propose circulant matrix transformation and stacking for speaker embeddings and countermeasure embeddings. With the stacking operation of newly-defined circulant embeddings, we almost explore all the possible interactions between speaker embeddings and countermeasure embeddings. Secondly, we attempt different convolution neural networks to selectively fuse the embeddings' salient regions into channels with convolution kernels. Finally, we design parallel attention in 1D convolution neural networks to learn the global correlation in channel dimensions as well as to learn the important parts in feature dimensions. Meanwhile, we embed squeeze-and-excitation attention in 2D convolutional neural networks to learn the global dependence among speaker embeddings and countermeasure embeddings. Experimental results demonstrate that all the above methods are effective. After fusion of four well-trained models enhanced by the mentioned methods, the best SASV-EER, SPF-EER and SV-EER we achieve are 0.559\%, 0.354\% and 0.857\% on the evaluation set respectively. Together with the above contributions, our submission system achieves the fifth place in this challenge.
【8】 BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model
标题:他能预测对比焦点吗?使用语言模型预测和控制神经TTS中的突出度
链接:https://arxiv.org/abs/2207.01718
作者:Brooke Stephenson,Laurent Besacier,Laurent Girin,Thomas Hueber摘要:最近的几项研究测试了transformer语言模型表示法在文本语音合成(TTS)中推断韵律特征的使用。虽然这些研究总体上探讨了韵律,但在这项工作中,我们特别关注对比焦点对人称代词的预测。这是一项特别具有挑战性的任务,因为它通常需要语义、话语和/或语用知识才能正确预测。我们收集了一个包含对比焦点的话语语料库,并评估了经过微调以预测量化声学突出特征的BERT模型在这些样本上的准确性。我们还研究了过去的话语如何为这种预测提供相关信息。此外,我们评估了基于声学突出特征的TTS模型中代词突出的可控性。摘要:Several recent studies have tested the use of transformer language model representations to infer prosodic features for text-to-speech synthesis (TTS). While these studies have explored prosody in general, in this work, we look specifically at the prediction of contrastive focus on personal pronouns. This is a particularly challenging task as it often requires semantic, discursive and/or pragmatic knowledge to predict correctly. We collect a corpus of utterances containing contrastive focus and we evaluate the accuracy of a BERT model, finetuned to predict quantized acoustic prominence features, on these samples. We also investigate how past utterances can provide relevant information for this prediction. Furthermore, we evaluate the controllability of pronoun prominence in a TTS model conditioned on acoustic prominence features.
【9】 An adaptive music generation architecture for games based on the deep learning Transformer mode
标题:一种基于深度学习变换模式的自适应游戏音乐生成架构
链接:https://arxiv.org/abs/2207.01698
* 与cs.SD语音【5】为同一篇
作者:Gustavo Amaral Costa dos Santos,Augusto Baffa,Jean-Pierre Briot,Bruno Feijó,Antonio Luz Furtado摘要:本文提出了一种基于Transformer深度学习模型的视频游戏音乐生成架构。该系统按照目前设计视频游戏音乐的作曲家使用的标准分层策略,生成不同层次的音乐。根据唤醒价模型,音乐适应演奏者的心理环境。我们的动机是根据玩家的口味定制音乐,玩家可以通过一组训练音乐示例来选择自己喜欢的音乐风格。我们讨论了当前的局限性和未来的前景,例如音乐组件的协作和交互控制。摘要:This paper presents an architecture for generating music for video games based on the Transformer deep learning model. The system generates music in various layers, following the standard layering strategy currently used by composers designing video game music. The music is adaptive to the psychological context of the player, according to the arousal-valence model. Our motivation is to customize music according to the player's tastes, who can select his preferred style of music through a set of training examples of music. We discuss current limitations and prospects for the future, such as collaborative and interactive control of the musical components.
【10】 Stochastic Restoration of Heavily Compressed Musical Audio using Generative Adversarial Networks
标题:基于产生式对抗网络的重压缩音乐音频随机恢复
链接:https://arxiv.org/abs/2207.01667
* 与cs.SD语音【6】为同一篇
作者:Stefan Lattner,Javier Nistal摘要:有损音频编解码器通过删除人类感知中往往听不见的信息来压缩(和解压缩)数字音频流。在高压缩率下,此类编解码器可能会在音频信号中引入各种损伤。许多作品都使用深度学习技术解决了音频增强和压缩伪影消除的问题。然而,只有少数作品解决了音乐领域中严重压缩音频信号的恢复问题。在这种情况下,没有唯一的解决方案来恢复原始信号。因此,在本研究中,我们测试了生成对抗网络(GAN)架构的随机生成器。这种随机发生器以高度压缩的音乐音频信号为条件,有朝一日可能会产生与高质量版本无法区分的输出。因此,本研究可能有助于深入了解更高效的音乐数据存储和传输。我们在16、32和64kbit/s的MP3压缩音频信号上训练随机和确定性生成器。我们利用客观指标和听力测试对不同的实验进行了广泛的评估。我们发现,与MP3版本相比,该模型可以提高16和32kbit/s音频信号的质量,并且随机生成器能够生成比确定性生成器更接近原始信号的输出。摘要:Lossy audio codecs compress (and decompress) digital audio streams by removing information that tends to be inaudible in human perception. Under high compression rates, such codecs may introduce a variety of impairments in the audio signal. Many works have tackled the problem of audio enhancement and compression artifact removal using deep learning techniques. However, only a few works tackle the restoration of heavily compressed audio signals in the musical domain. In such a scenario, there is no unique solution for the restoration of the original signal. Therefore, in this study, we test a stochastic generator of a Generative Adversarial Network (GAN) architecture for this task. Such a stochastic generator, conditioned on highly compressed musical audio signals, could one day generate outputs indistinguishable from high-quality releases. Therefore, the present study may yield insights into more efficient musical data storage and transmission. We train stochastic and deterministic generators on MP3-compressed audio signals with 16, 32, and 64 kbit/s. We perform an extensive evaluation of the different experiments utilizing objective metrics and listening tests. We find that the models can improve the quality of the audio signals over the MP3 versions for 16 and 32 kbit/s and that the stochastic generators are capable of generating outputs that are closer to the original signals than those of the deterministic generators.机器翻译,仅供参考