今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理4篇。本文经arXiv每日学术速递授权转载
【1】 BigVGAN: A Universal Neural Vocoder with Large-Scale Training标题:BigVGAN:一种大规模训练的通用神经声码器作者:Sang-gil Lee,Wei Ping,Boris Ginsburg,Bryan Catanzaro,Sungroh Yoon机构:Data Science & AI Lab., Seoul National University (SNU), NVIDIA, AIIS, ASRI, INMC, ISRC, NSI, and Interdisciplinary Program in AI, SNU备注:Listen to audio samples from BigVGAN at: this https URL摘要:尽管基于生成对抗网络(GAN)的声码器最近取得了进展,该模型根据mel频谱图生成原始波形,但为不同录音环境中的众多扬声器合成高保真音频仍然具有挑战性。在这项工作中,我们提出了BigVGAN,这是一种通用的声码器,在零炮设置的各种不可见条件下都能很好地推广。我们在发生器中引入了周期非线性和抗混叠表示,这为波形合成带来了所需的电感偏置,并显著改善了音频质量。基于我们改进的发生器和最先进的鉴别器,我们以112M参数的最大规模训练GAN声码器,这在文献中是前所未有的。特别是,我们确定并解决了特定于这种规模的训练不稳定性,同时保持高保真输出,而不过度正则化。我们的BigVGAN在各种发行外场景中实现了最先进的零拍性能,包括在看不见(甚至嘈杂)的录制环境中使用新的扬声器、新颖的语言、唱歌的声音、音乐和器乐音频。我们将在以下位置发布代码和模型:https://github.com/NVIDIA/BigVGAN摘要:Despite recent progress in generative adversarial network(GAN)-based vocoders, where the model generates raw waveform conditioned on mel spectrogram, it is still challenging to synthesize high-fidelity audio for numerous speakers across varied recording environments. In this work, we present BigVGAN, a universal vocoder that generalizes well under various unseen conditions in zero-shot setting. We introduce periodic nonlinearities and anti-aliased representation into the generator, which brings the desired inductive bias for waveform synthesis and significantly improves audio quality. Based on our improved generator and the state-of-the-art discriminators, we train our GAN vocoder at the largest scale up to 112M parameters, which is unprecedented in the literature. In particular, we identify and address the training instabilities specific to such scale, while maintaining high-fidelity output without over-regularization. Our BigVGAN achieves the state-of-the-art zero-shot performance for various out-of-distribution scenarios, including new speakers, novel languages, singing voices, music and instrumental audio in unseen (even noisy) recording environments. We will release our code and model at: https://github.com/NVIDIA/BigVGAN
【2】 Revisiting End-to-End Speech-to-Text Translation From Scratch
标题:从零开始重新审视端到端的语音到文本翻译
链接:https://arxiv.org/abs/2206.04571
作者:Biao Zhang,Barry Haddow,Rico Sennrich机构:IntroductionEnd-to-end (E 2E) speech-to-text translation (ST) is the taskof translating a source-language audio directly to a foreign 1School of Informatics, University of Edinburgh 2Departmentof Computational Linguistics, University of Zurich摘要:端到端(E2E)语音到文本翻译(ST)通常依赖于通过语音识别或文本翻译任务使用源转录本对其编码器和/或解码器进行预训练,否则翻译性能会大幅下降。然而,成绩单并不总是可用的,文献中很少研究这种预训练对E2E ST的重要性。在本文中,我们重新探讨了这个问题,并探讨了仅在语音翻译对上训练的E2E ST的质量可以在多大程度上得到提高。我们重新检查了之前证明对ST有益的几种技术,并提供了一套最佳实践,使基于Transformer的E2E ST系统从零开始训练。此外,我们还提出了参数化距离惩罚,以便于在语音自我注意模型中进行局部建模。在涵盖23种语言的四个基准上,我们的实验表明,在不使用任何成绩单或预训练的情况下,拟议的系统达到甚至优于以前采用预训练的研究,尽管差距仍然存在于(极)低的资源环境中。最后,我们讨论了神经声学特征建模,其中设计了一个神经模型来直接从原始语音信号中提取声学特征,目的是简化归纳偏差,并在描述语音时为模型增加自由度。我们首次证明了它的可行性,并在ST任务上显示了令人鼓舞的结果。摘要:End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially. However, transcripts are not always available, and how significant such pretraining is for E2E ST has rarely been studied in the literature. In this paper, we revisit this question and explore the extent to which the quality of E2E ST trained on speech-translation pairs alone can be improved. We reexamine several techniques proven beneficial to ST previously, and offer a set of best practices that biases a Transformer-based E2E ST system toward training from scratch. Besides, we propose parameterized distance penalty to facilitate the modeling of locality in the self-attention model for speech. On four benchmarks covering 23 languages, our experiments show that, without using any transcripts or pretraining, the proposed system reaches and even outperforms previous studies adopting pretraining, although the gap remains in (extremely) low-resource settings. Finally, we discuss neural acoustic feature modeling, where a neural model is designed to extract acoustic features from raw speech signals directly, with the goal to simplify inductive biases and add freedom to the model in describing speech. For the first time, we demonstrate its feasibility and show encouraging results on ST tasks.
【3】 Face-Dubbing++: Lip-Synchronous, Voice Preserving Translation of Videos
标题:人脸配音++:视频的唇形同步、保声翻译
链接:https://arxiv.org/abs/2206.04523
作者:Alexander Waibel,Moritz Behr,Fevziye Irem Eyiokur,Dogucan Yaman,Tuan-Nam Nguyen,Carlos Mullov,Mehmet Arif Demirtas,Alperen Kantarcı,Stefan Constantin,Hazım Kemal Ekenel机构:Karlsruhe Institute of Technology,Carnegie Mellon University,Istanbul Technical University摘要:在本文中,我们提出了一种神经端到端系统,用于视频的语音保持、唇部同步翻译。该系统设计为结合多个组件模型,生成以目标语言说话的原始说话人的视频,该视频与目标语音保持唇部同步,同时重点关注原始说话人的语音、语音特征和面部视频。管道从自动语音识别开始,包括重点检测,然后是翻译模型。然后,通过一个文本到语音模型合成翻译后的文本,该模型重新创建从原始句子映射的原始重点。然后,使用语音转换模型将生成的合成语音映射回原始说话人的语音。最后,为了使说话人的嘴唇与翻译的音频同步,基于条件生成对抗网络的模型生成关于输入人脸图像以及语音转换模型的输出的自适应嘴唇运动帧。最后,系统将生成的视频与转换后的音频相结合,生成最终输出。结果是一段视频,一个演讲者用另一种语言说话,但实际上并不知道。为了评估我们的设计,我们提供了完整系统的用户研究以及单个组件的单独评估。由于没有可用的数据集来评估整个系统,我们收集了一个测试集,并在此测试集上评估我们的系统。结果表明,我们的系统能够在保留原说话人特征的同时,生成说话人讲目标语言的令人信服的视频。将共享收集的数据集。摘要:In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos. The system is designed to combine multiple component models and produces a video of the original speaker speaking in the target language that is lip-synchronous with the target speech, yet maintains emphases in speech, voice characteristics, face video of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by a translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases mapped from the original sentence. The resulting synthetic voice is then mapped back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a conditional generative adversarial network-based model generates frames of adapted lip movements with respect to the input face image as well as the output of the voice conversion model. In the end, the system combines the generated video with the converted audio to produce the final output. The result is a video of a speaker speaking in another language without actually knowing it. To evaluate our design, we present a user study of the complete system as well as separate evaluations of the single components. Since there is no available dataset to evaluate our whole system, we collect a test set and evaluate our system on this test set. The results indicate that our system is able to generate convincing videos of the original speaker speaking the target language while preserving the original speaker's characteristics. The collected dataset will be shared.
【4】 Context-based out-of-vocabulary word recovery for ASR systems in Indian languages
标题:印度语言ASR系统中基于上下文的词汇外词汇恢复
链接:https://arxiv.org/abs/2206.04305
作者:Arun Baby,Saranya Vinnaitherthan,Akhil Kerhalkar,Pranav Jawale,Sharath Adavanne,Nagaraj Adiga机构:Zapr Media Labs, India摘要:对于自动语音识别(ASR)系统来说,词汇外(OOV)词的检测和恢复一直是一个挑战。许多现有的方法侧重于通过修改声学和语言模型以及将上下文词巧妙地集成到模型中来建模OOV词。为了训练如此复杂的模型,我们需要大量的数据,包括上下文、额外的训练时间和增加的模型大小。然而,在获得ASR转录来恢复基于上下文的OOV单词后,对其后处理方法的研究还不多。在这项工作中,我们提出了一种后处理技术来提高基于上下文的OOV恢复的性能。我们创建了一个声学增强的语言模型,在电话层面上用OOV单词列表制作了一个子图。我们提出了两种方法来确定合适的代价函数,以根据上下文检索OOV单词。代价函数是基于语音和声学知识定义的,用于匹配和恢复解码中正确的上下文单词。从单词和句子两个层面对所提出的成本函数的有效性进行了评估。评估结果表明,该方法可以在多个类别中平均恢复50%的基于上下文的OOV单词。摘要:Detecting and recovering out-of-vocabulary (OOV) words is always challenging for Automatic Speech Recognition (ASR) systems. Many existing methods focus on modeling OOV words by modifying acoustic and language models and integrating context words cleverly into models. To train such complex models, we need a large amount of data with context words, additional training time, and increased model size. However, after getting the ASR transcription to recover context-based OOV words, the post-processing method has not been explored much. In this work, we propose a post-processing technique to improve the performance of context-based OOV recovery. We created an acoustically boosted language model with a sub-graph made at phone level with an OOV words list. We proposed two methods to determine a suitable cost function to retrieve the OOV words based on the context. The cost function is defined based on phonetic and acoustic knowledge for matching and recovering the correct context words in the decode. The effectiveness of the proposed cost function is evaluated at both word-level and sentence-level. The evaluation results show that this approach can recover an average of 50% context-based OOV words across multiple categories.
【1】 Context-based out-of-vocabulary word recovery for ASR systems in Indian languages
标题:印度语言ASR系统中基于上下文的词汇外词汇恢复
链接:https://arxiv.org/abs/2206.04305
作者:Arun Baby,Saranya Vinnaitherthan,Akhil Kerhalkar,Pranav Jawale,Sharath Adavanne,Nagaraj Adiga机构:Zapr Media Labs, India摘要:对于自动语音识别(ASR)系统来说,词汇外(OOV)词的检测和恢复一直是一个挑战。许多现有的方法侧重于通过修改声学和语言模型以及将上下文词巧妙地集成到模型中来建模OOV词。为了训练如此复杂的模型,我们需要大量的数据,包括上下文、额外的训练时间和增加的模型大小。然而,在获得ASR转录来恢复基于上下文的OOV单词后,对其后处理方法的研究还不多。在这项工作中,我们提出了一种后处理技术来提高基于上下文的OOV恢复的性能。我们创建了一个声学增强的语言模型,在电话层面上用OOV单词列表制作了一个子图。我们提出了两种方法来确定合适的代价函数,以根据上下文检索OOV单词。代价函数是基于语音和声学知识定义的,用于匹配和恢复解码中正确的上下文单词。从单词和句子两个层面对所提出的成本函数的有效性进行了评估。评估结果表明,该方法可以在多个类别中平均恢复50%的基于上下文的OOV单词。摘要:Detecting and recovering out-of-vocabulary (OOV) words is always challenging for Automatic Speech Recognition (ASR) systems. Many existing methods focus on modeling OOV words by modifying acoustic and language models and integrating context words cleverly into models. To train such complex models, we need a large amount of data with context words, additional training time, and increased model size. However, after getting the ASR transcription to recover context-based OOV words, the post-processing method has not been explored much. In this work, we propose a post-processing technique to improve the performance of context-based OOV recovery. We created an acoustically boosted language model with a sub-graph made at phone level with an OOV words list. We proposed two methods to determine a suitable cost function to retrieve the OOV words based on the context. The cost function is defined based on phonetic and acoustic knowledge for matching and recovering the correct context words in the decode. The effectiveness of the proposed cost function is evaluated at both word-level and sentence-level. The evaluation results show that this approach can recover an average of 50% context-based OOV words across multiple categories.
【2】 BigVGAN: A Universal Neural Vocoder with Large-Scale Training
标题:BigVGAN:一种大规模训练的通用神经声码器
链接:https://arxiv.org/abs/2206.04658
作者:Sang-gil Lee,Wei Ping,Boris Ginsburg,Bryan Catanzaro,Sungroh Yoon机构:Data Science & AI Lab., Seoul National University (SNU), NVIDIA, AIIS, ASRI, INMC, ISRC, NSI, and Interdisciplinary Program in AI, SNU备注:Listen to audio samples from BigVGAN at: this https URL摘要:尽管基于生成对抗网络(GAN)的声码器最近取得了进展,该模型根据mel频谱图生成原始波形,但为不同录音环境中的众多扬声器合成高保真音频仍然具有挑战性。在这项工作中,我们提出了BigVGAN,这是一种通用的声码器,在零炮设置的各种不可见条件下都能很好地推广。我们在发生器中引入了周期非线性和抗混叠表示,这为波形合成带来了所需的电感偏置,并显著改善了音频质量。基于我们改进的发生器和最先进的鉴别器,我们以112M参数的最大规模训练GAN声码器,这在文献中是前所未有的。特别是,我们确定并解决了特定于这种规模的训练不稳定性,同时保持高保真输出,而不过度正则化。我们的BigVGAN在各种发行外场景中实现了最先进的零拍性能,包括在看不见(甚至嘈杂)的录制环境中使用新的扬声器、新颖的语言、唱歌的声音、音乐和器乐音频。我们将在以下位置发布代码和模型:https://github.com/NVIDIA/BigVGAN摘要:Despite recent progress in generative adversarial network(GAN)-based vocoders, where the model generates raw waveform conditioned on mel spectrogram, it is still challenging to synthesize high-fidelity audio for numerous speakers across varied recording environments. In this work, we present BigVGAN, a universal vocoder that generalizes well under various unseen conditions in zero-shot setting. We introduce periodic nonlinearities and anti-aliased representation into the generator, which brings the desired inductive bias for waveform synthesis and significantly improves audio quality. Based on our improved generator and the state-of-the-art discriminators, we train our GAN vocoder at the largest scale up to 112M parameters, which is unprecedented in the literature. In particular, we identify and address the training instabilities specific to such scale, while maintaining high-fidelity output without over-regularization. Our BigVGAN achieves the state-of-the-art zero-shot performance for various out-of-distribution scenarios, including new speakers, novel languages, singing voices, music and instrumental audio in unseen (even noisy) recording environments. We will release our code and model at: https://github.com/NVIDIA/BigVGAN
【3】 Revisiting End-to-End Speech-to-Text Translation From Scratch
标题:从零开始重新审视端到端的语音到文本翻译
链接:https://arxiv.org/abs/2206.04571
作者:Biao Zhang,Barry Haddow,Rico Sennrich机构:IntroductionEnd-to-end (E 2E) speech-to-text translation (ST) is the taskof translating a source-language audio directly to a foreign 1School of Informatics, University of Edinburgh 2Departmentof Computational Linguistics, University of Zurich摘要:端到端(E2E)语音到文本翻译(ST)通常依赖于通过语音识别或文本翻译任务使用源转录本对其编码器和/或解码器进行预训练,否则翻译性能会大幅下降。然而,成绩单并不总是可用的,文献中很少研究这种预训练对E2E ST的重要性。在本文中,我们重新探讨了这个问题,并探讨了仅在语音翻译对上训练的E2E ST的质量可以在多大程度上得到提高。我们重新检查了之前证明对ST有益的几种技术,并提供了一套最佳实践,使基于Transformer的E2E ST系统从零开始训练。此外,我们还提出了参数化距离惩罚,以便于在语音自我注意模型中进行局部建模。在涵盖23种语言的四个基准上,我们的实验表明,在不使用任何成绩单或预训练的情况下,拟议的系统达到甚至优于以前采用预训练的研究,尽管差距仍然存在于(极)低的资源环境中。最后,我们讨论了神经声学特征建模,其中设计了一个神经模型来直接从原始语音信号中提取声学特征,目的是简化归纳偏差,并在描述语音时为模型增加自由度。我们首次证明了它的可行性,并在ST任务上显示了令人鼓舞的结果。摘要:End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially. However, transcripts are not always available, and how significant such pretraining is for E2E ST has rarely been studied in the literature. In this paper, we revisit this question and explore the extent to which the quality of E2E ST trained on speech-translation pairs alone can be improved. We reexamine several techniques proven beneficial to ST previously, and offer a set of best practices that biases a Transformer-based E2E ST system toward training from scratch. Besides, we propose parameterized distance penalty to facilitate the modeling of locality in the self-attention model for speech. On four benchmarks covering 23 languages, our experiments show that, without using any transcripts or pretraining, the proposed system reaches and even outperforms previous studies adopting pretraining, although the gap remains in (extremely) low-resource settings. Finally, we discuss neural acoustic feature modeling, where a neural model is designed to extract acoustic features from raw speech signals directly, with the goal to simplify inductive biases and add freedom to the model in describing speech. For the first time, we demonstrate its feasibility and show encouraging results on ST tasks.
【4】 Face-Dubbing++: Lip-Synchronous, Voice Preserving Translation of Videos
标题:人脸配音++:视频的唇形同步、保声翻译
链接:https://arxiv.org/abs/2206.04523
作者:Alexander Waibel,Moritz Behr,Fevziye Irem Eyiokur,Dogucan Yaman,Tuan-Nam Nguyen,Carlos Mullov,Mehmet Arif Demirtas,Alperen Kantarcı,Stefan Constantin,Hazım Kemal Ekenel机构:Karlsruhe Institute of Technology,Carnegie Mellon University,Istanbul Technical University摘要:在本文中,我们提出了一种神经端到端系统,用于视频的语音保持、唇部同步翻译。该系统设计为结合多个组件模型,生成以目标语言说话的原始说话人的视频,该视频与目标语音保持唇部同步,同时重点关注原始说话人的语音、语音特征和面部视频。管道从自动语音识别开始,包括重点检测,然后是翻译模型。然后,通过一个文本到语音模型合成翻译后的文本,该模型重新创建从原始句子映射的原始重点。然后,使用语音转换模型将生成的合成语音映射回原始说话人的语音。最后,为了使说话人的嘴唇与翻译的音频同步,基于条件生成对抗网络的模型生成关于输入人脸图像以及语音转换模型的输出的自适应嘴唇运动帧。最后,系统将生成的视频与转换后的音频相结合,生成最终输出。结果是一段视频,一个演讲者用另一种语言说话,但实际上并不知道。为了评估我们的设计,我们提供了完整系统的用户研究以及单个组件的单独评估。由于没有可用的数据集来评估整个系统,我们收集了一个测试集,并在此测试集上评估我们的系统。结果表明,我们的系统能够在保留原说话人特征的同时,生成说话人讲目标语言的令人信服的视频。将共享收集的数据集。摘要:In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos. The system is designed to combine multiple component models and produces a video of the original speaker speaking in the target language that is lip-synchronous with the target speech, yet maintains emphases in speech, voice characteristics, face video of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by a translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases mapped from the original sentence. The resulting synthetic voice is then mapped back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a conditional generative adversarial network-based model generates frames of adapted lip movements with respect to the input face image as well as the output of the voice conversion model. In the end, the system combines the generated video with the converted audio to produce the final output. The result is a video of a speaker speaking in another language without actually knowing it. To evaluate our design, we present a user study of the complete system as well as separate evaluations of the single components. Since there is no available dataset to evaluate our whole system, we collect a test set and evaluate our system on this test set. The results indicate that our system is able to generate convincing videos of the original speaker speaking the target language while preserving the original speaker's characteristics. The collected dataset will be shared.
机器翻译,仅供参考