声明:工作以来主要从事TTS工作,平时看些文章做些笔记。文章中难免存在错误的地方,还望大家海涵。平时搜集一些资料,方便查阅学习:

TTS 论文列表 http://yqli.tech/page/tts_paper.html TTS

开源数据 http://yqli.tech/page/data.html

如转载,请标明出处。


本文整理一下GAN在语音合成声学模型上的研究,统计可能不全。GAN在语音合成中的应用更多是在声码器部分,而在声学模型部分的研究很少。使用GAN的目的更多解决过平滑问题,从而使韵律更好,情感表达更丰富。本文搜集近几年的文章,做简短的总结:


基于tacotron

1)TFGAN: A Lightweight Library for Generative Adversarial Networks  (2017)

https://ai.googleblog.com/2017/12/tfgan-lightweight-library-for.html


2) A new gan-based end-to-end tts training algorithm (2019)

https://arxiv.org/pdf/1904.04775.pdf


3) A New End-to-End Long-Time Speech Synthesis System Based on Tacotron2 (2019)

https://sci-hub.se/https://doi.org/10.1145/3364908.3365292


非taoctron实验

4) Statistical parametric speech synthesis incorporating generative adversarial networks (2018)

https://sci-hub.se/https://doi.org/10.1109/TASLP.2017.2761547


5) Wasserstein GAN and waveform loss-based acoustic model training for multi-speaker text-to-speech synthesis systems using a WaveNet vocoder (2018)

https://arxiv.org/pdf/1807.11679.pdf


结合图片处理pix2pixHD

6) Reducing over-smoothness in speech synthesis using Generative Adversarial Networks (2018)

https://arxiv.org/pdf/1810.10989.pdf


7)  High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram(2019)

https://arxiv.org/pdf/1912.01167.pdf


歌唱合成SVS

8) WGANSing: A Multi-Voice Singing Voice Synthesizer Based on the Wasserstein-GAN (2019)

https://arxiv.org/pdf/1903.10729.pdf


9) HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis (2020)

https://arxiv.org/pdf/2009.01776.pdf


其它 (该文章没找到下载方式,付费不起,只能先放着)

10)  GAN acoustic model for Kazakh speech synthesis(2021)


一、 基于tacotron

基于tacotron的实验主要在tacotron系统上进行的优化。


1 TFGAN: A Lightweight Library for Generative Adversarial Networks  (2017)

本部分是google的blog中对tacotron 1进行的实验,解决tacotron1产生的声谱过平滑问题。


2 A new gan-based end-to-end tts training algorithm (2019)
本文在tacotorn2上添加discriminator,主要解决teacher forcing和free running两种方式产生feature之间的gap,从而提高合成的音频质量。



3 A New End-to-End Long-Time Speech Synthesis System Based on Tacotron2 (2019)

添加了三个分辨器:prosody, sound quality和alignments discriminator。tacotron2 删除mse损失函数而使用ganloss。



二、 非tacotron实验
基于DNN网络设计生成器和分辨器来生成声学特征。
4 Statistical parametric speech synthesis incorporating generative adversarial networks (2018)


5 Wasserstein GAN and waveform loss-based acoustic model training for multi-speaker text-to-speech synthesis systems using a WaveNet vocoder (2018)



三 、结合图片处理pix2pixHD
本部分的研究把mel特征当成图片处理,从而解决过平滑的问题(图片模糊)。主要先通过声学模型生成声学特征,然后使用图片处理pix2pixHD网络进行处理,从而使图片更清晰,则代表韵律更丰富。

6 Reducing over-smoothness in speech synthesis using Generative Adversarial Networks (2018)

提出使用gan的话可以解决过平滑问题,韵律更好
把tacotron2生成的mel当成图片进行pix2pixhd处理


7 High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram(2019)



四 歌唱合成SVS
该部分主要使用GAN来进行SVS,其中第九篇使用sub-frequency gan可进行借鉴
8 WGANSing: A Multi-Voice Singing Voice Synthesizer Based on the Wasserstein-GAN (2019)


9 HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis (2020)


五 其它
10GAN acoustic model for Kazakh speech synthesis (2021)

今年的文章,没找到免费的pdf,没钱下载,读者有渠道下载后可以分享我一份。