本文经arXiv每日学术速递授权转载
【1】 The evolution of inharmonicity and noisiness in contemporary popular music
标题: 当代流行音乐不和谐与噪音的演变
作者:Emmanuel Deruty,David Meredith,Stefan Lattner
备注:43 pages, 23 figures
链接:点击下载PDF文件
【2】 Hearing Your Blood Sugar: Non-Invasive Glucose Measurement Through Simple Vocal Signals, Transforming any Speech into a Sensor with Machine Learning
标题: 聆听血糖:通过简单的声音信号进行无创血糖测量,通过机器学习将任何语音转化为传感器
作者:Nihat Ahmadli,Mehmet Ali Sarsil,Onur Ergen
备注:5 figure and 5 tables. This manuscript is a pre-print to be submitted to a journal orand a conference. arXiv admin note: substantial text overlap with arXiv:2402.13812
链接:点击下载PDF文件
【3】 Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
标题: 通过对抗流匹配优化加速高保真波生成
作者:Sang-Hoon Lee,Ha-Yeong Choi,Seong-Whan Lee
备注:9 pages, 9 tables, 1 figure,
链接:点击下载PDF文件
【4】 Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
标题: 通过稀有和歧义词的上下文化增强基于大语言模型的语音识别
作者:Kento Nozawa,Takashi Masuko,Toru Taniguchi
备注:13 pages, 1 figure, and 7 tables
链接:点击下载PDF文件
标题: 通过稀有和歧义词的上下文化增强基于大语言模型的语音识别
作者:Kento Nozawa,Takashi Masuko,Toru Taniguchi
备注:13 pages, 1 figure, and 7 tables
链接:点击下载PDF文件
【2】 Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
标题: 推进对比音频预训练的多粒度对齐
作者:Yiming Li,Zhifang Guo,Xiangdong Wang,Hong Liu
备注:ACM MM 2024 (Oral)
链接:点击下载PDF文件
【3】 The evolution of inharmonicity and noisiness in contemporary popular music
标题: 当代流行音乐不和谐与噪音的演变
作者:Emmanuel Deruty,David Meredith,Stefan Lattner
备注:43 pages, 23 figures
链接:点击下载PDF文件
【4】 Hearing Your Blood Sugar: Non-Invasive Glucose Measurement Through Simple Vocal Signals, Transforming any Speech into a Sensor with Machine Learning
标题: 聆听血糖:通过简单的声音信号进行无创血糖测量,通过机器学习将任何语音转化为传感器
作者:Nihat Ahmadli,Mehmet Ali Sarsil,Onur Ergen
备注:5 figure and 5 tables. This manuscript is a pre-print to be submitted to a journal orand a conference. arXiv admin note: substantial text overlap with arXiv:2402.13812
链接:点击下载PDF文件
【5】 Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
标题: 通过对抗流匹配优化加速高保真波生成
作者:Sang-Hoon Lee,Ha-Yeong Choi,Seong-Whan Lee
备注:9 pages, 9 tables, 1 figure,
链接:点击下载PDF文件
标题: 当代流行音乐不和谐与噪音的演变
作者:Emmanuel Deruty,David Meredith,Stefan Lattner
备注:43 pages, 23 figures
链接:点击下载PDF文件
摘要:许多西方古典音乐使用基于声学共振的乐器。这种乐器产生谐波或准谐波的声音。另一方面,自20世纪70年代初以来,流行音乐主要是在录音室制作的。因此,流行音乐并不一定是以和声或准和声为基础的。在这项研究中,我们使用修改后的MPEG-7功能,探索和验证自1961年以来流行音乐中噪音和不和谐的使用方式。我们将这种演变置于其他广泛的音乐类别的背景下,包括西方古典钢琴音乐,西方古典管弦乐和音乐协奏曲。我们提出了新的功能,使我们能够区分噪声和不和谐造成的相对离散的偏音之间的相互作用之间的不和谐。当我们从这些新特征的角度来审视当代流行音乐的历史时,我们发现自1961年以来的这段时期可以分为三个阶段。从1961年到1972年,不和谐性稳步增加,但噪音没有显著增加。从1972年到1986年,非谐性和噪声都有所增加。然后,自1986年以来,今天的流行音乐的不和谐性和噪音都在稳步下降,与60年代的音乐相比,噪音明显减少,但更不和谐。我们将这些观察到的趋势与这一时期音乐制作实践的发展联系起来,并通过对某些关键艺术家和曲目的重点分析来说明这些趋势。摘要:Much of Western classical music uses instruments based on acoustic resonance. Such instruments produce harmonic or quasi-harmonic sounds. On the other hand, since the early 1970s, popular music has largely been produced in the recording studio. As a result, popular music is not bound to be based on harmonic or quasi-harmonic sounds. In this study, we use modified MPEG-7 features to explore and characterise the way in which the use of noise and inharmonicity has evolved in popular music since 1961. We set this evolution in the context of other broad categories of music, including Western classical piano music, Western classical orchestral music, and musique concr ete. We propose new features that allow us to distinguish between inharmonicity resulting from noise and inharmonicity resulting from interactions between relatively discrete partials. When the history of contemporary popular music is viewed through the lens of these new features, we find that the period since 1961 can be divided into three phases. From 1961 to 1972, there was a steady increase in inharmonicity but no significant increase in noise. From 1972 to 1986, both inharmonicity and noise increased. Then, since 1986, there has been a steady decrease in both inharmonicity and noise to today's popular music which is significantly less noisy but more inharmonic than the music of the sixties. We relate these observed trends to the development of music production practice over the period and illustrate them with focused analyses of certain key artists and tracks.
【2】 Hearing Your Blood Sugar: Non-Invasive Glucose Measurement Through Simple Vocal Signals, Transforming any Speech into a Sensor with Machine Learning
标题: 聆听血糖:通过简单的声音信号进行无创血糖测量,通过机器学习将任何语音转化为传感器
作者:Nihat Ahmadli,Mehmet Ali Sarsil,Onur Ergen
备注:5 figure and 5 tables. This manuscript is a pre-print to be submitted to a journal orand a conference. arXiv admin note: substantial text overlap with arXiv:2402.13812
链接:点击下载PDF文件
摘要:有效的糖尿病管理在很大程度上依赖于血糖水平的持续监测,传统上通过侵入性和不舒服的方法来实现。虽然已经探索了各种非侵入性技术,例如光学、微波和电化学方法,但由于与复杂性、准确性和成本相关的问题,没有一种技术能够有效地取代这些侵入性技术。在这项研究中,我们提出了一种变革性的和直接的方法,利用语音分析来预测血糖水平。我们的研究调查了血糖波动和发声特征之间的关系,强调了血管动力学在发声过程中的影响。通过应用先进的机器学习算法,我们分析了声音信号的变化,并建立了与血糖水平的显着相关性。我们使用人工智能开发了一个预测模型,基于参与者的语音记录和相应的葡萄糖测量,利用逻辑回归和岭正则化。我们的研究结果表明,语音分析可以作为一个可行的非侵入性的替代血糖监测。这种创新的方法不仅有可能简化和降低与糖尿病管理相关的成本,而且旨在通过提供一种无痛和用户友好的血糖水平监测方法来提高糖尿病患者的生活质量。摘要:Effective diabetes management relies heavily on the continuous monitoring of blood glucose levels, traditionally achieved through invasive and uncomfortable methods. While various non-invasive techniques have been explored, such as optical, microwave, and electrochemical approaches, none have effectively supplanted these invasive technologies due to issues related to complexity, accuracy, and cost. In this study, we present a transformative and straightforward method that utilizes voice analysis to predict blood glucose levels. Our research investigates the relationship between fluctuations in blood glucose and vocal characteristics, highlighting the influence of blood vessel dynamics during voice production. By applying advanced machine learning algorithms, we analyzed vocal signal variations and established a significant correlation with blood glucose levels. We developed a predictive model using artificial intelligence, based on voice recordings and corresponding glucose measurements from participants, utilizing logistic regression and Ridge regularization. Our findings indicate that voice analysis may serve as a viable non-invasive alternative for glucose monitoring. This innovative approach not only has the potential to streamline and reduce the costs associated with diabetes management but also aims to enhance the quality of life for individuals living with diabetes by providing a painless and user-friendly method for monitoring blood sugar levels.
【3】 Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
标题: 通过对抗流匹配优化加速高保真波生成
作者:Sang-Hoon Lee,Ha-Yeong Choi,Seong-Whan Lee
备注:9 pages, 9 tables, 1 figure,
链接:点击下载PDF文件
摘要:本文介绍了PeriodWave-Turbo,一个高保真,高效的波形生成模型,通过对抗流匹配优化。最近,条件流匹配(CFM)生成模型已成功地用于波形生成任务,利用单个矢量场估计目标进行训练。虽然这些模型可以生成高保真的波形信号,但与基于GAN的模型相比,它们需要更多的ODE步骤,后者只需要一个生成步骤。此外,由于噪声矢量场估计,生成的样本通常缺乏高频信息,这无法确保高频再现。为了解决这一限制,我们通过结合固定步长的生成器修改来增强预训练的基于CFM的生成模型。我们利用重建损失和对抗反馈来加速高保真波形生成。通过对抗性流匹配优化,它只需要1,000步微调,就可以在各种客观指标上实现最先进的性能。此外,我们显着减少推理速度从16步到2或4步。此外,通过将PeriodWave的主干参数从29 M扩展到70 M以提高泛化能力,PeriodWave-Turbo实现了前所未有的性能,在LibriTTS数据集上的语音质量感知评估(PESQ)得分为4.454。音频样本、源代码和检查点将在https: github.com sh-lee-prml PeriodWave上提供。摘要:This paper introduces PeriodWave-Turbo, a high-fidelity and high-efficient waveform generation model via adversarial flow matching optimization. Recently, conditional flow matching (CFM) generative models have been successfully adopted for waveform generation tasks, leveraging a single vector field estimation objective for training. Although these models can generate high-fidelity waveform signals, they require significantly more ODE steps compared to GAN-based models, which only need a single generation step. Additionally, the generated samples often lack high-frequency information due to noisy vector field estimation, which fails to ensure high-frequency reproduction. To address this limitation, we enhance pre-trained CFM-based generative models by incorporating a fixed-step generator modification. We utilized reconstruction losses and adversarial feedback to accelerate high-fidelity waveform generation. Through adversarial flow matching optimization, it only requires 1,000 steps of fine-tuning to achieve state-of-the-art performance across various objective metrics. Moreover, we significantly reduce inference speed from 16 steps to 2 or 4 steps. Additionally, by scaling up the backbone of PeriodWave from 29M to 70M parameters for improved generalization, PeriodWave-Turbo achieves unprecedented performance, with a perceptual evaluation of speech quality (PESQ) score of 4.454 on the LibriTTS dataset. Audio samples, source code and checkpoints will be available at https: github.com sh-lee-prml PeriodWave.
【4】 Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
标题: 通过稀有和歧义词的上下文化增强基于大语言模型的语音识别
作者:Kento Nozawa,Takashi Masuko,Toru Taniguchi
备注:13 pages, 1 figure, and 7 tables
链接:点击下载PDF文件
摘要:我们开发了一个基于大语言模型(LLM)的自动语音识别(ASR)系统,可以通过提供关键字作为文本提示中的先验信息来上下文化。我们采用仅解码器架构,并使用我们的内部LLM,PLaMo-100 B,从头开始使用以日语和英语文本为主的数据集进行预训练作为解码器。我们采用预训练的Whisper编码器作为音频编码器,来自音频编码器的音频嵌入通过适配器层投影到文本嵌入空间,并与从文本提示转换的文本嵌入连接,以形成解码器的输入。通过在文本提示中提供关键字作为先验信息,我们可以将我们基于LLM的ASR系统置于上下文中,而无需修改模型架构以准确地转录输入音频中的歧义单词。实验结果表明,向解码器提供关键字可以显著提高生僻词和歧义词的识别性能。摘要:We develop a large language model (LLM) based automatic speech recognition (ASR) system that can be contextualized by providing keywords as prior information in text prompts. We adopt decoder-only architecture and use our in-house LLM, PLaMo-100B, pre-trained from scratch using datasets dominated by Japanese and English texts as the decoder. We adopt a pre-trained Whisper encoder as an audio encoder, and the audio embeddings from the audio encoder are projected to the text embedding space by an adapter layer and concatenated with text embeddings converted from text prompts to form inputs to the decoder. By providing keywords as prior information in the text prompts, we can contextualize our LLM-based ASR system without modifying the model architecture to transcribe ambiguous words in the input audio accurately. Experimental results demonstrate that providing keywords to the decoder can significantly improve the recognition performance of rare and ambiguous words.
eess.AS音频处理
【1】 Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words标题: 通过稀有和歧义词的上下文化增强基于大语言模型的语音识别
作者:Kento Nozawa,Takashi Masuko,Toru Taniguchi
备注:13 pages, 1 figure, and 7 tables
链接:点击下载PDF文件
摘要:我们开发了一个基于大语言模型(LLM)的自动语音识别(ASR)系统,可以通过提供关键字作为文本提示中的先验信息来上下文化。我们采用仅解码器架构,并使用我们的内部LLM,PLaMo-100 B,从头开始使用以日语和英语文本为主的数据集进行预训练作为解码器。我们采用预训练的Whisper编码器作为音频编码器,来自音频编码器的音频嵌入通过适配器层投影到文本嵌入空间,并与从文本提示转换的文本嵌入连接,以形成解码器的输入。通过在文本提示中提供关键字作为先验信息,我们可以将我们基于LLM的ASR系统置于上下文中,而无需修改模型架构以准确地转录输入音频中的歧义单词。实验结果表明,向解码器提供关键字可以显著提高生僻词和歧义词的识别性能。摘要:We develop a large language model (LLM) based automatic speech recognition (ASR) system that can be contextualized by providing keywords as prior information in text prompts. We adopt decoder-only architecture and use our in-house LLM, PLaMo-100B, pre-trained from scratch using datasets dominated by Japanese and English texts as the decoder. We adopt a pre-trained Whisper encoder as an audio encoder, and the audio embeddings from the audio encoder are projected to the text embedding space by an adapter layer and concatenated with text embeddings converted from text prompts to form inputs to the decoder. By providing keywords as prior information in the text prompts, we can contextualize our LLM-based ASR system without modifying the model architecture to transcribe ambiguous words in the input audio accurately. Experimental results demonstrate that providing keywords to the decoder can significantly improve the recognition performance of rare and ambiguous words.
【2】 Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
标题: 推进对比音频预训练的多粒度对齐
作者:Yiming Li,Zhifang Guo,Xiangdong Wang,Hong Liu
备注:ACM MM 2024 (Oral)
链接:点击下载PDF文件
摘要:最近的进展已经见证了音频语言联合学习,如CLAP,这表明在多模态理解任务中取得了很大成功。这些模型通常将单模态的局部表示,即框架或单词特征,聚合成全局表示,在全局表示上使用对比损失来实现粗粒度的跨模态对齐。然而,与文本的帧级对应可能会被忽视,这使得它在可解释性和细粒度挑战方面不成立,这也可能会损害粗粒度任务的性能。在这项工作中,我们的目标是在大规模对比预训练中提高粗粒度和细粒度的音频语言对齐。为了统一两种模态的粒度和潜在分布,采用共享码本来表示具有公共基的多模态全局特征,并对每个码字进行正则化以编码模态共享语义,弥合了帧和词特征之间的差距。在此基础上,引入局部感知块来净化局部模式,并设计硬负引导损失来增强对齐。在11个zero-shot粗粒度和细粒度任务上的实验表明,我们的模型不仅显著超过了基线CLAP,而且与当前SOTA作品相比,还产生了更好的或有竞争力的结果。摘要:Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-modal local representations, namely frame or word features, into global ones, on which the contrastive loss is employed to reach coarse-grained cross-modal alignment. However, frame-level correspondence with texts may be ignored, making it ill-posed on explainability and fine-grained challenges which may also undermine performances on coarse-grained tasks. In this work, we aim to improve both coarse- and fine-grained audio-language alignment in large-scale contrastive pre-training. To unify the granularity and latent distribution of two modalities, a shared codebook is adopted to represent multi-modal global features with common bases, and each codeword is regularized to encode modality-shared semantics, bridging the gap between frame and word features. Based on it, a locality-aware block is involved to purify local patterns, and a hard-negative guided loss is devised to boost alignment. Experiments on eleven zero-shot coarse- and fine-grained tasks suggest that our model not only surpasses the baseline CLAP significantly but also yields superior or competitive results compared to current SOTA works.
【3】 The evolution of inharmonicity and noisiness in contemporary popular music
标题: 当代流行音乐不和谐与噪音的演变
作者:Emmanuel Deruty,David Meredith,Stefan Lattner
备注:43 pages, 23 figures
链接:点击下载PDF文件
摘要:许多西方古典音乐使用基于声学共振的乐器。这种乐器产生谐波或准谐波的声音。另一方面,自20世纪70年代初以来,流行音乐主要是在录音室制作的。因此,流行音乐并不一定是以和声或准和声为基础的。在这项研究中,我们使用修改后的MPEG-7功能,探索和验证自1961年以来流行音乐中噪音和不和谐的使用方式。我们将这种演变置于其他广泛音乐类别的背景下,包括西方古典钢琴音乐、西方古典管弦乐和协奏曲。我们提出了新的功能,使我们能够区分噪声和不和谐造成的相对离散的偏音之间的相互作用之间的不和谐。当通过这些新特征的视角来看待当代流行音乐的历史时,我们发现自1961年以来的时期可以分为三个阶段。从1961年到1972年,不和谐性稳步增加,但噪音没有显著增加。从1972年到1986年,非谐性和噪声都有所增加。然后,自1986年以来,今天的流行音乐的不和谐性和噪音都在稳步下降,与60年代的音乐相比,噪音明显减少,但更不和谐。我们将这些观察到的趋势与这一时期音乐制作实践的发展联系起来,并通过对某些关键艺术家和曲目的重点分析来说明这些趋势。摘要:Much of Western classical music uses instruments based on acoustic resonance. Such instruments produce harmonic or quasi-harmonic sounds. On the other hand, since the early 1970s, popular music has largely been produced in the recording studio. As a result, popular music is not bound to be based on harmonic or quasi-harmonic sounds. In this study, we use modified MPEG-7 features to explore and characterise the way in which the use of noise and inharmonicity has evolved in popular music since 1961. We set this evolution in the context of other broad categories of music, including Western classical piano music, Western classical orchestral music, and musique concr ete. We propose new features that allow us to distinguish between inharmonicity resulting from noise and inharmonicity resulting from interactions between relatively discrete partials. When the history of contemporary popular music is viewed through the lens of these new features, we find that the period since 1961 can be divided into three phases. From 1961 to 1972, there was a steady increase in inharmonicity but no significant increase in noise. From 1972 to 1986, both inharmonicity and noise increased. Then, since 1986, there has been a steady decrease in both inharmonicity and noise to today's popular music which is significantly less noisy but more inharmonic than the music of the sixties. We relate these observed trends to the development of music production practice over the period and illustrate them with focused analyses of certain key artists and tracks.
【4】 Hearing Your Blood Sugar: Non-Invasive Glucose Measurement Through Simple Vocal Signals, Transforming any Speech into a Sensor with Machine Learning
标题: 聆听血糖:通过简单的声音信号进行无创血糖测量,通过机器学习将任何语音转化为传感器
作者:Nihat Ahmadli,Mehmet Ali Sarsil,Onur Ergen
备注:5 figure and 5 tables. This manuscript is a pre-print to be submitted to a journal orand a conference. arXiv admin note: substantial text overlap with arXiv:2402.13812
链接:点击下载PDF文件
摘要:有效的糖尿病管理在很大程度上依赖于血糖水平的持续监测,传统上通过侵入性和不舒服的方法来实现。虽然已经探索了各种非侵入性技术,例如光学、微波和电化学方法,但由于与复杂性、准确性和成本相关的问题,没有一种技术能够有效地取代这些侵入性技术。在这项研究中,我们提出了一种变革性的和直接的方法,利用语音分析来预测血糖水平。我们的研究调查了血糖波动和发声特征之间的关系,强调了血管动力学在发声过程中的影响。通过应用先进的机器学习算法,我们分析了声音信号的变化,并建立了与血糖水平的显着相关性。我们使用人工智能开发了一个预测模型,基于参与者的语音记录和相应的葡萄糖测量,利用逻辑回归和岭正则化。我们的研究结果表明,语音分析可以作为一个可行的非侵入性的替代血糖监测。这种创新的方法不仅有可能简化和降低与糖尿病管理相关的成本,而且旨在通过提供一种无痛和用户友好的血糖水平监测方法来提高糖尿病患者的生活质量。摘要:Effective diabetes management relies heavily on the continuous monitoring of blood glucose levels, traditionally achieved through invasive and uncomfortable methods. While various non-invasive techniques have been explored, such as optical, microwave, and electrochemical approaches, none have effectively supplanted these invasive technologies due to issues related to complexity, accuracy, and cost. In this study, we present a transformative and straightforward method that utilizes voice analysis to predict blood glucose levels. Our research investigates the relationship between fluctuations in blood glucose and vocal characteristics, highlighting the influence of blood vessel dynamics during voice production. By applying advanced machine learning algorithms, we analyzed vocal signal variations and established a significant correlation with blood glucose levels. We developed a predictive model using artificial intelligence, based on voice recordings and corresponding glucose measurements from participants, utilizing logistic regression and Ridge regularization. Our findings indicate that voice analysis may serve as a viable non-invasive alternative for glucose monitoring. This innovative approach not only has the potential to streamline and reduce the costs associated with diabetes management but also aims to enhance the quality of life for individuals living with diabetes by providing a painless and user-friendly method for monitoring blood sugar levels.
【5】 Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
标题: 通过对抗流匹配优化加速高保真波生成
作者:Sang-Hoon Lee,Ha-Yeong Choi,Seong-Whan Lee
备注:9 pages, 9 tables, 1 figure,
链接:点击下载PDF文件
摘要:本文介绍了PeriodWave-Turbo,一个高保真,高效的波形生成模型,通过对抗流匹配优化。最近,条件流匹配(CFM)生成模型已成功地用于波形生成任务,利用单个矢量场估计目标进行训练。虽然这些模型可以生成高保真的波形信号,但与基于GAN的模型相比,它们需要更多的ODE步骤,后者只需要一个生成步骤。此外,由于噪声矢量场估计,生成的样本通常缺乏高频信息,这无法确保高频再现。为了解决这一限制,我们通过结合固定步长的生成器修改来增强预训练的基于CFM的生成模型。我们利用重建损失和对抗反馈来加速高保真波形生成。通过对抗性流匹配优化,它只需要1,000步微调,就可以在各种客观指标上实现最先进的性能。此外,我们显着减少推理速度从16步到2或4步。此外,通过将PeriodWave的主干参数从29 M扩展到70 M以提高泛化能力,PeriodWave-Turbo实现了前所未有的性能,在LibriTTS数据集上的语音质量感知评估(PESQ)得分为4.454。音频样本、源代码和检查点将在https: github.com sh-lee-prml PeriodWave上提供。摘要:This paper introduces PeriodWave-Turbo, a high-fidelity and high-efficient waveform generation model via adversarial flow matching optimization. Recently, conditional flow matching (CFM) generative models have been successfully adopted for waveform generation tasks, leveraging a single vector field estimation objective for training. Although these models can generate high-fidelity waveform signals, they require significantly more ODE steps compared to GAN-based models, which only need a single generation step. Additionally, the generated samples often lack high-frequency information due to noisy vector field estimation, which fails to ensure high-frequency reproduction. To address this limitation, we enhance pre-trained CFM-based generative models by incorporating a fixed-step generator modification. We utilized reconstruction losses and adversarial feedback to accelerate high-fidelity waveform generation. Through adversarial flow matching optimization, it only requires 1,000 steps of fine-tuning to achieve state-of-the-art performance across various objective metrics. Moreover, we significantly reduce inference speed from 16 steps to 2 or 4 steps. Additionally, by scaling up the backbone of PeriodWave from 29M to 70M parameters for improved generalization, PeriodWave-Turbo achieves unprecedented performance, with a perceptual evaluation of speech quality (PESQ) score of 4.454 on the LibriTTS dataset. Audio samples, source code and checkpoints will be available at https: github.com sh-lee-prml PeriodWave.
机器翻译,仅供参考
