本文经arXiv每日学术速递授权转载
标题: GRAFX:PyTorch中音频处理图形的开源库
作者:Sungho Lee,Marco Martínez-Ramírez,Wei-Hsiang Liao,Stefan Uhlich,Giorgio Fabbro,Kyogu Lee,Yuki Mitsufuji
备注:Accepted to DAFx 2024 demo
链接:点击下载PDF文件
摘要:我们介绍了GRAFX,这是一个开源库,旨在处理PyTorch中的音频处理图。除了各种库功能外,我们还介绍了在GPU中对输入图形、信号和处理器参数进行高效并行计算的技术细节。然后,我们展示了它在音乐混合场景下的示例使用,其中大型图中每个可微处理器的参数通过梯度下降进行优化。该代码可在https: github.com sh-lee97 grafx上获得。摘要:We present GRAFX, an open-source library designed for handling audio processing graphs in PyTorch. Along with various library functionalities, we describe technical details on the efficient parallel computation of input graphs, signals, and processor parameters in GPU. Then, we show its example use under a music mixing scenario, where parameters of every differentiable processor in a large graph are optimized via gradient descent. The code is available at https: github.com sh-lee97 grafx.
【2】 Self-Supervised Learning for Multi-Channel Neural Transducer
标题: 多通道神经传感器的自我监督学习
作者:Atsushi Kojima
链接:点击下载PDF文件
摘要:自监督学习,例如wav 2 vec 2.0框架,可以显著提高端到端自动语音识别(ASR)的准确性。Wav 2 vec 2.0已应用于单通道端到端ASR模型。在这项工作中,我们探索了一种基于wav 2 vec 2.0框架的多通道端到端ASR模型的自监督学习方法。作为多通道端到端ASR模型,我们重点研究了多通道神经传感器。在预训练中,我们比较了三种不同的特征量化方法来训练多声道一致性音频编码器:联合量化,特征量化和声道量化。在精调中,我们训练了多通道共形换能器。所有实验均使用远场内部和CHiME-4数据集进行。实验结果表明,特征量化是最有效的方法。我们观察到,与远场内部数据集没有任何预训练的模型相比,字符错误率相对降低了66%。摘要:Self-supervised learning, such as with the wav2vec 2.0 framework significantly improves the accuracy of end-to-end automatic speech recognition (ASR). Wav2vec 2.0 has been applied to single-channel end-to-end ASR models. In this work, we explored a self-supervised learning method for a multi-channel end-to-end ASR model based on the wav2vec 2.0 framework. As the multi-channel end-to-end ASR model, we focused on a multi-channel neural transducer. In pre-training, we compared three different methods for feature quantization to train a multi-channel conformer audio encoder: joint quantization, feature-wise quantization and channel-wise quantization. In fine-tuning, we trained the multi-channel conformer-transducer. All experiments were conducted using the far-field in-house and CHiME-4 datasets. The results of the experiments showed that feature-wise quantization was the most effective among the methods. We observed a 66% relative reduction in character error rate compared with the model without any pre-training for the far-field in-house dataset.
【3】 Automatic Voice Identification after Speech Resynthesis using PPG
标题: 使用PGP进行语音再合成后的自动语音识别
作者:Thibault Gaudier,Marie Tahon,Anthony Larcher,Yannick Estève
Journal-ref:Speaker and Language Recognition Workshop - Odyssey, Jun 2024, Qu{'e}bec (Canada), Canada
链接:点击下载PDF文件
摘要:语音再合成是一种通用的任务,我们希望将音频与另一个音频作为输入进行合成,这对于媒体监视器和记者来说是有应用的。在语音再合成所解决的不同任务中,语音转换保留了语言信息,同时修改了说话人的身份,而语音编辑保留了说话人的身份,但修改了一些单词。在这两种情况下,语音后验图(Phonetic PosteriorGrams,PPG)是音素的帧级概率表示,通常被认为是与说话人无关的。本文提出了一种基于PPG的语音再合成系统。感知评估评估它产生正确的音频质量。然后,我们证明了自动说话人验证模型在用PPG重新合成之后不能恢复源说话人,即使当该模型在合成数据上训练时也是如此。摘要:Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, voice conversion preserves the linguistic information while modifying the identity of the speaker, and speech edition preserves the identity of the speaker but some words are modified.In both cases, we need to disentangle speaker and phonetic contents in intermediate representations.Phonetic PosteriorGrams (PPG) are a frame-level probabilistic representation of phonemes, and are usually considered speaker-independent.This paper presents a PPG-based speech resynthesis system.A perceptive evaluation assesses that it produces correct audio quality.Then, we demonstrate that an automatic speaker verification model is not able to recover the source speaker after re-synthesis with PPG, even when the model is trained on synthetic data.
【4】 Text Conditioned Symbolic Drumbeat Generation using Latent Diffusion Models
标题: 使用潜在扩散模型的文本条件符号鼓乐生成
作者:Pushkar Jajoria,James McDermott
链接:点击下载PDF文件
摘要:本研究提出一种基于文本的潜在扩散模型生成鼓点的方法。它使用从训练数据文件名中提取的信息条件文本。通过在多模态网络中进行对比学习来预训练文本和鼓点编码器,按照CLIP进行对齐,我们将文本和音乐的模态紧密对齐。此外,我们还研究了一种基于多热文本编码的替代文本编码器。受音乐多分辨率性质的启发,我们提出了一种新的LSTM变体MultiResolutionLSTM,旨在独立地在各种分辨率下运行。与图像空间中的最新LDM一样,它通过在预训练的无条件自动编码器提供的潜在空间中运行扩散来加速生成过程。我们通过测量距离(在二进制钢琴和潜在空间中)与训练数据集和生成的鼓点之间的距离来展示生成的鼓点的原创性和多样性。我们还通过听力测试来评估生成的鼓点,重点是质量,提示文本的适用性和新颖性。我们表明,生成的鼓点是新颖的,易于提示文本,并在质量上与人类音乐家创造的那些。摘要:This study introduces a text-conditioned approach to generating drumbeats with Latent Diffusion Models (LDMs). It uses informative conditioning text extracted from training data filenames. By pretraining a text and drumbeat encoder through contrastive learning within a multimodal network, aligned following CLIP, we align the modalities of text and music closely. Additionally, we examine an alternative text encoder based on multihot text encodings. Inspired by musics multi-resolution nature, we propose a novel LSTM variant, MultiResolutionLSTM, designed to operate at various resolutions independently. In common with recent LDMs in the image space, it speeds up the generation process by running diffusion in a latent space provided by a pretrained unconditional autoencoder. We demonstrate the originality and variety of the generated drumbeats by measuring distance (both over binary pianorolls and in the latent space) versus the training dataset and among the generated drumbeats. We also assess the generated drumbeats through a listening test focused on questions of quality, aptness for the prompt text, and novelty. We show that the generated drumbeats are novel and apt to the prompt text, and comparable in quality to those created by human musicians.
eess.AS音频处理
【1】 GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch标题: GRAFX:PyTorch中音频处理图形的开源库
作者:Sungho Lee,Marco Martínez-Ramírez,Wei-Hsiang Liao,Stefan Uhlich,Giorgio Fabbro,Kyogu Lee,Yuki Mitsufuji
备注:Accepted to DAFx 2024 demo
链接:点击下载PDF文件
摘要:我们介绍了GRAFX,这是一个开源库,旨在处理PyTorch中的音频处理图。除了各种库功能外,我们还介绍了在GPU中对输入图形、信号和处理器参数进行高效并行计算的技术细节。然后,我们展示了它在音乐混合场景下的示例使用,其中大型图中每个可微处理器的参数通过梯度下降进行优化。该代码可在https: github.com sh-lee97 grafx上获得。摘要:We present GRAFX, an open-source library designed for handling audio processing graphs in PyTorch. Along with various library functionalities, we describe technical details on the efficient parallel computation of input graphs, signals, and processor parameters in GPU. Then, we show its example use under a music mixing scenario, where parameters of every differentiable processor in a large graph are optimized via gradient descent. The code is available at https: github.com sh-lee97 grafx.
【2】 Self-Supervised Learning for Multi-Channel Neural Transducer
标题: 多通道神经传感器的自我监督学习
作者:Atsushi Kojima
链接:点击下载PDF文件
摘要:自监督学习,例如wav 2 vec 2.0框架,可以显著提高端到端自动语音识别(ASR)的准确性。Wav 2 vec 2.0已应用于单通道端到端ASR模型。在这项工作中,我们探索了一种基于wav 2 vec 2.0框架的多通道端到端ASR模型的自监督学习方法。作为多通道端到端ASR模型,我们重点研究了多通道神经传感器。在预训练中,我们比较了三种不同的特征量化方法来训练多声道一致性音频编码器:联合量化,特征量化和声道量化。在精调中,我们训练了多通道共形换能器。所有实验均使用远场内部和CHiME-4数据集进行。实验结果表明,特征量化是最有效的方法。我们观察到,与远场内部数据集没有任何预训练的模型相比,字符错误率相对降低了66%。摘要:Self-supervised learning, such as with the wav2vec 2.0 framework significantly improves the accuracy of end-to-end automatic speech recognition (ASR). Wav2vec 2.0 has been applied to single-channel end-to-end ASR models. In this work, we explored a self-supervised learning method for a multi-channel end-to-end ASR model based on the wav2vec 2.0 framework. As the multi-channel end-to-end ASR model, we focused on a multi-channel neural transducer. In pre-training, we compared three different methods for feature quantization to train a multi-channel conformer audio encoder: joint quantization, feature-wise quantization and channel-wise quantization. In fine-tuning, we trained the multi-channel conformer-transducer. All experiments were conducted using the far-field in-house and CHiME-4 datasets. The results of the experiments showed that feature-wise quantization was the most effective among the methods. We observed a 66% relative reduction in character error rate compared with the model without any pre-training for the far-field in-house dataset.
【3】 Automatic Voice Identification after Speech Resynthesis using PPG
标题: 使用PGP进行语音再合成后的自动语音识别
作者:Thibault Gaudier,Marie Tahon,Anthony Larcher,Yannick Estève
Journal-ref:Speaker and Language Recognition Workshop - Odyssey, Jun 2024, Qu{'e}bec (Canada), Canada
链接:点击下载PDF文件
摘要:语音再合成是一种通用的任务,我们希望将音频与另一个音频作为输入进行合成,这对于媒体监视器和记者来说是有应用的。在语音再合成所解决的不同任务中,语音转换保留了语言信息,同时修改了说话人的身份,而语音编辑保留了说话人的身份,但修改了一些单词。在这两种情况下,语音后验图(Phonetic PosteriorGrams,PPG)是音素的帧级概率表示,通常被认为是与说话人无关的。本文提出了一种基于PPG的语音再合成系统。感知评估评估它产生正确的音频质量。然后,我们证明了自动说话人验证模型在用PPG重新合成之后不能恢复源说话人,即使当该模型在合成数据上训练时也是如此。摘要:Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, voice conversion preserves the linguistic information while modifying the identity of the speaker, and speech edition preserves the identity of the speaker but some words are modified.In both cases, we need to disentangle speaker and phonetic contents in intermediate representations.Phonetic PosteriorGrams (PPG) are a frame-level probabilistic representation of phonemes, and are usually considered speaker-independent.This paper presents a PPG-based speech resynthesis system.A perceptive evaluation assesses that it produces correct audio quality.Then, we demonstrate that an automatic speaker verification model is not able to recover the source speaker after re-synthesis with PPG, even when the model is trained on synthetic data.
【4】 Text Conditioned Symbolic Drumbeat Generation using Latent Diffusion Models
标题: 使用潜在扩散模型的文本条件符号鼓乐生成
作者:Pushkar Jajoria,James McDermott
链接:点击下载PDF文件
摘要:本研究提出一种基于文本的潜在扩散模型生成鼓点的方法。它使用从训练数据文件名中提取的信息条件文本。通过在多模态网络中进行对比学习来预训练文本和鼓点编码器,按照CLIP进行对齐,我们将文本和音乐的模态紧密对齐。此外,我们还研究了一种基于多热文本编码的替代文本编码器。受音乐多分辨率本质的启发,我们提出了一种新型LSTM变体MultiResolutionLSTM,旨在独立地在各种分辨率下运行。与图像空间中的最新LDM一样,它通过在预训练的无条件自动编码器提供的潜在空间中运行扩散来加速生成过程。我们通过测量距离(在二进制钢琴和潜在空间中)与训练数据集和生成的鼓点之间的距离来展示生成的鼓点的原创性和多样性。我们还通过听力测试来评估生成的鼓点,重点是质量,提示文本的适用性和新颖性。我们表明,生成的鼓点是新颖的,易于提示文本,并在质量上与人类音乐家创造的那些。摘要:This study introduces a text-conditioned approach to generating drumbeats with Latent Diffusion Models (LDMs). It uses informative conditioning text extracted from training data filenames. By pretraining a text and drumbeat encoder through contrastive learning within a multimodal network, aligned following CLIP, we align the modalities of text and music closely. Additionally, we examine an alternative text encoder based on multihot text encodings. Inspired by musics multi-resolution nature, we propose a novel LSTM variant, MultiResolutionLSTM, designed to operate at various resolutions independently. In common with recent LDMs in the image space, it speeds up the generation process by running diffusion in a latent space provided by a pretrained unconditional autoencoder. We demonstrate the originality and variety of the generated drumbeats by measuring distance (both over binary pianorolls and in the latent space) versus the training dataset and among the generated drumbeats. We also assess the generated drumbeats through a listening test focused on questions of quality, aptness for the prompt text, and novelty. We show that the generated drumbeats are novel and apt to the prompt text, and comparable in quality to those created by human musicians.
机器翻译,仅供参考
