今日论文合集:cs.SD语音7篇,eess.AS音频处理7篇。本文经arXiv每日学术速递授权转载
【1】Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
标题:Llama—VITS:利用语义感知增强TTS合成作者:Xincan Feng,Akifumi Yoshimoto备注:9 pages, 2 figures, 4 tables摘要:自然语言处理(NLP)的最新进展已经看到大规模语言模型(LLM)在为各种目的生成高质量文本方面表现出色。值得注意的是,在文本到语音(TTS)系统中,集成BERT的语义令牌生成强调了语义内容在产生连贯的语音输出的重要性。尽管如此,LLM在增强TTS合成方面的具体效用仍然相当有限。本研究介绍了一种创新的方法,Llama-VITS,它通过使用LLM丰富文本的语义内容来增强TTS合成。Llama-VITS将Llama 2的语义嵌入与领先的端到端TTS框架VITS模型集成在一起。通过利用Llama 2进行主要语音合成过程,我们的实验表明,Llama-VITS匹配原始VITS(ORI-VITS)的自然度,并且在LJSpeech数据集上结合了BERT(BERT-VITS),这是一个中性清晰语音的大量集合。此外,我们的方法显着增强了ECOV_DB_bea_sem数据集上的情感表达能力,这是从ECOV_DB数据集中精选的情感一致的语音,突出了其生成情感语音的潜力。摘要:Recent advancements in Natural Language Processing (NLP) have seen Large-scale Language Models (LLMs) excel at producing high-quality text for various purposes. Notably, in Text-To-Speech (TTS) systems, the integration of BERT for semantic token generation has underscored the importance of semantic content in producing coherent speech outputs. Despite this, the specific utility of LLMs in enhancing TTS synthesis remains considerably limited. This research introduces an innovative approach, Llama-VITS, which enhances TTS synthesis by enriching the semantic content of text using LLM. Llama-VITS integrates semantic embeddings from Llama2 with the VITS model, a leading end-to-end TTS framework. By leveraging Llama2 for the primary speech synthesis process, our experiments demonstrate that Llama-VITS matches the naturalness of the original VITS (ORI-VITS) and those incorporate BERT (BERT-VITS), on the LJSpeech dataset, a substantial collection of neutral, clear speech. Moreover, our method significantly enhances emotive expressiveness on the EmoV_DB_bea_sem dataset, a curated selection of emotionally consistent speech from the EmoV_DB dataset, highlighting its potential to generate emotive speech.【2】 Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment作者:Yuka Hashizume,Li Li,Atsushi Miyashita,Tomoki Toda摘要:为了实现一个灵活的推荐和检索系统,它是可取的计算音乐相似性,通过集中在多个部分元素的音乐作品,并允许用户选择他们想要focuson.A先前的研究提出使用多个单独的网络计算音乐相似性的基础上,每个乐器的声音,但它是不切实际的,使用每个信号作为一个查询搜索系统。使用分离的乐器声音替代地导致由于伪影而导致的较低准确性。在本文中,我们提出了一种方法来计算相似性集中在每个乐器的声音与一个单一的网络,以混合的声音作为输入,而不是单独的乐器的声音。具体来说,我们为每个仪器设计了一个具有解纠缠维度的单一相似性嵌入空间,由条件相似性网络提取,该网络通过使用掩码的三重丢失进行训练。实验结果表明:(1)与使用单独的声音作为输入的单个网络相比,该方法可以获得更准确的特征表示;(2)每个子嵌入空间都可以保持相应乐器的特征;(3)通过该方法选择关注每个乐器声音的相似音乐作品可以获得人类的同意,特别是在鼓和吉他中。摘要:To achieve a flexible recommendation and retrieval system, it is desirable to calculate music similarity by focusing on multiple partial elements of musical pieces and allowing the users to select the element they want to focus on. A previous study proposed using multiple individual networks for calculating music similarity based on each instrumental sound, but it is impractical to use each signal as a query in search systems. Using separated instrumental sounds alternatively resulted in less accuracy due to artifacts. In this paper, we propose a method to compute similarities focusing on each instrumental sound with a single network that takes mixed sounds as input instead of individual instrumental sounds. Specifically, we design a single similarity embedding space with disentangled dimensions for each instrument, extracted by Conditional Similarity Networks, which is trained by the triplet loss using masks. Experimental results have shown that (1) the proposed method can obtain more accurate feature representation than using individual networks using separated sounds as input, (2) each sub-embedding space can hold the characteristics of the corresponding instrument, and (3) the selection of similar musical pieces focusing on each instrumental sound by the proposed method can obtain human consent, especially in drums and guitar.
【3】 VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing标题:VoiceShop:一个统一的语音到语音框架,用于保持身份的Zero-Shot语音编辑作者:Philip Anastassiou,Zhenyu Tang,Kainan Peng,Dongya Jia,Jiaxin Li,Ming Tu,Yuping Wang,Yuxuan Wang,Mingbo Ma摘要:我们提出了VoiceShop,一种新颖的语音到语音的框架,可以修改语音的多个属性,如年龄,性别,口音和语音风格,在一个单一的向前传递,同时保留输入扬声器的音色。以前的作品已被限制到专门的模型,只能单独编辑这些属性,并遭受以下缺陷:转换效果的幅度是弱的,没有zero-shot能力的分布扬声器,或合成的输出表现出音色泄漏,改变扬声器的感知身份。我们的工作在一个简单的模块化框架中提出了这些问题的解决方案,该框架基于条件扩散骨干模型,具有可选的标准化基于流和序列到序列的扬声器属性编辑模块,其组件可以在推理过程中组合或删除,以满足各种任务,而无需额外的模型微调。音频样本可在https://voiceshopai.github.io上获得摘要:We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass while preserving the input speaker's timbre. Previous works have been constrained to specialized models that can only edit these attributes individually and suffer from the following pitfalls: the magnitude of the conversion effect is weak, there is no zero-shot capability for out-of-distribution speakers, or the synthesized outputs exhibit timbre leakage which changes the speaker's perceived identity. Our work proposes solutions for each of these issues in a simple modular framework based on a conditional diffusion backbone model with optional normalizing flow-based and sequence-to-sequence speaker attribute-editing modules, whose components can be combined or removed during inference to meet a wide array of tasks without additional model finetuning. Audio samples are available at https://voiceshopai.github.io
【4】 Efficient Sound Field Reconstruction with Conditional Invertible Neural Networks作者:Xenofon Karakonstantis,Efren Fernandez-Grande,Peter Gerstoft摘要:在这项研究中,我们介绍了一种估计混响环境中的声场使用条件可逆神经网络(CINN)的方法。声场重建可能会受到实验误差、有限的空间数据、模型不匹配和长推理时间的阻碍,从而导致潜在的缺陷和延长的表征。此外,管理固有的不确定性的复杂性往往会增加计算需求,或者在模型中被忽略。我们的方法旨在平衡精度和计算效率,同时结合不确定性估计,以定制重建的特定需求。通过用随机波场的蒙特卡罗模拟训练CINN,我们的方法减少了对大量数据集的依赖,并能够从稀疏的实验数据中进行推断。CINN被证明是通用的重建房间脉冲响应(RIR),作为一个最大后验概率估计的似然模型或作为一个近似的后验分布,通过摊销贝叶斯推断。与传统的贝叶斯方法相比,CINN以更高的效率实现了相似的精度,并且不需要其适应不同的声场条件。摘要:In this study, we introduce a method for estimating sound fields in reverberant environments using a conditional invertible neural network (CINN). Sound field reconstruction can be hindered by experimental errors, limited spatial data, model mismatches, and long inference times, leading to potentially flawed and prolonged characterizations. Further, the complexity of managing inherent uncertainties often escalates computational demands or is neglected in models. Our approach seeks to balance accuracy and computational efficiency, while incorporating uncertainty estimates to tailor reconstructions to specific needs. By training a CINN with Monte Carlo simulations of random wave fields, our method reduces the dependency on extensive datasets and enables inference from sparse experimental data. The CINN proves versatile at reconstructing Room Impulse Responses (RIRs), by acting either as a likelihood model for maximum a posteriori estimation or as an approximate posterior distribution through amortized Bayesian inference. Compared to traditional Bayesian methods, the CINN achieves similar accuracy with greater efficiency and without requiring its adaptation to distinct sound field conditions.
【5】 Towards Efficient and Real-Time Piano Transcription Using Neural Autoregressive Models作者:Taegyun Kwon,Dasaem Jeong,Juhan Nam备注:11 pages, 8 figures, preprint摘要:近年来,神经网络设计的进步和大规模标记数据集的可用性导致钢琴转录模型的准确性显着提高。然而,大多数以前的工作集中在高性能的离线转录,忽略了故意考虑模型的大小。这项工作的目标是实现钢琴转录的实时推理,同时确保高性能和轻量级。为此,我们提出了卷积递归神经网络的新架构,重新设计了现有的自回归钢琴转录模型。首先,我们通过向CNN模块添加频率调节的FilLM层来扩展声学模块,以适应频率轴上的卷积滤波器。其次,我们通过使用一个音高方向的LSTM来改进音符状态序列建模,该LSTM专注于音符中的音符状态转换。此外,我们增加了自回归连接与增强的递归上下文。使用这些组件,我们提出了两种类型的模型,一个高性能和高紧凑性。通过大量的实验,我们表明,所提出的模型是最先进的模型在MAESTRO数据集上的注意准确性。我们还调查了有效的模型大小和实时推理延迟逐渐精简的架构。最后,我们对未知钢琴数据集进行了交叉数据评估和深入分析,以阐明所提出的组件在音符长度和音高范围方面的效果。摘要:In recent years, advancements in neural network designs and the availability of large-scale labeled datasets have led to significant improvements in the accuracy of piano transcription models. However, most previous work focused on high-performance offline transcription, neglecting deliberate consideration of model size. The goal of this work is to implement real-time inference for piano transcription while ensuring both high performance and lightweight. To this end, we propose novel architectures for convolutional recurrent neural networks, redesigning an existing autoregressive piano transcription model. First, we extend the acoustic module by adding a frequency-conditioned FiLM layer to the CNN module to adapt the convolutional filters on the frequency axis. Second, we improve note-state sequence modeling by using a pitchwise LSTM that focuses on note-state transitions within a note. In addition, we augment the autoregressive connection with an enhanced recursive context. Using these components, we propose two types of models; one for high performance and the other for high compactness. Through extensive experiments, we show that the proposed models are comparable to state-of-the-art models in terms of note accuracy on the MAESTRO dataset. We also investigate the effective model size and real-time inference latency by gradually streamlining the architecture. Finally, we conduct cross-data evaluation on unseen piano datasets and in-depth analysis to elucidate the effect of the proposed components in the view of note length and pitch range.【6】 What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy Conditions标题:LEARNable前端(LEAF)学习什么?使每通道能量归一化(PCEN)适应噪声条件作者:Hanyu Meng,Vidhyasaharan Sethu,Eliathamby Ambikairajah摘要:LEArnable前端(LEAF)在各种语音处理系统中的应用越来越受到人们的关注。然而,缺乏对实际学到的东西以及培训前端不同组成部分的相对重要性的分析。在本文中,我们调查这个问题上的关键字定位,基于语音的情感识别和语言识别任务,并发现用于频谱分解的滤波器和用于估计频谱能量变化的低通滤波器没有表现出学习和每通道能量归一化(PCEN)是学习的关键组成部分。在此之后,我们探索了仅使用少量噪声数据调整PCEN层的潜力,使其能够学习更适合噪声条件的适当动态范围压缩。这反过来又使一个在干净的语音上训练的系统能够更准确地处理嘈杂的测试数据,正如本文报告的实验结果所证明的那样。摘要:There is increasing interest in the use of the LEArnable Front-end (LEAF) in a variety of speech processing systems. However, there is a dearth of analyses of what is actually learnt and the relative importance of training the different components of the front-end. In this paper, we investigate this question on keyword spotting, speech-based emotion recognition and language identification tasks and find that the filters for spectral decomposition and the low pass filter used to estimate spectral energy variations exhibit no learning and the per-channel energy normalisation (PCEN) is the key component that is learnt. Following this, we explore the potential of adapting only the PCEN layer with a small amount of noisy data to enable it to learn appropriate dynamic range compression that better suits the noise conditions. This in turn enables a system trained on clean speech to work more accurately on noisy test data as demonstrated by the experimental results reported in this paper.【7】 CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations标题:CoVoMix:为类似人类的多人对话推进Zero-Shot语音生成作者:Leying Zhang,Yao Qian,Long Zhou,Shujie Liu,Dongmei Wang,Xiaofei Wang,Midia Yousefi,Yanmin Qian,Jinyu Li,Lei He,Sheng Zhao,Michael Zeng摘要:zero-shot文本到语音(TTS)建模的最新进展在生成高保真和多样化语音方面取得了重大进展。然而,对话生成以及在语音中实现类似人类的自然性仍然是该领域的一个挑战。在本文中,我们介绍了CoVoMix:会话语音混合生成,一个新的模型,zero-shot,类人,多说话人,多轮对话语音生成。CoVoMix能够首先将对话文本转换为多个离散令牌流,每个令牌流表示单个说话者的语义信息。这些令牌流然后被馈送到基于流匹配的声学模型中以生成混合的梅尔频谱图。最后,使用HiFi-GAN模型产生语音波形。此外,我们设计了一套全面的衡量对话建模和生成的有效性的指标。我们的实验结果表明,CoVoMix可以生成对话,不仅是人类一样的自然性和连贯性,但也涉及多个谈话者参与多轮对话。这些对话,在一个单一的通道内产生,其特点是无缝的语音过渡,包括重叠的语音,和适当的语言行为,如笑声。音频样本可在https://aka.ms/covomix上获得。摘要:Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge in the field. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix is capable of first converting dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. These dialogues, generated within a single channel, are characterized by seamless speech transitions, including overlapping speech, and appropriate paralinguistic behaviors such as laughter. Audio samples are available at https://aka.ms/covomix.
【1】 Efficient Sound Field Reconstruction with Conditional Invertible Neural Networks作者:Xenofon Karakonstantis,Efren Fernandez-Grande,Peter Gerstoft摘要:在这项研究中,我们介绍了一种估计混响环境中的声场使用条件可逆神经网络(CINN)的方法。声场重建可能会受到实验误差、有限的空间数据、模型不匹配和长推理时间的阻碍,从而导致潜在的缺陷和延长的表征。此外,管理固有的不确定性的复杂性往往会增加计算需求,或者在模型中被忽略。我们的方法旨在平衡精度和计算效率,同时结合不确定性估计,以定制重建的特定需求。通过用随机波场的蒙特卡罗模拟训练CINN,我们的方法减少了对大量数据集的依赖,并能够从稀疏的实验数据中进行推断。CINN被证明是通用的重建房间脉冲响应(RIR),作为一个最大后验概率估计的似然模型或作为一个近似的后验分布,通过摊销贝叶斯推断。与传统的贝叶斯方法相比,CINN以更高的效率实现了相似的精度,并且不需要其适应不同的声场条件。摘要:In this study, we introduce a method for estimating sound fields in reverberant environments using a conditional invertible neural network (CINN). Sound field reconstruction can be hindered by experimental errors, limited spatial data, model mismatches, and long inference times, leading to potentially flawed and prolonged characterizations. Further, the complexity of managing inherent uncertainties often escalates computational demands or is neglected in models. Our approach seeks to balance accuracy and computational efficiency, while incorporating uncertainty estimates to tailor reconstructions to specific needs. By training a CINN with Monte Carlo simulations of random wave fields, our method reduces the dependency on extensive datasets and enables inference from sparse experimental data. The CINN proves versatile at reconstructing Room Impulse Responses (RIRs), by acting either as a likelihood model for maximum a posteriori estimation or as an approximate posterior distribution through amortized Bayesian inference. Compared to traditional Bayesian methods, the CINN achieves similar accuracy with greater efficiency and without requiring its adaptation to distinct sound field conditions.
【2】 Towards Efficient and Real-Time Piano Transcription Using Neural Autoregressive Models作者:Taegyun Kwon,Dasaem Jeong,Juhan Nam备注:11 pages, 8 figures, preprint摘要:近年来,神经网络设计的进步和大规模标记数据集的可用性导致钢琴转录模型的准确性显着提高。然而,大多数以前的工作集中在高性能的离线转录,忽略了故意考虑模型的大小。这项工作的目标是实现钢琴转录的实时推理,同时确保高性能和轻量级。为此,我们提出了卷积递归神经网络的新架构,重新设计了现有的自回归钢琴转录模型。首先,我们通过向CNN模块添加频率调节的FilLM层来扩展声学模块,以适应频率轴上的卷积滤波器。其次,我们通过使用一个音高方向的LSTM来改进音符状态序列建模,该LSTM专注于音符中的音符状态转换。此外,我们增加了自回归连接与增强的递归上下文。使用这些组件,我们提出了两种类型的模型,一个高性能和高紧凑性。通过大量的实验,我们表明,所提出的模型是最先进的模型在MAESTRO数据集上的注意准确性。我们还调查了有效的模型大小和实时推理延迟逐渐精简的架构。最后,我们对未知钢琴数据集进行了交叉数据评估和深入分析,以阐明所提出的组件在音符长度和音高范围方面的效果。摘要:In recent years, advancements in neural network designs and the availability of large-scale labeled datasets have led to significant improvements in the accuracy of piano transcription models. However, most previous work focused on high-performance offline transcription, neglecting deliberate consideration of model size. The goal of this work is to implement real-time inference for piano transcription while ensuring both high performance and lightweight. To this end, we propose novel architectures for convolutional recurrent neural networks, redesigning an existing autoregressive piano transcription model. First, we extend the acoustic module by adding a frequency-conditioned FiLM layer to the CNN module to adapt the convolutional filters on the frequency axis. Second, we improve note-state sequence modeling by using a pitchwise LSTM that focuses on note-state transitions within a note. In addition, we augment the autoregressive connection with an enhanced recursive context. Using these components, we propose two types of models; one for high performance and the other for high compactness. Through extensive experiments, we show that the proposed models are comparable to state-of-the-art models in terms of note accuracy on the MAESTRO dataset. We also investigate the effective model size and real-time inference latency by gradually streamlining the architecture. Finally, we conduct cross-data evaluation on unseen piano datasets and in-depth analysis to elucidate the effect of the proposed components in the view of note length and pitch range.【3】 What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy Conditions标题:LEARNable前端(LEAF)学习什么?使每通道能量归一化(PCEN)适应噪声条件作者:Hanyu Meng,Vidhyasaharan Sethu,Eliathamby Ambikairajah摘要:LEArnable前端(LEAF)在各种语音处理系统中的应用越来越受到人们的关注。然而,缺乏对实际学到的东西以及培训前端不同组成部分的相对重要性的分析。在本文中,我们调查这个问题上的关键字定位,基于语音的情感识别和语言识别任务,并发现用于频谱分解的滤波器和用于估计频谱能量变化的低通滤波器没有表现出学习和每通道能量归一化(PCEN)是学习的关键组成部分。在此之后,我们探索了仅使用少量噪声数据调整PCEN层的潜力,使其能够学习更适合噪声条件的适当动态范围压缩。这反过来又使一个在干净的语音上训练的系统能够更准确地处理嘈杂的测试数据,正如本文报告的实验结果所证明的那样。摘要:There is increasing interest in the use of the LEArnable Front-end (LEAF) in a variety of speech processing systems. However, there is a dearth of analyses of what is actually learnt and the relative importance of training the different components of the front-end. In this paper, we investigate this question on keyword spotting, speech-based emotion recognition and language identification tasks and find that the filters for spectral decomposition and the low pass filter used to estimate spectral energy variations exhibit no learning and the per-channel energy normalisation (PCEN) is the key component that is learnt. Following this, we explore the potential of adapting only the PCEN layer with a small amount of noisy data to enable it to learn appropriate dynamic range compression that better suits the noise conditions. This in turn enables a system trained on clean speech to work more accurately on noisy test data as demonstrated by the experimental results reported in this paper.【4】 CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations标题:CoVoMix:为类似人类的多人对话推进Zero-Shot语音生成作者:Leying Zhang,Yao Qian,Long Zhou,Shujie Liu,Dongmei Wang,Xiaofei Wang,Midia Yousefi,Yanmin Qian,Jinyu Li,Lei He,Sheng Zhao,Michael Zeng摘要:zero-shot文本到语音(TTS)建模的最新进展在生成高保真和多样化语音方面取得了重大进展。然而,对话生成以及在语音中实现类似人类的自然性仍然是该领域的一个挑战。在本文中,我们介绍了CoVoMix:会话语音混合生成,一个新的模型,zero-shot,类人,多说话人,多轮对话语音生成。CoVoMix能够首先将对话文本转换为多个离散令牌流,每个令牌流表示单个说话者的语义信息。这些令牌流然后被馈送到基于流匹配的声学模型中以生成混合的梅尔频谱图。最后,使用HiFi-GAN模型产生语音波形。此外,我们设计了一套全面的衡量对话建模和生成的有效性的指标。我们的实验结果表明,CoVoMix可以生成对话,不仅是人类一样的自然性和连贯性,但也涉及多个谈话者参与多轮对话。这些对话,在一个单一的通道内产生,其特点是无缝的语音过渡,包括重叠的语音,和适当的语言行为,如笑声。音频样本可在https://aka.ms/covomix上获得。摘要:Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to be a challenge in the field. In this paper, we introduce CoVoMix: Conversational Voice Mixture Generation, a novel model for zero-shot, human-like, multi-speaker, multi-round dialogue speech generation. CoVoMix is capable of first converting dialogue text into multiple streams of discrete tokens, with each token stream representing semantic information for individual talkers. These token streams are then fed into a flow-matching based acoustic model to generate mixed mel-spectrograms. Finally, the speech waveforms are produced using a HiFi-GAN model. Furthermore, we devise a comprehensive set of metrics for measuring the effectiveness of dialogue modeling and generation. Our experimental results show that CoVoMix can generate dialogues that are not only human-like in their naturalness and coherence but also involve multiple talkers engaging in multiple rounds of conversation. These dialogues, generated within a single channel, are characterized by seamless speech transitions, including overlapping speech, and appropriate paralinguistic behaviors such as laughter. Audio samples are available at https://aka.ms/covomix.
【5】 Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness标题:Llama—VITS:利用语义感知增强TTS合成作者:Xincan Feng,Akifumi Yoshimoto备注:9 pages, 2 figures, 4 tables摘要:自然语言处理(NLP)的最新进展已经看到大规模语言模型(LLM)在为各种目的生成高质量文本方面表现出色。值得注意的是,在文本到语音(TTS)系统中,集成BERT的语义令牌生成强调了语义内容在产生连贯的语音输出的重要性。尽管如此,LLM在增强TTS合成方面的具体效用仍然相当有限。本研究介绍了一种创新的方法,Llama-VITS,它通过使用LLM丰富文本的语义内容来增强TTS合成。Llama-VITS将Llama 2的语义嵌入与领先的端到端TTS框架VITS模型集成在一起。通过利用Llama 2进行主要语音合成过程,我们的实验表明,Llama-VITS匹配原始VITS(ORI-VITS)的自然度,并且在LJSpeech数据集上结合了BERT(BERT-VITS),这是一个中性清晰语音的大量集合。此外,我们的方法显着增强了ECOV_DB_bea_sem数据集上的情感表达能力,这是从ECOV_DB数据集中精选的情感一致的语音,突出了其生成情感语音的潜力。摘要:Recent advancements in Natural Language Processing (NLP) have seen Large-scale Language Models (LLMs) excel at producing high-quality text for various purposes. Notably, in Text-To-Speech (TTS) systems, the integration of BERT for semantic token generation has underscored the importance of semantic content in producing coherent speech outputs. Despite this, the specific utility of LLMs in enhancing TTS synthesis remains considerably limited. This research introduces an innovative approach, Llama-VITS, which enhances TTS synthesis by enriching the semantic content of text using LLM. Llama-VITS integrates semantic embeddings from Llama2 with the VITS model, a leading end-to-end TTS framework. By leveraging Llama2 for the primary speech synthesis process, our experiments demonstrate that Llama-VITS matches the naturalness of the original VITS (ORI-VITS) and those incorporate BERT (BERT-VITS), on the LJSpeech dataset, a substantial collection of neutral, clear speech. Moreover, our method significantly enhances emotive expressiveness on the EmoV_DB_bea_sem dataset, a curated selection of emotionally consistent speech from the EmoV_DB dataset, highlighting its potential to generate emotive speech.
【6】 Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment作者:Yuka Hashizume,Li Li,Atsushi Miyashita,Tomoki Toda摘要:为了实现一个灵活的推荐和检索系统,它是可取的计算音乐相似性,通过集中在多个部分元素的音乐作品,并允许用户选择他们想要focuson.A先前的研究提出使用多个单独的网络计算音乐相似性的基础上,每个乐器的声音,但它是不切实际的,使用每个信号作为一个查询搜索系统。使用分离的乐器声音替代地导致由于伪影而导致的较低准确性。在本文中,我们提出了一种方法来计算相似性集中在每个乐器的声音与一个单一的网络,以混合的声音作为输入,而不是单独的乐器的声音。具体来说,我们为每个仪器设计了一个具有解纠缠维度的单一相似性嵌入空间,由条件相似性网络提取,该网络通过使用掩码的三重丢失进行训练。实验结果表明:(1)与使用单独的声音作为输入的单个网络相比,该方法可以获得更准确的特征表示;(2)每个子嵌入空间都可以保持相应乐器的特征;(3)通过该方法选择关注每个乐器声音的相似音乐作品可以获得人类的同意,特别是在鼓和吉他中。摘要:To achieve a flexible recommendation and retrieval system, it is desirable to calculate music similarity by focusing on multiple partial elements of musical pieces and allowing the users to select the element they want to focus on. A previous study proposed using multiple individual networks for calculating music similarity based on each instrumental sound, but it is impractical to use each signal as a query in search systems. Using separated instrumental sounds alternatively resulted in less accuracy due to artifacts. In this paper, we propose a method to compute similarities focusing on each instrumental sound with a single network that takes mixed sounds as input instead of individual instrumental sounds. Specifically, we design a single similarity embedding space with disentangled dimensions for each instrument, extracted by Conditional Similarity Networks, which is trained by the triplet loss using masks. Experimental results have shown that (1) the proposed method can obtain more accurate feature representation than using individual networks using separated sounds as input, (2) each sub-embedding space can hold the characteristics of the corresponding instrument, and (3) the selection of similar musical pieces focusing on each instrumental sound by the proposed method can obtain human consent, especially in drums and guitar.
【7】 VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing标题:VoiceShop:一个统一的语音到语音框架,用于保持身份的Zero-Shot语音编辑作者:Philip Anastassiou,Zhenyu Tang,Kainan Peng,Dongya Jia,Jiaxin Li,Ming Tu,Yuping Wang,Yuxuan Wang,Mingbo Ma摘要:我们提出了VoiceShop,一种新颖的语音到语音的框架,可以修改语音的多个属性,如年龄,性别,口音和语音风格,在一个单一的向前传递,同时保留输入扬声器的音色。以前的作品已被限制到专门的模型,只能单独编辑这些属性,并遭受以下缺陷:转换效果的幅度是弱的,没有zero-shot能力的分布扬声器,或合成的输出表现出音色泄漏,改变扬声器的感知身份。我们的工作在一个简单的模块化框架中提出了这些问题的解决方案,该框架基于条件扩散骨干模型,具有可选的标准化基于流和序列到序列的扬声器属性编辑模块,其组件可以在推理过程中组合或删除,以满足各种任务,而无需额外的模型微调。音频样本可在https://voiceshopai.github.io上获得摘要:We present VoiceShop, a novel speech-to-speech framework that can modify multiple attributes of speech, such as age, gender, accent, and speech style, in a single forward pass while preserving the input speaker's timbre. Previous works have been constrained to specialized models that can only edit these attributes individually and suffer from the following pitfalls: the magnitude of the conversion effect is weak, there is no zero-shot capability for out-of-distribution speakers, or the synthesized outputs exhibit timbre leakage which changes the speaker's perceived identity. Our work proposes solutions for each of these issues in a simple modular framework based on a conditional diffusion backbone model with optional normalizing flow-based and sequence-to-sequence speaker attribute-editing modules, whose components can be combined or removed during inference to meet a wide array of tasks without additional model finetuning. Audio samples are available at https://voiceshopai.github.io