今日论文合集:cs.SD语音19篇,eess.AS音频处理21篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Sequence-to-Sequence Multi-Modal Speech In-Painting
标题: 绘画中的序列到序列多模式语音
作者:Mahsa Kadkhodaei Elyaderani,Shahram Shirani
链接:点击下载PDF文件
摘要:语音修复是使用可靠的上下文信息重新生成丢失的音频内容的任务。尽管最近在音频修复的多模态感知方面进行了各种研究,但仍然需要在语音修复中有效地注入视觉和听觉信息。在本文中,我们介绍了一种新的序列到序列模型,利用视觉信息,通过编码器-解码器架构内画音频信号。编码器在面部记录中扮演唇读器的角色,解码器将编码器的输出以及失真的音频频谱图恢复为原始语音。我们的模型优于一个只音频的语音在绘画模型,并与最近的多模态语音在画家的语音质量和可懂度指标的失真300毫秒至1500毫秒的持续时间,这证明了引入的多模态语音在绘画的有效性。摘要:Speech in-painting is the task of regenerating missing audio contents using reliable context information. Despite various recent studies in multi-modal perception of audio in-painting, there is still a need for an effective infusion of visual and auditory information in speech in-painting. In this paper, we introduce a novel sequence-to-sequence model that leverages the visual information to in-paint audio signals via an encoder-decoder architecture. The encoder plays the role of a lip-reader for facial recordings and the decoder takes both encoder outputs as well as the distorted audio spectrograms to restore the original speech. Our model outperforms an audio-only speech in-painting model and has comparable results with a recent multi-modal speech in-painter in terms of speech quality and intelligibility metrics for distortions of 300 ms to 1500 ms duration, which proves the effectiveness of the introduced multi-modality in speech in-painting.

【2】 animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
标题: animal2vec和MeerKAT:用于罕见事件原始音频输入的自我监督Transformer和用于生物声学的大规模参考数据集
作者:Julian C. Schäfer-Zimmermann,Vlad Demartsev,Baptiste Averly,Kiran Dhanjal-Adams,Mathieu Duteil,Gabriella Gall,Marius Faiß,Lily Johnson-Ulrich,Dan Stowell,Marta B. Manser,Marie A. Roch,Ariana Strandburg-Peshkin
备注:Code available at: this https URL | Dataset available at: this https URL
链接:点击下载PDF文件
摘要:生物声学研究为动物的行为、生态和保护提供了宝贵的见解。大多数生物声学数据集由长时间的记录组成,其中感兴趣的事件(如发声)非常罕见。分析这些数据集对研究人员提出了巨大的挑战,深度学习技术已经成为标准方法。他们的适应仍然具有挑战性,重点是为计算机视觉设计的模型,其中音频波形被设计成用于训练和推理的频谱表示。我们从两个方面改进了生物声学深度学习的现状:首先,我们提出了animal 2 vec框架:一个完全可解释的Transformer模型和为稀疏和不平衡的生物声学数据量身定制的自监督训练方案。其次,我们公开发布了MeerKAT:Meerkat Kalahari Audio Transcripts,这是一个大型数据集,包含通过部署在自由放养的猫鼬上的生物记录器收集的音频,长度超过1068 h,其中184 h有12个时间分辨的发声类型类,每个类的分辨率为ms,使其成为陆地哺乳动物中最大的公开可用标记数据集。此外,我们将animal 2 vec与NIPS 4 Bplus鸟鸣数据集进行了基准测试。我们报告了这两个数据集的最新结果,并评估了标记训练数据的animal 2 vec的Few-Shot能力。最后,我们进行消融研究,以突出我们的架构和香草Transformer基线人类产生的声音之间的差异。animal 2 vec允许研究人员对大量稀疏的生物声学数据进行分类,即使可用的地面实况信息很少。此外,MeerKAT数据集是第一个大规模的毫秒分辨率语料库,用于在预训练 微调范式中对生物声学模型进行基准测试。我们相信这为生物声学的新参考点奠定了基础。摘要:Bioacoustic research provides invaluable insights into the behavior, ecology, and conservation of animals. Most bioacoustic datasets consist of long recordings where events of interest, such as vocalizations, are exceedingly rare. Analyzing these datasets poses a monumental challenge to researchers, where deep learning techniques have emerged as a standard method. Their adaptation remains challenging, focusing on models conceived for computer vision, where the audio waveforms are engineered into spectrographic representations for training and inference. We improve the current state of deep learning in bioacoustics in two ways: First, we present the animal2vec framework: a fully interpretable transformer model and self-supervised training scheme tailored for sparse and unbalanced bioacoustic data. Second, we openly publish MeerKAT: Meerkat Kalahari Audio Transcripts, a large-scale dataset containing audio collected via biologgers deployed on free-ranging meerkats with a length of over 1068h, of which 184h have twelve time-resolved vocalization-type classes, each with ms-resolution, making it the largest publicly-available labeled dataset on terrestrial mammals. Further, we benchmark animal2vec against the NIPS4Bplus birdsong dataset. We report new state-of-the-art results on both datasets and evaluate the few-shot capabilities of animal2vec of labeled training data. Finally, we perform ablation studies to highlight the differences between our architecture and a vanilla transformer baseline for human-produced sounds. animal2vec allows researchers to classify massive amounts of sparse bioacoustic data even with little ground truth information available. In addition, the MeerKAT dataset is the first large-scale, millisecond-resolution corpus for benchmarking bioacoustic models in the pretrain finetune paradigm. We believe this sets the stage for a new reference point for bioacoustics.

【3】 Searching For Music Mixing Graphs: A Pruning Approach
标题: 搜索音乐混音图:修剪方法
作者:Sungho Lee,Marco A. Martínez-Ramírez,Wei-Hsiang Liao,Stefan Uhlich,Giorgio Fabbro,Kyogu Lee,Yuki Mitsufuji
备注:Accepted to DAFx 2024
链接:点击下载PDF文件
摘要:音乐混合是合成的-专家结合多个音频处理器,以实现从干源轨道的凝聚力混合。我们提出了一种方法来从输入和输出音频逆向工程这个过程。首先,我们创建一个混合控制台,将所有可用的处理器应用到每个链。然后,在初始控制台参数优化之后,我们在移除冗余处理器和微调之间交替。我们实现这一目标,通过差异化的实现处理器和修剪。因此,我们找到了一个稀疏的混合图,实现了几乎相同的匹配质量的完整的混合控制台。我们将此过程应用于来自各种数据集的干混对,并收集也可用于训练音乐混合应用的神经网络的图形。摘要:Music mixing is compositional -- experts combine multiple audio processors to achieve a cohesive mix from dry source tracks. We propose a method to reverse engineer this process from the input and output audio. First, we create a mixing console that applies all available processors to every chain. Then, after the initial console parameter optimization, we alternate between removing redundant processors and fine-tuning. We achieve this through differentiable implementation of both processors and pruning. Consequently, we find a sparse mixing graph that achieves nearly identical matching quality of the full mixing console. We apply this procedure to dry-mix pairs from various datasets and collect graphs that also can be used to train neural networks for music mixing applications.

【4】 Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
标题: 具有高效分层Transformer的生成预训练语音语言模型
作者:Yongxin Zhu,Dan Su,Liqiang He,Linli Xu,Dong Yu
备注:Accept in ACL2024-main
链接:点击下载PDF文件
摘要:虽然语音语言模型的最新进展已经取得了重大进展,但它们在神经音频编解码器的长声学序列建模方面面临着巨大的挑战。在本文中,我们介绍了 textbf{G}生成 textbf{P}再训练 textbf{S}语音 textbf{T} transformer(GPST),一个分层的Transformer设计有效的语音语言建模。GPST将音频波形量化为两种不同类型的离散语音表示,并将其集成在分层Transformer架构中,从而实现统一的一级生成过程并增强Hi-Res音频生成能力。通过以端到端的无监督方式对大型语音语料库进行训练,GPST可以生成具有不同说话人身份的语法一致的语音。在简短的3秒提示下,GPST可以产生自然和连贯的个性化语音,展示了上下文学习能力。此外,我们的方法可以很容易地扩展到口语跨语言的语音生成,通过将多语言的语义令牌和通用的声学令牌。实验结果表明,GPST显着优于现有的语音语言模型的词错误率,语音质量,说话人相似度。查看 url{https: Github.io GPST demo}获取演示示例。摘要:While recent advancements in speech language models have achieved significant progress, they face remarkable challenges in modeling the long acoustic sequences of neural audio codecs. In this paper, we introduce textbf{G}enerative textbf{P}re-trained textbf{S}peech textbf{T}ransformer (GPST), a hierarchical transformer designed for efficient speech language modeling. GPST quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a hierarchical transformer architecture, allowing for a unified one-stage generation process and enhancing Hi-Res audio generation capabilities. By training on large corpora of speeches in an end-to-end unsupervised manner, GPST can generate syntactically consistent speech with diverse speaker identities. Given a brief 3-second prompt, GPST can produce natural and coherent personalized speech, demonstrating in-context learning abilities. Moreover, our approach can be easily extended to spoken cross-lingual speech generation by incorporating multi-lingual semantic tokens and universal acoustic tokens. Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality, and speaker similarity. See url{https: youngsheen.github.io GPST demo} for demo samples.

【5】 Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach
标题: 鲁棒的多模式语音绘画:序列到序列方法
作者:Mahsa Kadkhodaei Elyaderani,Shahram Shirani
链接:点击下载PDF文件
摘要:从上下文重建语音音频的缺失部分的过程被称为语音修复。人类对语音的感知本质上是多模态的,涉及音频和视觉(AV)提示。在本文中,我们介绍和研究了序列到序列(seq2seq)的语音修复模型,结合AV功能。我们的方法将AV语音修复技术扩展到音频和视觉数据可能共同损坏的场景。为了实现这一目标,我们采用了多模态训练范式,该范式可以在涉及声学和视觉失真的各种条件下提高模型的鲁棒性。这使得我们的失真感知模型成为现实世界挑战性环境的合理解决方案。我们将我们的方法与现有的基于变换器和基于递归神经网络的模型进行比较,这些模型试图重建从几毫秒到一秒多的缺失语音间隙。我们的实验结果表明,我们的新seq2seq架构优于最先进的Transformer解决方案的38.8%,提高语音质量和7.14%,提高语音清晰度。我们利用一个多任务学习框架,同时执行唇读(将视频组件转录为文本),同时重建相关语音的缺失部分。摘要:The process of reconstructing missing parts of speech audio from context is called speech in-painting. Human perception of speech is inherently multi-modal, involving both audio and visual (AV) cues. In this paper, we introduce and study a sequence-to-sequence (seq2seq) speech in-painting model that incorporates AV features. Our approach extends AV speech in-painting techniques to scenarios where both audio and visual data may be jointly corrupted. To achieve this, we employ a multi-modal training paradigm that boosts the robustness of our model across various conditions involving acoustic and visual distortions. This makes our distortion-aware model a plausible solution for real-world challenging environments. We compare our method with existing transformer-based and recurrent neural network-based models, which attempt to reconstruct missing speech gaps ranging from a few milliseconds to over a second. Our experimental results demonstrate that our novel seq2seq architecture outperforms the state-of-the-art transformer solution by 38.8% in terms of enhancing speech quality and 7.14% in terms of improving speech intelligibility. We exploit a multi-task learning framework that simultaneously performs lip-reading (transcribing video components to text) while reconstructing missing parts of the associated speech.

【6】 YODAS: Youtube-Oriented Dataset for Audio and Speech
标题: YODAS:面向YouTube的音频和语音数据集
作者:Xinjian Li,Shinnosuke Takamichi,Takaaki Saeki,William Chen,Sayaka Shiota,Shinji Watanabe
备注:ASRU 2023
链接:点击下载PDF文件
摘要:在这项研究中,我们介绍了YODAS(面向YouTube的音频和语音数据集),这是一个大规模的多语言数据集,目前包含超过50万小时的语音数据,来自100多种语言,来自标记和未标记的YouTube语音数据集。标记的子集,包括手动或自动字幕,有助于监督模型训练。相反,未标记的子集适合于自监督学习应用。YODAS是其规模的第一个公开可用的数据集,它是根据知识共享许可证分发的。我们介绍了收集方法用于YODAS,这有助于大规模的语音数据集的建设。随后,我们对数据集中包含的语音、文本进行了全面的分析。最后,我们描述了前15种语言的语音识别基线。摘要:In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled YouTube speech datasets. The labeled subsets, including manual or automatic subtitles, facilitate supervised model training. Conversely, the unlabeled subsets are apt for self-supervised learning applications. YODAS is distinctive as the first publicly available dataset of its scale, and it is distributed under a Creative Commons license. We introduce the collection methodology utilized for YODAS, which contributes to the large-scale speech dataset construction. Subsequently, we provide a comprehensive analysis of speech, text contained within the dataset. Finally, we describe the speech recognition baselines over the top-15 languages.

【7】 Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
标题: 具有参数化和非参数化CNN的原始波声学模型的语音误差分析
作者:Erfan Loweimi,Andrea Carmantini,Peter Bell,Steve Renals,Zoran Cvetkovic
备注:5 pages, 6 figures, 3 tables
链接:点击下载PDF文件
摘要:在本文中,我们分析了错误模式的原始波形声学模型在TIMIT的电话识别任务。我们的分析超越了传统的电话错误率(PER)指标。我们将音素分为三组:{塞擦音,双元音,摩擦音,鼻音,爆破音,半元音,元音,沉默},{辅音,元音+,沉默}和{浊音,清音,沉默},并计算每个类别中每个广义语音类的PER。我们还构建了一个混淆矩阵,为每个类别使用的替代错误和比较的混淆模式与过滤器组和Wav 2 vec 2.0系统。我们的原始波形声学模型由参数(Sinc 2Net)或非参数CNN和双向LSTM组成,在TIMIT Dev Test集上实现了低至13.7% 15.2%的PER,优于文献中原始波形模型的PER。我们还研究了迁移学习对语音错误模式和混淆矩阵的影响。它将开发 测试集的PER降低到11.8% 13.7%。摘要:In this paper, we analyse the error patterns of the raw waveform acoustic models in TIMIT's phone recognition task. Our analysis goes beyond the conventional phone error rate (PER) metric. We categorise the phones into three groups: {affricate, diphthong, fricative, nasal, plosive, semi-vowel, vowel, silence}, {consonant, vowel+, silence}, and {voiced, unvoiced, silence} and, compute the PER for each broad phonetic class in each category. We also construct a confusion matrix for each category using the substitution errors and compare the confusion patterns with those of the Filterbank and Wav2vec 2.0 systems. Our raw waveform acoustic models consists of parametric (Sinc2Net) or non-parametric CNNs and Bidirectional LSTMs, achieving down to 13.7% 15.2% PERs on TIMIT Dev Test sets, outperforming reported PERs for raw waveform models in the literature. We also investigate the impact of transfer learning from WSJ on the phonetic error patterns and confusion matrices. It reduces the PER to 11.8% 13.7% on the Dev Test sets.

【8】 Enhanced Classification of Heart Sounds Using Mel Frequency Cepstral Coefficients: A Comparative Study of Single and Ensemble Classifier Strategies
标题: 使用Mel频率倒谱系数增强心形分类:单一分类器和集合分类器策略的比较研究
作者:Amir Masoud Rahmani,Amir Haider,Parisa Khoshvaght,Mohammad Adeli,Entesar Gemeay,Yazeed Alkhrijah,Mokhtar Mohammadi,Mehdi Hosseinzadeh
链接:点击下载PDF文件
摘要:本文探讨了梅尔频率倒谱系数(MFCC)检测异常心音图使用两种分类策略的有效性:一个单一的分类器和一个集成分类器的方法。心音图分为S1、收缩期、S2和舒张期,每段估计13个MFCC,每拍产生52个MFCC。在单分类器策略中,将来自九个连续搏动的MFCC平均以分类心音图。相反,集合分类器策略采用9个分类器来单独评估心跳是否正常或异常,总体分类基于多数投票。这两种方法都在公开的心音图数据库上进行了测试。结果表明,集成分类器的策略实现了更高的准确性相比,单分类器的方法,建立MFCC更有效的比其他功能,包括时间,时间频率和统计特征,在类似的研究中评估。摘要:This paper explores the efficacy of Mel Frequency Cepstral Coefficients (MFCCs) in detecting abnormal phonocardiograms using two classification strategies: a single-classifier and an ensemble-classifier approach. Phonocardiograms were segmented into S1, systole, S2, and diastole intervals, with thirteen MFCCs estimated from each segment, yielding 52 MFCCs per beat. In the single-classifier strategy, the MFCCs from nine consecutive beats were averaged to classify phonocardiograms. Conversely, the ensemble-classifier strategy employed nine classifiers to individually assess beats as normal or abnormal, with the overall classification based on the majority vote. Both methods were tested on a publicly available phonocardiogram database. Results demonstrated that the ensemble-classifier strategy achieved higher accuracy compared to the single-classifier approach, establishing MFCCs as more effective than other features, including time, time-frequency, and statistical features, evaluated in similar studies.

【9】 Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
标题: 利用人类反馈增强Zero-Shot文本到语音合成
作者:Chen Chen,Yuchen Hu,Wen Wu,Helin Wang,Eng Siong Chng,Chao Zhang
备注:19 pages, Preprint
链接:点击下载PDF文件
摘要:近年来,文本到语音(TTS)技术取得了令人印象深刻的进步,特别是在大规模训练数据集方面,展示了人类水平的语音质量和对看不见的扬声器的令人印象深刻的zero-shot能力。然而,尽管人类的主观评价,如平均意见分数(MOS),仍然是评估合成语音质量的黄金标准,即使是最先进的TTS方法也将人类反馈与导致不匹配的训练目标和评价指标的训练隔离开来。在这项工作中,我们研究了一个新的主题,将主观的人的评价到TTS训练循环。受最近成功的人类反馈强化学习的启发,我们提出了一个全面的采样-注释-学习框架,专为TTS优化,即不确定性感知优化(UNO)。具体而言,UNO消除了奖励模型或偏好数据的需要,通过直接最大化语音生成的效用,同时考虑主观人类语音感知和评估中固有的可变性的不确定性。主观和客观评价的实验结果表明,UNO显着提高了zero-shot的TTS模型的MOS,字错误率,和说话人相似度方面的性能。此外,我们提出了一个显着的能力UNO,它可以适应所需的说话风格的情感TTS无缝和灵活。摘要:In recent years, text-to-speech (TTS) technology has witnessed impressive advancements, particularly with large-scale training datasets, showcasing human-level speech quality and impressive zero-shot capabilities on unseen speakers. However, despite human subjective evaluations, such as the mean opinion score (MOS), remaining the gold standard for assessing the quality of synthetic speech, even state-of-the-art TTS approaches have kept human feedback isolated from training that resulted in mismatched training objectives and evaluation metrics. In this work, we investigate a novel topic of integrating subjective human evaluation into the TTS training loop. Inspired by the recent success of reinforcement learning from human feedback, we propose a comprehensive sampling-annotating-learning framework tailored to TTS optimization, namely uncertainty-aware optimization (UNO). Specifically, UNO eliminates the need for a reward model or preference data by directly maximizing the utility of speech generations while considering the uncertainty that lies in the inherent variability in subjective human speech perception and evaluations. Experimental results of both subjective and objective evaluations demonstrate that UNO considerably improves the zero-shot performance of TTS models in terms of MOS, word error rate, and speaker similarity. Additionally, we present a remarkable ability of UNO that it can adapt to the desired speaking style in emotional TTS seamlessly and flexibly.

【10】 Intelligent Text-Conditioned Music Generation
标题: 智能文本条件音乐生成
作者:Zhouyao Xie,Nikhil Yadala,Xinyi Chen,Jing Xi Liu
链接:点击下载PDF文件
摘要:CLIP(对比图像预训练)是一种多模态神经网络,在(文本,图像)对上训练,以预测给定图像的最相关文本标题。它已被广泛用于图像生成,将其输出与生成模型(如VQGAN)连接,最著名的例子是OpenAI的DALLE-2。在这个项目中,我们采用类似的方法来弥合自然语言和音乐之间的差距。我们的模型分为两个步骤:首先,我们训练一个CLIP类模型对文本和音乐对对比损失,以对齐一段音乐与其最可能的文本标题。然后,我们将对齐模型与音乐解码器相结合来生成音乐。据我们所知,这是第一次尝试文本条件下的深层音乐生成。我们的实验表明,它是可能的训练的文本音乐对齐模型使用对比度损失和训练解码器生成音乐从文本提示。摘要:CLIP (Contrastive Language-Image Pre-Training) is a multimodal neural network trained on (text, image) pairs to predict the most relevant text caption given an image. It has been used extensively in image generation by connecting its output with a generative model such as VQGAN, with the most notable example being OpenAI's DALLE-2. In this project, we apply a similar approach to bridge the gap between natural language and music. Our model is split into two steps: first, we train a CLIP-like model on pairs of text and music over contrastive loss to align a piece of music with its most probable text caption. Then, we combine the alignment model with a music decoder to generate music. To the best of our knowledge, this is the first attempt at text-conditioned deep music generation. Our experiments show that it is possible to train the text-music alignment model using contrastive loss and train a decoder to generate music from text prompts.

【11】 Recent Advances in End-to-End Simultaneous Speech Translation
标题: 端到端同步语音翻译的最新进展
作者:Xiaoqian Liu,Guoqiang Hu,Yangfan Du,Erfeng He,YingFeng Luo,Chen Xu,Tong Xiao,Jingbo Zhu
链接:点击下载PDF文件
摘要:同步语音翻译(SimulST)是一项要求很高的任务,涉及在实时生成翻译的同时连续处理语音输入。本文提供了一个全面的概述,在SimulST研究的最新发展,重点是四个主要的挑战。首先,与处理冗长且连续的语音流相关联的复杂性构成了重大障碍。其次,由于需要立即翻译输出,满足实时要求存在固有的困难。第三,在翻译质量和延迟限制之间取得平衡仍然是一个关键挑战。最后,注释数据的稀缺性给任务增加了另一层复杂性。通过我们对这些挑战的探索和提出的解决方案,我们的目标是为SimulST研究的当前前景提供有价值的见解,并为未来的探索提出有前途的方向。摘要:Simultaneous speech translation (SimulST) is a demanding task that involves generating translations in real-time while continuously processing speech input. This paper offers a comprehensive overview of the recent developments in SimulST research, focusing on four major challenges. Firstly, the complexities associated with processing lengthy and continuous speech streams pose significant hurdles. Secondly, satisfying real-time requirements presents inherent difficulties due to the need for immediate translation output. Thirdly, striking a balance between translation quality and latency constraints remains a critical challenge. Finally, the scarcity of annotated data adds another layer of complexity to the task. Through our exploration of these challenges and the proposed solutions, we aim to provide valuable insights into the current landscape of SimulST research and suggest promising directions for future exploration.

【12】 Frieren: Efficient Video-to-Audio Generation with Rectified Flow Matching
标题: Frieren:具有纠正流匹配的高效视频到音频生成
作者:Yongqi Wang,Wenxiang Guo,Rongjie Huang,Jiawei Huang,Zehan Wang,Fuming You,Ruiqi Li,Zhou Zhao
链接:点击下载PDF文件
摘要:视频到音频(Video-to-Audio,V2 A)生成的目标是从无声视频中合成内容匹配的音频,如何构建具有高生成质量、高效率和视听时间同步的V2 A模型仍然是一个挑战。我们提出了Frieren,一个基于整流匹配的V2 A模型。Frieren将条件传输向量场从噪声回归到具有直线路径的潜在频谱图,并通过求解ODE进行采样,在音频质量方面优于自回归和基于分数的模型。通过采用基于前馈Transformer的非自回归向量场估计器和具有强时间对齐的通道级跨模态特征融合,我们的模型生成与输入视频高度同步的音频。此外,通过回流和引导向量场的一步蒸馏,我们的模型可以在几个,甚至只有一个采样步骤生成像样的音频。实验表明,Frieren在VGGSound上的生成质量和时间对齐方面都达到了最先进的性能,对齐准确率达到97.22%,并且在基于强扩散的基线上,初始分数提高了6.2%。音频样本可在http: frieren-v2a.github.io上获得。摘要:Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a V2A model based on rectified flow matching. Frieren regresses the conditional transport vector field from noise to spectrogram latent with straight paths and conducts sampling by solving ODE, outperforming autoregressive and score-based models in terms of audio quality. By employing a non-autoregressive vector field estimator based on a feed-forward transformer and channel-level cross-modal feature fusion with strong temporal alignment, our model generates audio that is highly synchronized with the input video. Furthermore, through reflow and one-step distillation with guided vector field, our model can generate decent audio in a few, or even only one sampling step. Experiments indicate that Frieren achieves state-of-the-art performance in both generation quality and temporal alignment on VGGSound, with alignment accuracy reaching 97.22%, and 6.2% improvement in inception score over the strong diffusion-based baseline. Audio samples are available at http: frieren-v2a.github.io .

【13】 Creative Text-to-Audio Generation via Synthesizer Programming
标题: 通过合成器编程创造性的文本到音频生成
作者:Manuel Cherep,Nikhil Singh,Jessica Shand
备注:Accepted to ICML 2024
链接:点击下载PDF文件
摘要:神经音频合成方法现在允许用自然语言指定想法。然而,这些方法产生的结果不容易调整,因为它们基于大的潜在空间和高达数十亿的不可解释的参数。我们提出了一种文本到音频生成方法,利用虚拟模块化的声音合成器,只有78个参数。合成器由于其灵活性和直观的控制,长期以来一直被熟练的声音设计师用于音乐和电影等媒体。我们的方法,CTAG,迭代更新合成器的参数,以产生高质量的文本提示,可以很容易地检查和调整音频渲染。以这种方式产生的声音也更加抽象,在细粒度的声学细节上捕捉基本的概念特征,类似于简单的草图如何生动地传达视觉概念。我们的研究结果显示了CTAG如何产生独特的声音,被认为是艺术的,但与最近的神经音频合成模型相似,将其定位为一个有价值的补充工具。摘要:Neural audio synthesis methods now allow specifying ideas in natural language. However, these methods produce results that cannot be easily tweaked, as they are based on large latent spaces and up to billions of uninterpretable parameters. We propose a text-to-audio generation method that leverages a virtual modular sound synthesizer with only 78 parameters. Synthesizers have long been used by skilled sound designers for media like music and film due to their flexibility and intuitive controls. Our method, CTAG, iteratively updates a synthesizer's parameters to produce high-quality audio renderings of text prompts that can be easily inspected and tweaked. Sounds produced this way are also more abstract, capturing essential conceptual features over fine-grained acoustic details, akin to how simple sketches can vividly convey visual concepts. Our results show how CTAG produces sounds that are distinctive, perceived as artistic, and yet similarly identifiable to recent neural audio synthesis models, positioning it as a valuable and complementary tool.

【14】 A Survey of Deep Learning Audio Generation Methods
标题: 深度学习音频生成方法综述
作者:Matej Božić,Marko Horvat
备注:14 pages, 2 figures
链接:点击下载PDF文件
摘要:本文介绍了用于音频生成的深度学习模型开发的三个不同方面的典型技术。在文章的第一部分中,我们从基本的音频波形开始解释音频表示。然后,我们进展到频域,重点是人类听觉的属性,最后介绍一个相对较新的发展。本文的主要部分重点解释基本和扩展的深度学习架构变体,以及它们在音频生成领域的实际应用。解决了以下架构:1)自动编码器2)生成对抗网络3)规范化流4)Transformer网络5)扩散模型。最后,我们将研究音频生成中常用的四种不同的评估指标。本文旨在为该领域的新手读者和初学者提供一个全面的了解音频生成方法的当前技术水平以及可以为未来研究探索的相关研究。摘要:This article presents a review of typical techniques used in three distinct aspects of deep learning model development for audio generation. In the first part of the article, we provide an explanation of audio representations, beginning with the fundamental audio waveform. We then progress to the frequency domain, with an emphasis on the attributes of human hearing, and finally introduce a relatively recent development. The main part of the article focuses on explaining basic and extended deep learning architecture variants, along with their practical applications in the field of audio generation. The following architectures are addressed: 1) Autoencoders 2) Generative adversarial networks 3) Normalizing flows 4) Transformer networks 5) Diffusion models. Lastly, we will examine four distinct evaluation metrics that are commonly employed in audio generation. This article aims to offer novice readers and beginners in the field a comprehensive understanding of the current state of the art in audio generation methods as well as relevant studies that can be explored for future research.

【15】 Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
标题: 多语言韵律迁移:比较监督学习和迁移学习
作者:Arnav Goel,Medha Hira,Anubha Gupta
备注:7 pages, Accepted to ICLR 2024 - Tiny Track
链接:点击下载PDF文件
摘要:语音合成系统中的韵律转换领域正在迅速发展。这项研究的重点是评估学习方法,使预训练的单语文本到语音(TTS)模型适应多语言条件,即,监督微调(SFT)和迁移学习(TL)。这种比较使用了三个不同的指标:平均意见评分(MOS),识别准确率(RA)和梅尔倒谱系数失真(MCD)。结果表明,与SFT相比,TL导致显著增强的性能,平均MOS高出1.53点,RA增加37.5%,MCD提高约7.8点。这些发现有助于为低资源语言建立TTS模型。摘要:The field of prosody transfer in speech synthesis systems is rapidly advancing. This research is focused on evaluating learning methods for adapting pre-trained monolingual text-to-speech (TTS) models to multilingual conditions, i.e., Supervised Fine-Tuning (SFT) and Transfer Learning (TL). This comparison utilizes three distinct metrics: Mean Opinion Score (MOS), Recognition Accuracy (RA), and Mel Cepstral Distortion (MCD). Results demonstrate that, in comparison to SFT, TL leads to significantly enhanced performance, with an average MOS higher by 1.53 points, a 37.5% increase in RA, and approximately a 7.8-point improvement in MCD. These findings are instrumental in helping build TTS models for low-resource languages.

【16】 CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
标题: CrossVoice:使用迁移学习的跨语言韵律保留Cascade-S2 ST
作者:Medha Hira,Arnav Goel,Anubha Gupta
备注:8 pages, Accepted at ICLR 2024 - Tiny Track
链接:点击下载PDF文件
摘要:本文介绍了CrossVoice,一种新型的基于级联的语音到语音翻译(S2 ST)系统,采用先进的ASR,MT和TTS技术,通过迁移学习实现跨语言的韵律保留。我们进行了全面的实验,比较CrossVoice与直接S2 ST系统,在基准数据集CVSS-T和IndicTTS上显示了Fisher Es-En,VoxPopuli Fr-En和韵律保留等任务的BLEU分数提高。CrossVoice合成的语音平均意见得分为3.75分(满分为4分),在基准测试中与人类语音不相上下,突出了基于级联的系统和迁移学习在多语言S2 ST中的有效性。摘要:This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.

【17】 ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
标题: Control Speech:通过脱钩编解码器实现同时Zero-Shot说话人克隆和Zero-Shot语言风格控制
作者:Shengpeng Ji,Jialong Zuo,Minghui Fang,Siqi Zheng,Qian Chen,Wen Wang,Ziyue Jiang,Hai Huang,Xize Cheng,Rongjie Huang,Zhou Zhao
链接:点击下载PDF文件
摘要:在本文中,我们提出了ControlSpeech,一个文本到语音(TTS)系统,能够完全克隆说话人的声音,使任意控制和调整说话风格,仅仅基于几秒钟的音频提示和一个简单的文本风格描述提示。现有的zero-shot TTS模型和可控TTS模型要么只能模仿说话者的语音而没有进一步的控制和调整能力,要么与说话者特定的语音生成无关。因此,ControlSpeech专注于一个更具挑战性的新任务--一个同时具有可控音色、内容和风格的TTS系统。ControlSpeech将语音提示、内容提示和风格提示作为输入,并利用双向注意和基于掩码的并行解码在离散解耦编解码器空间中捕获相应的编解码器表示。此外,我们发现了多对多映射方式下的文本风格可控性问题,并提出了风格混合语义密度(SMSD)模型来解决这个问题。SMSD模块基于高斯混合密度网络,增强风格语义信息的细粒度划分和采样能力,生成风格更加多样的语音。在实验方面,我们提供了一个可控的模型工具包称为ControlToolkit与一个新的风格可控的数据集,一些复制的基线模型,并提出了新的指标来评估控制能力和质量的生成音频在ControlSpeech。相关的消融研究验证了ControlSpeech中各个成分的必要性。我们希望ControlSpeech能够建立下一个可控语音合成的基础范例。相关代码和演示可在https: github.com jishengpeng ControlSpeech上获得。摘要:In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style, merely based on a few seconds of audio prompt and a simple textual style description prompt. Prior zero-shot TTS models and controllable TTS models either could only mimic the speaker's voice without further control and adjustment capabilities or were unrelated to speaker-specific voice generation. Therefore, ControlSpeech focuses on a more challenging new task-a TTS system with controllable timbre, content, and style at the same time. ControlSpeech takes speech prompts, content prompts, and style prompts as inputs and utilizes bidirectional attention and mask-based parallel decoding to capture corresponding codec representations in a discrete decoupling codec space. Moreover, we discovered the issue of text style controllability in a many-to-many mapping fashion and proposed the Style Mixture Semantic Density (SMSD) model to resolve this problem. SMSD module which is based on Gaussian mixture density networks, is designed to enhance the fine-grained partitioning and sampling capabilities of style semantic information and generate speech with more diverse styles. In terms of experiments, we make available a controllable model toolkit called ControlToolkit with a new style controllable dataset, some replicated baseline models and propose new metrics to evaluate both the control capability and the quality of generated audio in ControlSpeech. The relevant ablation studies validate the necessity of each component in ControlSpeech is necessary. We hope that ControlSpeech can establish the next foundation paradigm of controllable speech synthesis. The relevant code and demo are available at https: github.com jishengpeng ControlSpeech .

【18】 Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
标题: 使用多层VAE和对抗训练的文本到语音中的口音转换
作者:Jan Melechovsky,Ambuj Mehrish,Berrak Sisman,Dorien Herremans
备注:Under review
链接:点击下载PDF文件
摘要:随着全球化的迅速发展,建立具有包容性和代表性的语音技术的必要性怎么强调都不为过。口音是语音的一个重要方面,在构建包容性语音合成器时需要考虑。包容性语音技术旨在消除对特定群体的任何偏见,例如某些口音的人。我们注意到,最先进的文本到语音(TTS)系统目前可能不适合所有人,无论他们的背景如何,因为它们旨在生成高质量的语音而不关注口音。在本文中,我们提出了一种TTS模型,该模型利用具有对抗学习的多级变分自动编码器来解决TTS中的重音语音合成和转换,并展望未来更具包容性的系统。我们通过客观指标和主观听力测试来评估性能。结果显示,与基线相比,口音转换能力有所提高。摘要:With rapid globalization, the need to build inclusive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers. Inclusive speech technology aims to erase any biases towards specific groups, such as people of certain accent. We note that state-of-the-art Text-to-Speech (TTS) systems may currently not be suitable for all people, regardless of their background, as they are designed to generate high-quality voices without focusing on accent. In this paper, we propose a TTS model that utilizes a Multi-Level Variational Autoencoder with adversarial learning to address accented speech synthesis and conversion in TTS, with a vision for more inclusive systems in the future. We evaluate the performance through both objective metrics and subjective listening tests. The results show an improvement in accent conversion ability compared to the baseline.

【19】 AudioLCM: Text-to-Audio Generation with Latent Consistency Models
标题: AudioRCM:具有潜在一致性模型的文本到音频生成
作者:Huadai Liu,Rongjie Huang,Yang Liu,Hengyuan Cao,Jialei Wang,Xize Cheng,Siqi Zheng,Zhou Zhao
链接:点击下载PDF文件
摘要:潜在扩散模型(LDMs)的最新进展将其推向了各种生成任务的最前沿。然而,它们的迭代采样过程带来了巨大的计算负担,导致生成速度缓慢,并限制了它们在文本到音频生成部署中的应用。在这项工作中,我们引入AudioLCM,一种新的基于一致性的模型,专为高效和高质量的文本到音频生成。AudioLCM将一致性模型集成到生成过程中,通过从任何时间步的任何点到轨迹初始点的映射,促进快速推理。为了克服LDM中固有的收敛问题,减少样本迭代,我们提出了引导潜在一致性蒸馏与多步常微分方程(ODE)求解器。这一创新将时间表从数千步缩短到数十步,同时保持样品质量,从而实现快速收敛和高质量生成。此外,为了优化基于变压器的神经网络架构的性能,我们将LLaMA开创的先进技术集成到Transformers的基础框架中。该架构支持稳定高效的训练,确保文本到音频合成的稳健性能。文本到声音生成和文本到音乐合成任务的实验结果表明,AudioLCM只需要2次迭代来合成高保真音频,同时它保持了与使用数百个步骤的最先进模型竞争的样本质量。AudioLCM能够在单个NVIDIA 4090Ti GPU上实现比实时快333倍的采样速度,使生成模型实际上适用于文本到音频生成部署。我们广泛的初步分析表明,AudioLCM中的每个设计都是有效的。摘要:Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slow generation speeds and limiting their application in text-to-audio generation deployment. In this work, we introduce AudioLCM, a novel consistency-based model tailored for efficient and high-quality text-to-audio generation. AudioLCM integrates Consistency Models into the generation process, facilitating rapid inference through a mapping from any point at any time step to the trajectory's initial point. To overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver. This innovation shortens the time schedule from thousands to dozens of steps while maintaining sample quality, thereby achieving fast convergence and high-quality generation. Furthermore, to optimize the performance of transformer-based neural network architectures, we integrate the advanced techniques pioneered by LLaMA into the foundational framework of transformers. This architecture supports stable and efficient training, ensuring robust performance in text-to-audio synthesis. Experimental results on text-to-sound generation and text-to-music synthesis tasks demonstrate that AudioLCM needs only 2 iterations to synthesize high-fidelity audios, while it maintains sample quality competitive with state-of-the-art models using hundreds of steps. AudioLCM enables a sampling speed of 333x faster than real-time on a single NVIDIA 4090Ti GPU, making generative models practically applicable to text-to-audio generation deployment. Our extensive preliminary analysis shows that each design in AudioLCM is effective.


eess.AS音频处理
【1】 ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control With Decoupled Codec
标题: Control Speech:通过脱钩编解码器实现同时Zero-Shot说话人克隆和Zero-Shot语言风格控制
作者:Shengpeng Ji,Jialong Zuo,Minghui Fang,Siqi Zheng,Qian Chen,Wen Wang,Ziyue Jiang,Hai Huang,Xize Cheng,Rongjie Huang,Zhou Zhao
链接:点击下载PDF文件
摘要:在本文中,我们提出了ControlSpeech,一个文本到语音(TTS)系统,能够完全克隆说话人的声音,使任意控制和调整说话风格,仅仅基于几秒钟的音频提示和一个简单的文本风格描述提示。现有的zero-shot TTS模型和可控TTS模型要么只能模仿说话者的语音而没有进一步的控制和调整能力,要么与说话者特定的语音生成无关。因此,ControlSpeech专注于一个更具挑战性的新任务-同时具有可控音色,内容和风格的TTS系统。ControlSpeech将语音提示、内容提示和风格提示作为输入,并利用双向注意和基于掩码的并行解码在离散解耦编解码器空间中捕获相应的编解码器表示。此外,我们发现了多对多映射方式下的文本风格可控性问题,并提出了风格混合语义密度(SMSD)模型来解决这个问题。SMSD模块基于高斯混合密度网络,增强风格语义信息的细粒度划分和采样能力,生成风格更加多样的语音。在实验方面,我们提供了一个可控的模型工具包称为ControlToolkit与一个新的风格可控的数据集,一些复制的基线模型,并提出了新的指标来评估控制能力和质量的生成音频在ControlSpeech。相关的消融研究验证了ControlSpeech中各个成分的必要性。我们希望ControlSpeech能够建立下一个可控语音合成的基础范例。相关代码和演示可在https: github.com jishengpeng ControlSpeech上获得。摘要:In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style, merely based on a few seconds of audio prompt and a simple textual style description prompt. Prior zero-shot TTS models and controllable TTS models either could only mimic the speaker's voice without further control and adjustment capabilities or were unrelated to speaker-specific voice generation. Therefore, ControlSpeech focuses on a more challenging new task-a TTS system with controllable timbre, content, and style at the same time. ControlSpeech takes speech prompts, content prompts, and style prompts as inputs and utilizes bidirectional attention and mask-based parallel decoding to capture corresponding codec representations in a discrete decoupling codec space. Moreover, we discovered the issue of text style controllability in a many-to-many mapping fashion and proposed the Style Mixture Semantic Density (SMSD) model to resolve this problem. SMSD module which is based on Gaussian mixture density networks, is designed to enhance the fine-grained partitioning and sampling capabilities of style semantic information and generate speech with more diverse styles. In terms of experiments, we make available a controllable model toolkit called ControlToolkit with a new style controllable dataset, some replicated baseline models and propose new metrics to evaluate both the control capability and the quality of generated audio in ControlSpeech. The relevant ablation studies validate the necessity of each component in ControlSpeech is necessary. We hope that ControlSpeech can establish the next foundation paradigm of controllable speech synthesis. The relevant code and demo are available at https: github.com jishengpeng ControlSpeech .

【2】 Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
标题: 使用多层VAE和对抗训练的文本到语音中的口音转换
作者:Jan Melechovsky,Ambuj Mehrish,Berrak Sisman,Dorien Herremans
备注:Under review
链接:点击下载PDF文件
摘要:随着全球化的迅速发展,建立具有包容性和代表性的语音技术的必要性怎么强调都不为过。口音是语音的一个重要方面,在构建包容性语音合成器时需要考虑。包容性语音技术旨在消除对特定群体的任何偏见,例如某些口音的人。我们注意到,最先进的文本到语音(TTS)系统目前可能不适合所有人,无论他们的背景如何,因为它们旨在生成高质量的语音而不关注口音。在本文中,我们提出了一种TTS模型,该模型利用具有对抗学习的多级变分自动编码器来解决TTS中的重音语音合成和转换,并展望未来更具包容性的系统。我们通过客观指标和主观听力测试来评估性能。结果显示,与基线相比,口音转换能力有所提高。摘要:With rapid globalization, the need to build inclusive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers. Inclusive speech technology aims to erase any biases towards specific groups, such as people of certain accent. We note that state-of-the-art Text-to-Speech (TTS) systems may currently not be suitable for all people, regardless of their background, as they are designed to generate high-quality voices without focusing on accent. In this paper, we propose a TTS model that utilizes a Multi-Level Variational Autoencoder with adversarial learning to address accented speech synthesis and conversion in TTS, with a vision for more inclusive systems in the future. We evaluate the performance through both objective metrics and subjective listening tests. The results show an improvement in accent conversion ability compared to the baseline.

【3】 Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
标题: Wav2Promise:零镜头和Few-Shot学习中的LLM端到端语音提示生成和调整
作者:Keqi Deng,Guangzhi Sun,Philip C. Woodland
链接:点击下载PDF文件
摘要:Wav 2 Prompt允许语音输入和基于文本的大型语言模型(LLM)之间的直接集成。Wav 2 Prompt使用一个简单的训练过程,只使用与训练自动语音识别(ASR)模型相同的数据。训练后,Wav 2 Prompt从语音中学习连续表示,并将其用作LLM提示。为了避免先前工作中发现的任务过度拟合问题并保留LLM的紧急能力,Wav 2 Prompt将LLM令牌嵌入作为训练目标,并利用连续的集成和激发机制进行显式语音文本对齐。因此,Wav 2Shot-LLM组合可以应用于zero-shot口语任务,诸如语音翻译(ST)、语音理解(SLU)、语音问答(SQA)和基于口语查询的QA(SQQA)。结果表明,对于这些zero-shot任务,Wav 2 Prompt执行类似于ASR-LLM级联,并且优于最近的先前工作。如果在Few-Shot场景中有相对少量的任务特定配对数据可用,则可以对Wav 2 Wavelt-LLM组合进行端到端(E2 E)微调。然后,Wav 2 Wavt-LLM组合相对于用于上述任务的ASR-LLM级联产生大大改进的结果。例如,对于使用BLOOMZ-7 B1 LLM的英语-法语ST,Wav 2bet-LLM组合比ASR-LLM级联增加了8.5个BLEU点。摘要:Wav2Prompt is proposed which allows straightforward integration between spoken input and a text-based large language model (LLM). Wav2Prompt uses a simple training process with only the same data used to train an automatic speech recognition (ASR) model. After training, Wav2Prompt learns continuous representations from speech and uses them as LLM prompts. To avoid task over-fitting issues found in prior work and preserve the emergent abilities of LLMs, Wav2Prompt takes LLM token embeddings as the training targets and utilises a continuous integrate-and-fire mechanism for explicit speech-text alignment. Therefore, a Wav2Prompt-LLM combination can be applied to zero-shot spoken language tasks such as speech translation (ST), speech understanding (SLU), speech question answering (SQA) and spoken-query-based QA (SQQA). It is shown that for these zero-shot tasks, Wav2Prompt performs similarly to an ASR-LLM cascade and better than recent prior work. If relatively small amounts of task-specific paired data are available in few-shot scenarios, the Wav2Prompt-LLM combination can be end-to-end (E2E) fine-tuned. The Wav2Prompt-LLM combination then yields greatly improved results relative to an ASR-LLM cascade for the above tasks. For instance, for English-French ST with the BLOOMZ-7B1 LLM, a Wav2Prompt-LLM combination gave a 8.5 BLEU point increase over an ASR-LLM cascade.

【4】 Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
标题: 用于空间声音复制的视频中视听说话者定位
作者:Davide Berghi,Philip J. B. Jackson
链接:点击下载PDF文件
摘要:基于对象的音频制作需要为每个点源对象定义位置元数据,包括声音场景前景中的关键元素。在许多媒体制作用例中,摄像机和麦克风都被用来录制,而人类的声音通常是一个关键因素。在这项研究中,我们检测和定位视频中的活动扬声器,便于自动提取的位置元数据的讲话者相对于摄像机的参考系。与视觉模态的集成,本研究扩展了我们以前的研究只集中在基于音频的主动说话人检测和定位。我们的实验比较了传统的视听方法,利用单声道音频,我们以前的音频方法,利用麦克风阵列的多通道录音,和一种新的视听方法集成的视觉和多通道音频主动扬声器检测。我们发现这两种模式的作用是相辅相成的。多通道音频克服了视觉遮挡的问题,与单通道音频的视听方法相比,检测误差减少了两位数。多通道音频和视觉的结合进一步提高了空间准确性,导致Tragic Talkers数据集的F1分数增加了4个百分点。未来的研究将评估该模型在嘈杂和高混响环境中的鲁棒性,以及解决屏幕外扬声器的问题。摘要:Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are employed to make recordings, and the human voice is often a key element. In this research, we detect and locate the active speaker in the video, facilitating the automatic extraction of the positional metadata of the talker relative to the camera's reference frame. With the integration of the visual modality, this study expands upon our previous investigation focused solely on audio-based active speaker detection and localization. Our experiments compare conventional audio-visual approaches for active speaker detection that leverage monaural audio, our previous audio-only method that leverages multichannel recordings from a microphone array, and a novel audio-visual approach integrating vision and multichannel audio. We found the role of the two modalities to complement each other. Multichannel audio, overcoming the problem of visual occlusions, provides a double-digit reduction in detection error compared to audio-visual methods with single-channel audio. The combination of multichannel audio and vision further enhances spatial accuracy, leading to a four-percentage point increase in F1 score on the Tragic Talkers dataset. Future investigations will assess the robustness of the model in noisy and highly reverberant environments, as well as tackle the problem of off-screen speakers.

【5】 AudioLCM: Text-to-Audio Generation with Latent Consistency Models
标题: AudioRCM:具有潜在一致性模型的文本到音频生成
作者:Huadai Liu,Rongjie Huang,Yang Liu,Hengyuan Cao,Jialei Wang,Xize Cheng,Siqi Zheng,Zhou Zhao
链接:点击下载PDF文件
摘要:潜在扩散模型(LDMs)的最新进展将其推向了各种生成任务的最前沿。然而,它们的迭代采样过程带来了巨大的计算负担,导致生成速度缓慢,并限制了它们在文本到音频生成部署中的应用。在这项工作中,我们引入AudioLCM,一种新的基于一致性的模型,专为高效和高质量的文本到音频生成。AudioLCM将一致性模型集成到生成过程中,通过从任何时间步的任何点到轨迹初始点的映射,促进快速推理。为了克服LDM中固有的收敛问题,减少样本迭代,我们提出了引导潜在一致性蒸馏与多步常微分方程(ODE)求解器。这一创新将时间表从数千步缩短到数十步,同时保持样品质量,从而实现快速收敛和高质量生成。此外,为了优化基于变压器的神经网络架构的性能,我们将LLaMA开创的先进技术集成到Transformers的基础框架中。该架构支持稳定高效的训练,确保文本到音频合成的稳健性能。文本到声音生成和文本到音乐合成任务的实验结果表明,AudioLCM只需要2次迭代来合成高保真音频,同时它保持了与使用数百个步骤的最先进模型竞争的样本质量。AudioLCM能够在单个NVIDIA 4090Ti GPU上实现比实时快333倍的采样速度,使生成模型实际上适用于文本到音频生成部署。我们广泛的初步分析表明,AudioLCM中的每个设计都是有效的。摘要:Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slow generation speeds and limiting their application in text-to-audio generation deployment. In this work, we introduce AudioLCM, a novel consistency-based model tailored for efficient and high-quality text-to-audio generation. AudioLCM integrates Consistency Models into the generation process, facilitating rapid inference through a mapping from any point at any time step to the trajectory's initial point. To overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver. This innovation shortens the time schedule from thousands to dozens of steps while maintaining sample quality, thereby achieving fast convergence and high-quality generation. Furthermore, to optimize the performance of transformer-based neural network architectures, we integrate the advanced techniques pioneered by LLaMA into the foundational framework of transformers. This architecture supports stable and efficient training, ensuring robust performance in text-to-audio synthesis. Experimental results on text-to-sound generation and text-to-music synthesis tasks demonstrate that AudioLCM needs only 2 iterations to synthesize high-fidelity audios, while it maintains sample quality competitive with state-of-the-art models using hundreds of steps. AudioLCM enables a sampling speed of 333x faster than real-time on a single NVIDIA 4090Ti GPU, making generative models practically applicable to text-to-audio generation deployment. Our extensive preliminary analysis shows that each design in AudioLCM is effective.

【6】 Enabling ASR for Low-Resource Languages: A Comprehensive Dataset Creation Approach
标题: 为低资源语言启用SVR:全面的数据集创建方法
作者:Ara Yeroyan,Nikolay Karpov
备注:13 pages, 10 figures (including ablation studies), to be published in 2024 IEEE Spoken Language Technology Workshop. Additionally, the associated software package can be accessed at (this https URL) for practical applications and further development
链接:点击下载PDF文件
摘要:近年来,自动语音识别(ASR)系统有了显著的改进,特别是在具有大量转录语音数据的语言中。然而,ASR系统往往对资源较少的低资源语言表现不佳,例如少数民族和区域语言。这项研究介绍了一种新的管道,旨在从有声读物中生成ASR训练数据集,这些数据集通常具有与长达数小时的音频相关的单个转录本。由于音频片段的长度很长,这些有声读物的共同结构构成了独特的挑战,而最佳的ASR训练需要4到15秒的片段。为了解决这个问题,我们提出了一种方法,可以有效地将音频与其相应的文本对齐,并将其分割成适合ASR训练的长度。我们的方法简化了低资源语言的ASR系统的数据准备,并通过涉及亚美尼亚语的案例研究演示了其应用程序。我们的方法,这是“便携式”的许多低资源的语言,不仅缓解了数据稀缺的问题,但也提高了性能的ASR模型代表性不足的语言。摘要:In recent years, automatic speech recognition (ASR) systems have significantly improved, especially in languages with a vast amount of transcribed speech data. However, ASR systems tend to perform poorly for low-resource languages with fewer resources, such as minority and regional languages. This study introduces a novel pipeline designed to generate ASR training datasets from audiobooks, which typically feature a single transcript associated with hours-long audios. The common structure of these audiobooks poses a unique challenge due to the extensive length of audio segments, whereas optimal ASR training requires segments ranging from 4 to 15 seconds. To address this, we propose a method for effectively aligning audio with its corresponding text and segmenting it into lengths suitable for ASR training. Our approach simplifies data preparation for ASR systems in low-resource languages and demonstrates its application through a case study involving the Armenian language. Our method, which is "portable" to many low-resource languages, not only mitigates the issue of data scarcity but also enhances the performance of ASR models for underrepresented languages.

【7】 Sequence-to-Sequence Multi-Modal Speech In-Painting
标题: 绘画中的序列到序列多模式语音
作者:Mahsa Kadkhodaei Elyaderani,Shahram Shirani
链接:点击下载PDF文件
摘要:语音修复是使用可靠的上下文信息重新生成丢失的音频内容的任务。尽管最近在音频修复的多模态感知方面进行了各种研究,但仍然需要在语音修复中有效地注入视觉和听觉信息。在本文中,我们介绍了一种新的序列到序列模型,利用视觉信息,通过编码器-解码器架构内画音频信号。编码器在面部记录中扮演唇读器的角色,解码器将编码器的输出以及失真的音频频谱图恢复为原始语音。我们的模型优于一个只音频的语音在绘画模型,并与最近的多模态语音在画家的语音质量和可懂度指标的失真300毫秒至1500毫秒的持续时间,这证明了引入的多模态语音在绘画的有效性。摘要:Speech in-painting is the task of regenerating missing audio contents using reliable context information. Despite various recent studies in multi-modal perception of audio in-painting, there is still a need for an effective infusion of visual and auditory information in speech in-painting. In this paper, we introduce a novel sequence-to-sequence model that leverages the visual information to in-paint audio signals via an encoder-decoder architecture. The encoder plays the role of a lip-reader for facial recordings and the decoder takes both encoder outputs as well as the distorted audio spectrograms to restore the original speech. Our model outperforms an audio-only speech in-painting model and has comparable results with a recent multi-modal speech in-painter in terms of speech quality and intelligibility metrics for distortions of 300 ms to 1500 ms duration, which proves the effectiveness of the introduced multi-modality in speech in-painting.

【8】 animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
标题: animal2vec和MeerKAT:用于罕见事件原始音频输入的自我监督Transformer和用于生物声学的大规模参考数据集
作者:Julian C. Schäfer-Zimmermann,Vlad Demartsev,Baptiste Averly,Kiran Dhanjal-Adams,Mathieu Duteil,Gabriella Gall,Marius Faiß,Lily Johnson-Ulrich,Dan Stowell,Marta B. Manser,Marie A. Roch,Ariana Strandburg-Peshkin
备注:Code available at: this https URL | Dataset available at: this https URL
链接:点击下载PDF文件
摘要:生物声学研究为动物的行为、生态和保护提供了宝贵的见解。大多数生物声学数据集由长时间的记录组成,其中感兴趣的事件(如发声)非常罕见。分析这些数据集对研究人员提出了巨大的挑战,深度学习技术已经成为标准方法。他们的适应仍然具有挑战性,重点是为计算机视觉设计的模型,其中音频波形被设计成用于训练和推理的频谱表示。我们从两个方面改进了生物声学深度学习的现状:首先,我们提出了animal 2 vec框架:一个完全可解释的Transformer模型和为稀疏和不平衡的生物声学数据量身定制的自监督训练方案。其次,我们公开发布了MeerKAT:Meerkat Kalahari Audio Transcripts,这是一个大型数据集,包含通过部署在自由放养的猫鼬上的生物记录器收集的音频,长度超过1068 h,其中184 h有12个时间分辨的发声类型类,每个类的分辨率为ms,使其成为陆地哺乳动物中最大的公开可用标记数据集。此外,我们将animal 2 vec与NIPS 4 Bplus鸟鸣数据集进行了基准测试。我们报告了这两个数据集的最新结果,并评估了标记训练数据的animal 2 vec的Few-Shot能力。最后,我们进行消融研究,以突出我们的架构和香草Transformer基线人类产生的声音之间的差异。animal 2 vec允许研究人员对大量稀疏的生物声学数据进行分类,即使可用的地面实况信息很少。此外,MeerKAT数据集是第一个大规模的毫秒分辨率语料库,用于在预训练 微调范式中对生物声学模型进行基准测试。我们相信这为生物声学的新参考点奠定了基础。摘要:Bioacoustic research provides invaluable insights into the behavior, ecology, and conservation of animals. Most bioacoustic datasets consist of long recordings where events of interest, such as vocalizations, are exceedingly rare. Analyzing these datasets poses a monumental challenge to researchers, where deep learning techniques have emerged as a standard method. Their adaptation remains challenging, focusing on models conceived for computer vision, where the audio waveforms are engineered into spectrographic representations for training and inference. We improve the current state of deep learning in bioacoustics in two ways: First, we present the animal2vec framework: a fully interpretable transformer model and self-supervised training scheme tailored for sparse and unbalanced bioacoustic data. Second, we openly publish MeerKAT: Meerkat Kalahari Audio Transcripts, a large-scale dataset containing audio collected via biologgers deployed on free-ranging meerkats with a length of over 1068h, of which 184h have twelve time-resolved vocalization-type classes, each with ms-resolution, making it the largest publicly-available labeled dataset on terrestrial mammals. Further, we benchmark animal2vec against the NIPS4Bplus birdsong dataset. We report new state-of-the-art results on both datasets and evaluate the few-shot capabilities of animal2vec of labeled training data. Finally, we perform ablation studies to highlight the differences between our architecture and a vanilla transformer baseline for human-produced sounds. animal2vec allows researchers to classify massive amounts of sparse bioacoustic data even with little ground truth information available. In addition, the MeerKAT dataset is the first large-scale, millisecond-resolution corpus for benchmarking bioacoustic models in the pretrain finetune paradigm. We believe this sets the stage for a new reference point for bioacoustics.

【9】 Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
标题: 具有高效分层Transformer的生成预训练语音语言模型
作者:Yongxin Zhu,Dan Su,Liqiang He,Linli Xu,Dong Yu
备注:Accept in ACL2024-main
链接:点击下载PDF文件
摘要:虽然语音语言模型的最新进展已经取得了重大进展,但它们在神经音频编解码器的长声学序列建模方面面临着巨大的挑战。在本文中,我们介绍了 textbf{G}生成 textbf{P}再训练 textbf{S}语音 textbf{T} transformer(GPST),一个分层的Transformer设计有效的语音语言建模。GPST将音频波形量化为两种不同类型的离散语音表示,并将其集成在分层Transformer架构中,从而实现统一的一级生成过程并增强Hi-Res音频生成能力。通过以端到端的无监督方式对大型语音语料库进行训练,GPST可以生成具有不同说话人身份的语法一致的语音。在简短的3秒提示下,GPST可以产生自然和连贯的个性化语音,展示了上下文学习能力。此外,我们的方法可以很容易地扩展到口语跨语言的语音生成,通过将多语言的语义令牌和通用的声学令牌。实验结果表明,GPST显着优于现有的语音语言模型的词错误率,语音质量,说话人相似度。查看 url{https: Github.io GPST demo}获取演示示例。摘要:While recent advancements in speech language models have achieved significant progress, they face remarkable challenges in modeling the long acoustic sequences of neural audio codecs. In this paper, we introduce textbf{G}enerative textbf{P}re-trained textbf{S}peech textbf{T}ransformer (GPST), a hierarchical transformer designed for efficient speech language modeling. GPST quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a hierarchical transformer architecture, allowing for a unified one-stage generation process and enhancing Hi-Res audio generation capabilities. By training on large corpora of speeches in an end-to-end unsupervised manner, GPST can generate syntactically consistent speech with diverse speaker identities. Given a brief 3-second prompt, GPST can produce natural and coherent personalized speech, demonstrating in-context learning abilities. Moreover, our approach can be easily extended to spoken cross-lingual speech generation by incorporating multi-lingual semantic tokens and universal acoustic tokens. Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality, and speaker similarity. See url{https: youngsheen.github.io GPST demo} for demo samples.

【10】 Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach
标题: 鲁棒的多模式语音绘画:序列到序列方法
作者:Mahsa Kadkhodaei Elyaderani,Shahram Shirani
链接:点击下载PDF文件
摘要:从上下文重建语音音频的缺失部分的过程被称为语音修复。人类对语音的感知本质上是多模态的,涉及音频和视觉(AV)提示。在本文中,我们介绍和研究了序列到序列(seq2seq)的语音修复模型,结合AV功能。我们的方法将AV语音修复技术扩展到音频和视觉数据可能共同损坏的场景。为了实现这一目标,我们采用了多模态训练范式,该范式可以在涉及声学和视觉失真的各种条件下提高模型的鲁棒性。这使得我们的失真感知模型成为现实世界挑战性环境的合理解决方案。我们将我们的方法与现有的基于变换器和基于递归神经网络的模型进行比较,这些模型试图重建从几毫秒到一秒多的缺失语音间隙。我们的实验结果表明,我们的新seq2seq架构优于最先进的Transformer解决方案的38.8%,提高语音质量和7.14%,提高语音清晰度。我们利用一个多任务学习框架,同时执行唇读(将视频组件转录为文本),同时重建相关语音的缺失部分。摘要:The process of reconstructing missing parts of speech audio from context is called speech in-painting. Human perception of speech is inherently multi-modal, involving both audio and visual (AV) cues. In this paper, we introduce and study a sequence-to-sequence (seq2seq) speech in-painting model that incorporates AV features. Our approach extends AV speech in-painting techniques to scenarios where both audio and visual data may be jointly corrupted. To achieve this, we employ a multi-modal training paradigm that boosts the robustness of our model across various conditions involving acoustic and visual distortions. This makes our distortion-aware model a plausible solution for real-world challenging environments. We compare our method with existing transformer-based and recurrent neural network-based models, which attempt to reconstruct missing speech gaps ranging from a few milliseconds to over a second. Our experimental results demonstrate that our novel seq2seq architecture outperforms the state-of-the-art transformer solution by 38.8% in terms of enhancing speech quality and 7.14% in terms of improving speech intelligibility. We exploit a multi-task learning framework that simultaneously performs lip-reading (transcribing video components to text) while reconstructing missing parts of the associated speech.

【11】 YODAS: Youtube-Oriented Dataset for Audio and Speech
标题: YODAS:面向YouTube的音频和语音数据集
作者:Xinjian Li,Shinnosuke Takamichi,Takaaki Saeki,William Chen,Sayaka Shiota,Shinji Watanabe
备注:ASRU 2023
链接:点击下载PDF文件
摘要:在这项研究中,我们介绍了YODAS(面向YouTube的音频和语音数据集),这是一个大规模的多语言数据集,目前包含超过50万小时的语音数据,来自100多种语言,来自标记和未标记的YouTube语音数据集。标记的子集,包括手动或自动字幕,有助于监督模型训练。相反,未标记的子集适合于自监督学习应用。YODAS是其规模的第一个公开可用的数据集,它是根据知识共享许可证分发的。我们介绍了收集方法用于YODAS,这有助于大规模的语音数据集的建设。随后,我们对数据集中包含的语音、文本进行了全面的分析。最后,我们描述了前15种语言的语音识别基线。摘要:In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled YouTube speech datasets. The labeled subsets, including manual or automatic subtitles, facilitate supervised model training. Conversely, the unlabeled subsets are apt for self-supervised learning applications. YODAS is distinctive as the first publicly available dataset of its scale, and it is distributed under a Creative Commons license. We introduce the collection methodology utilized for YODAS, which contributes to the large-scale speech dataset construction. Subsequently, we provide a comprehensive analysis of speech, text contained within the dataset. Finally, we describe the speech recognition baselines over the top-15 languages.

【12】 Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
标题: 具有参数化和非参数化CNN的原始波声学模型的语音误差分析
作者:Erfan Loweimi,Andrea Carmantini,Peter Bell,Steve Renals,Zoran Cvetkovic
备注:5 pages, 6 figures, 3 tables
链接:点击下载PDF文件
摘要:在本文中,我们分析了错误模式的原始波形声学模型在TIMIT的电话识别任务。我们的分析超越了传统的电话错误率(PER)指标。我们将音素分为三组:{塞擦音,双元音,摩擦音,鼻音,爆破音,半元音,元音,沉默},{辅音,元音+,沉默}和{浊音,清音,沉默},并计算每个类别中每个广义语音类的PER。我们还构建了一个混淆矩阵,为每个类别使用的替代错误和比较的混淆模式与过滤器组和Wav 2 vec 2.0系统。我们的原始波形声学模型由参数(Sinc 2Net)或非参数CNN和双向LSTM组成,在TIMIT Dev Test集上实现了低至13.7% 15.2%的PER,优于文献中原始波形模型的PER。我们还研究了迁移学习对语音错误模式和混淆矩阵的影响。它将开发 测试集的PER降低到11.8% 13.7%。摘要:In this paper, we analyse the error patterns of the raw waveform acoustic models in TIMIT's phone recognition task. Our analysis goes beyond the conventional phone error rate (PER) metric. We categorise the phones into three groups: {affricate, diphthong, fricative, nasal, plosive, semi-vowel, vowel, silence}, {consonant, vowel+, silence}, and {voiced, unvoiced, silence} and, compute the PER for each broad phonetic class in each category. We also construct a confusion matrix for each category using the substitution errors and compare the confusion patterns with those of the Filterbank and Wav2vec 2.0 systems. Our raw waveform acoustic models consists of parametric (Sinc2Net) or non-parametric CNNs and Bidirectional LSTMs, achieving down to 13.7% 15.2% PERs on TIMIT Dev Test sets, outperforming reported PERs for raw waveform models in the literature. We also investigate the impact of transfer learning from WSJ on the phonetic error patterns and confusion matrices. It reduces the PER to 11.8% 13.7% on the Dev Test sets.

【13】 Enhanced Classification of Heart Sounds Using Mel Frequency Cepstral Coefficients: A Comparative Study of Single and Ensemble Classifier Strategies
标题: 使用Mel频率倒谱系数增强心形分类:单一分类器和集合分类器策略的比较研究
作者:Amir Masoud Rahmani,Amir Haider,Parisa Khoshvaght,Mohammad Adeli,Entesar Gemeay,Yazeed Alkhrijah,Mokhtar Mohammadi,Mehdi Hosseinzadeh
链接:点击下载PDF文件
摘要:本文探讨了梅尔频率倒谱系数(MFCC)检测异常心音图使用两种分类策略的有效性:一个单一的分类器和一个集成分类器的方法。心音图分为S1、收缩期、S2和舒张期,每段估计13个MFCC,每拍产生52个MFCC。在单分类器策略中,将来自九个连续搏动的MFCC平均以分类心音图。相反,集合分类器策略采用9个分类器来单独评估心跳是否正常或异常,总体分类基于多数投票。这两种方法都在公开的心音图数据库上进行了测试。结果表明,集成分类器的策略实现了更高的准确性相比,单分类器的方法,建立MFCC更有效的比其他功能,包括时间,时间频率和统计特征,在类似的研究中评估。摘要:This paper explores the efficacy of Mel Frequency Cepstral Coefficients (MFCCs) in detecting abnormal phonocardiograms using two classification strategies: a single-classifier and an ensemble-classifier approach. Phonocardiograms were segmented into S1, systole, S2, and diastole intervals, with thirteen MFCCs estimated from each segment, yielding 52 MFCCs per beat. In the single-classifier strategy, the MFCCs from nine consecutive beats were averaged to classify phonocardiograms. Conversely, the ensemble-classifier strategy employed nine classifiers to individually assess beats as normal or abnormal, with the overall classification based on the majority vote. Both methods were tested on a publicly available phonocardiogram database. Results demonstrated that the ensemble-classifier strategy achieved higher accuracy compared to the single-classifier approach, establishing MFCCs as more effective than other features, including time, time-frequency, and statistical features, evaluated in similar studies.

【14】 Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
标题: 利用人类反馈增强Zero-Shot文本到语音合成
作者:Chen Chen,Yuchen Hu,Wen Wu,Helin Wang,Eng Siong Chng,Chao Zhang
备注:19 pages, Preprint
链接:点击下载PDF文件
摘要:近年来,文本到语音(TTS)技术取得了令人印象深刻的进步,特别是在大规模训练数据集方面,展示了人类水平的语音质量和对看不见的扬声器的令人印象深刻的zero-shot能力。然而,尽管人类的主观评价,如平均意见分数(MOS),仍然是评估合成语音质量的黄金标准,即使是最先进的TTS方法也将人类反馈与导致不匹配的训练目标和评价指标的训练隔离开来。在这项工作中,我们研究了一个新的主题,将主观的人的评价到TTS训练循环。受最近成功的人类反馈强化学习的启发,我们提出了一个全面的采样-注释-学习框架,专为TTS优化,即不确定性感知优化(UNO)。具体而言,UNO消除了奖励模型或偏好数据的需要,通过直接最大化语音生成的效用,同时考虑主观人类语音感知和评估中固有的可变性的不确定性。主观和客观评价的实验结果表明,UNO显着提高了zero-shot的TTS模型的MOS,字错误率,和说话人相似度方面的性能。此外,我们提出了一个显着的能力UNO,它可以适应所需的说话风格的情感TTS无缝和灵活。摘要:In recent years, text-to-speech (TTS) technology has witnessed impressive advancements, particularly with large-scale training datasets, showcasing human-level speech quality and impressive zero-shot capabilities on unseen speakers. However, despite human subjective evaluations, such as the mean opinion score (MOS), remaining the gold standard for assessing the quality of synthetic speech, even state-of-the-art TTS approaches have kept human feedback isolated from training that resulted in mismatched training objectives and evaluation metrics. In this work, we investigate a novel topic of integrating subjective human evaluation into the TTS training loop. Inspired by the recent success of reinforcement learning from human feedback, we propose a comprehensive sampling-annotating-learning framework tailored to TTS optimization, namely uncertainty-aware optimization (UNO). Specifically, UNO eliminates the need for a reward model or preference data by directly maximizing the utility of speech generations while considering the uncertainty that lies in the inherent variability in subjective human speech perception and evaluations. Experimental results of both subjective and objective evaluations demonstrate that UNO considerably improves the zero-shot performance of TTS models in terms of MOS, word error rate, and speaker similarity. Additionally, we present a remarkable ability of UNO that it can adapt to the desired speaking style in emotional TTS seamlessly and flexibly.

【15】 Intelligent Text-Conditioned Music Generation
标题: 智能文本条件音乐生成
作者:Zhouyao Xie,Nikhil Yadala,Xinyi Chen,Jing Xi Liu
链接:点击下载PDF文件
摘要:CLIP(对比图像预训练)是一种多模态神经网络,在(文本,图像)对上训练,以预测给定图像的最相关文本标题。它已被广泛用于图像生成,将其输出与生成模型(如VQGAN)连接,最著名的例子是OpenAI的DALLE-2。在这个项目中,我们采用类似的方法来弥合自然语言和音乐之间的差距。我们的模型分为两个步骤:首先,我们训练一个CLIP类模型对文本和音乐对对比损失,以对齐一段音乐与其最可能的文本标题。然后,我们将对齐模型与音乐解码器相结合来生成音乐。据我们所知,这是第一次尝试文本条件下的深层音乐生成。我们的实验表明,它是可能的训练的文本音乐对齐模型使用对比度损失和训练解码器生成音乐从文本提示。摘要:CLIP (Contrastive Language-Image Pre-Training) is a multimodal neural network trained on (text, image) pairs to predict the most relevant text caption given an image. It has been used extensively in image generation by connecting its output with a generative model such as VQGAN, with the most notable example being OpenAI's DALLE-2. In this project, we apply a similar approach to bridge the gap between natural language and music. Our model is split into two steps: first, we train a CLIP-like model on pairs of text and music over contrastive loss to align a piece of music with its most probable text caption. Then, we combine the alignment model with a music decoder to generate music. To the best of our knowledge, this is the first attempt at text-conditioned deep music generation. Our experiments show that it is possible to train the text-music alignment model using contrastive loss and train a decoder to generate music from text prompts.

【16】 Recent Advances in End-to-End Simultaneous Speech Translation
标题: 端到端同步语音翻译的最新进展
作者:Xiaoqian Liu,Guoqiang Hu,Yangfan Du,Erfeng He,YingFeng Luo,Chen Xu,Tong Xiao,Jingbo Zhu
链接:点击下载PDF文件
摘要:同步语音翻译(SimulST)是一项要求很高的任务,涉及在实时生成翻译的同时连续处理语音输入。本文提供了一个全面的概述,在SimulST研究的最新发展,重点是四个主要的挑战。首先,与处理冗长且连续的语音流相关联的复杂性构成了重大障碍。其次,由于需要立即翻译输出,满足实时要求存在固有的困难。第三,在翻译质量和延迟限制之间取得平衡仍然是一个关键挑战。最后,注释数据的稀缺性给任务增加了另一层复杂性。通过我们对这些挑战的探索和提出的解决方案,我们的目标是为SimulST研究的当前前景提供有价值的见解,并为未来的探索提出有前途的方向。摘要:Simultaneous speech translation (SimulST) is a demanding task that involves generating translations in real-time while continuously processing speech input. This paper offers a comprehensive overview of the recent developments in SimulST research, focusing on four major challenges. Firstly, the complexities associated with processing lengthy and continuous speech streams pose significant hurdles. Secondly, satisfying real-time requirements presents inherent difficulties due to the need for immediate translation output. Thirdly, striking a balance between translation quality and latency constraints remains a critical challenge. Finally, the scarcity of annotated data adds another layer of complexity to the task. Through our exploration of these challenges and the proposed solutions, we aim to provide valuable insights into the current landscape of SimulST research and suggest promising directions for future exploration.

【17】 Frieren: Efficient Video-to-Audio Generation with Rectified Flow Matching
标题: Frieren:具有纠正流匹配的高效视频到音频生成
作者:Yongqi Wang,Wenxiang Guo,Rongjie Huang,Jiawei Huang,Zehan Wang,Fuming You,Ruiqi Li,Zhou Zhao
链接:点击下载PDF文件
摘要:视频到音频(Video-to-Audio,V2 A)生成的目标是从无声视频中合成内容匹配的音频,如何构建具有高生成质量、高效率和视听时间同步的V2 A模型仍然是一个挑战。我们提出了Frieren,一个基于整流匹配的V2 A模型。Frieren将条件传输向量场从噪声回归到具有直线路径的潜在频谱图,并通过求解ODE进行采样,在音频质量方面优于自回归和基于分数的模型。通过采用基于前馈Transformer的非自回归向量场估计器和具有强时间对齐的通道级跨模态特征融合,我们的模型生成与输入视频高度同步的音频。此外,通过回流和引导向量场的一步蒸馏,我们的模型可以在几个,甚至只有一个采样步骤生成像样的音频。实验表明,Frieren在VGGSound上的生成质量和时间对齐方面都达到了最先进的性能,对齐准确率达到97.22%,并且在基于强扩散的基线上,初始分数提高了6.2%。音频样本可在http: frieren-v2a.github.io上获得。摘要:Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a V2A model based on rectified flow matching. Frieren regresses the conditional transport vector field from noise to spectrogram latent with straight paths and conducts sampling by solving ODE, outperforming autoregressive and score-based models in terms of audio quality. By employing a non-autoregressive vector field estimator based on a feed-forward transformer and channel-level cross-modal feature fusion with strong temporal alignment, our model generates audio that is highly synchronized with the input video. Furthermore, through reflow and one-step distillation with guided vector field, our model can generate decent audio in a few, or even only one sampling step. Experiments indicate that Frieren achieves state-of-the-art performance in both generation quality and temporal alignment on VGGSound, with alignment accuracy reaching 97.22%, and 6.2% improvement in inception score over the strong diffusion-based baseline. Audio samples are available at http: frieren-v2a.github.io .

【18】 Creative Text-to-Audio Generation via Synthesizer Programming
标题: 通过合成器编程创造性的文本到音频生成
作者:Manuel Cherep,Nikhil Singh,Jessica Shand
备注:Accepted to ICML 2024
链接:点击下载PDF文件
摘要:神经音频合成方法现在允许用自然语言指定想法。然而,这些方法产生的结果不容易调整,因为它们基于大的潜在空间和高达数十亿的不可解释的参数。我们提出了一种文本到音频生成方法,利用虚拟模块化的声音合成器,只有78个参数。合成器由于其灵活性和直观的控制,长期以来一直被熟练的声音设计师用于音乐和电影等媒体。我们的方法,CTAG,迭代更新合成器的参数,以产生高质量的文本提示,可以很容易地检查和调整音频渲染。以这种方式产生的声音也更加抽象,在细粒度的声学细节上捕捉基本的概念特征,类似于简单的草图如何生动地传达视觉概念。我们的研究结果显示了CTAG如何产生独特的声音,被认为是艺术的,但与最近的神经音频合成模型相似,将其定位为一个有价值的补充工具。摘要:Neural audio synthesis methods now allow specifying ideas in natural language. However, these methods produce results that cannot be easily tweaked, as they are based on large latent spaces and up to billions of uninterpretable parameters. We propose a text-to-audio generation method that leverages a virtual modular sound synthesizer with only 78 parameters. Synthesizers have long been used by skilled sound designers for media like music and film due to their flexibility and intuitive controls. Our method, CTAG, iteratively updates a synthesizer's parameters to produce high-quality audio renderings of text prompts that can be easily inspected and tweaked. Sounds produced this way are also more abstract, capturing essential conceptual features over fine-grained acoustic details, akin to how simple sketches can vividly convey visual concepts. Our results show how CTAG produces sounds that are distinctive, perceived as artistic, and yet similarly identifiable to recent neural audio synthesis models, positioning it as a valuable and complementary tool.

【19】 A Survey of Deep Learning Audio Generation Methods
标题: 深度学习音频生成方法综述
作者:Matej Božić,Marko Horvat
备注:14 pages, 2 figures
链接:点击下载PDF文件
摘要:本文介绍了用于音频生成的深度学习模型开发的三个不同方面的典型技术。在文章的第一部分中,我们从基本的音频波形开始解释音频表示。然后,我们进展到频域,重点是人类听觉的属性,最后介绍一个相对较新的发展。本文的主要部分重点解释基本和扩展的深度学习架构变体,以及它们在音频生成领域的实际应用。解决了以下架构:1)自动编码器2)生成对抗网络3)规范化流4)Transformer网络5)扩散模型。最后,我们将研究音频生成中常用的四种不同的评估指标。本文旨在为该领域的新手读者和初学者提供一个全面的了解音频生成方法的当前技术水平以及可以为未来研究探索的相关研究。摘要:This article presents a review of typical techniques used in three distinct aspects of deep learning model development for audio generation. In the first part of the article, we provide an explanation of audio representations, beginning with the fundamental audio waveform. We then progress to the frequency domain, with an emphasis on the attributes of human hearing, and finally introduce a relatively recent development. The main part of the article focuses on explaining basic and extended deep learning architecture variants, along with their practical applications in the field of audio generation. The following architectures are addressed: 1) Autoencoders 2) Generative adversarial networks 3) Normalizing flows 4) Transformer networks 5) Diffusion models. Lastly, we will examine four distinct evaluation metrics that are commonly employed in audio generation. This article aims to offer novice readers and beginners in the field a comprehensive understanding of the current state of the art in audio generation methods as well as relevant studies that can be explored for future research.

【20】 Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
标题: 多语言韵律迁移:比较监督学习和迁移学习
作者:Arnav Goel,Medha Hira,Anubha Gupta
备注:7 pages, Accepted to ICLR 2024 - Tiny Track
链接:点击下载PDF文件
摘要:语音合成系统中的韵律转换领域正在迅速发展。这项研究的重点是评估学习方法,使预训练的单语文本到语音(TTS)模型适应多语言条件,即,监督微调(SFT)和迁移学习(TL)。这种比较使用了三个不同的指标:平均意见评分(MOS),识别准确率(RA)和梅尔倒谱系数失真(MCD)。结果表明,与SFT相比,TL导致显著增强的性能,平均MOS高出1.53点,RA增加37.5%,MCD提高约7.8点。这些发现有助于为低资源语言建立TTS模型。摘要:The field of prosody transfer in speech synthesis systems is rapidly advancing. This research is focused on evaluating learning methods for adapting pre-trained monolingual text-to-speech (TTS) models to multilingual conditions, i.e., Supervised Fine-Tuning (SFT) and Transfer Learning (TL). This comparison utilizes three distinct metrics: Mean Opinion Score (MOS), Recognition Accuracy (RA), and Mel Cepstral Distortion (MCD). Results demonstrate that, in comparison to SFT, TL leads to significantly enhanced performance, with an average MOS higher by 1.53 points, a 37.5% increase in RA, and approximately a 7.8-point improvement in MCD. These findings are instrumental in helping build TTS models for low-resource languages.

【21】 CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
标题: CrossVoice:使用迁移学习的跨语言韵律保留Cascade-S2 ST
作者:Medha Hira,Arnav Goel,Anubha Gupta
备注:8 pages, Accepted at ICLR 2024 - Tiny Track
链接:点击下载PDF文件
摘要:本文介绍了CrossVoice,一种新型的基于级联的语音到语音翻译(S2 ST)系统,采用先进的ASR,MT和TTS技术,通过迁移学习实现跨语言的韵律保留。我们进行了全面的实验,比较CrossVoice与直接S2 ST系统,在基准数据集CVSS-T和IndicTTS上显示了Fisher Es-En,VoxPopuli Fr-En和韵律保留等任务的BLEU分数提高。CrossVoice合成的语音平均意见得分为3.75分(满分为4分),在基准测试中与人类语音不相上下,突出了基于级联的系统和迁移学习在多语言S2 ST中的有效性。摘要:This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.


机器翻译,仅供参考