今日论文合集:cs.SD语音8篇,eess.AS音频处理10篇。本文经arXiv每日学术速递授权转载
【1】Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
标题:Tango 2:通过直接偏好优化调整基于扩散的文本到音频生成作者:Navonil Majumder,Chia-Yu Hung,Deepanway Ghosal,Wei-Ning Hsu,Rada Mihalcea,Soujanya Poria摘要:生成式多模态内容在大部分内容创作领域越来越普遍,因为它有可能让艺术家和媒体人员通过快速将他们的想法变为现实来创建预制作模型。从文本提示生成音频是音乐和电影行业中这种过程的一个重要方面。最近的许多基于扩散的文本到音频模型专注于在大量的文本-音频对数据集上训练日益复杂的扩散模型。这些模型没有明确地关注概念或事件的存在以及它们在输出音频中相对于输入提示的时间顺序。我们的假设集中在音频生成的这些方面如何在有限数据的情况下提高音频生成性能。因此,在这项工作中,使用现有的文本到音频模型Tango,我们综合创建了一个偏好数据集,其中每个提示都有一个获胜者音频输出和一些失败者音频输出供扩散模型学习。从理论上讲,loser输出中有一些来自提示符的概念丢失或顺序不正确。我们使用我们的偏好数据集上的扩散DPO(直接偏好优化)损失来微调公开可用的Tango文本到音频模型,并表明它在自动和手动评估指标方面都比Tango和AudioLDM 2改进了音频输出。摘要:Generative multimodal content is increasingly prevalent in much of the content creation arena, as it has the potential to allow artists and media personnel to create pre-production mockups by quickly bringing their ideas to life. The generation of audio from text prompts is an important aspect of such processes in the music and film industry. Many of the recent diffusion-based text-to-audio models focus on training increasingly sophisticated diffusion models on a large set of datasets of prompt-audio pairs. These models do not explicitly focus on the presence of concepts or events and their temporal ordering in the output audio with respect to the input prompt. Our hypothesis is focusing on how these aspects of audio generation could improve audio generation performance in the presence of limited data. As such, in this work, using an existing text-to-audio model Tango, we synthetically create a preference dataset where each prompt has a winner audio output and some loser audio outputs for the diffusion model to learn from. The loser outputs, in theory, have some concepts from the prompt missing or in an incorrect order. We fine-tune the publicly available Tango text-to-audio model using diffusion-DPO (direct preference optimization) loss on our preference dataset and show that it leads to improved audio output over Tango and AudioLDM2, in terms of both automatic- and manual-evaluation metrics.
【2】 Scoring Intervals using Non-hierarchical Transformer For Automatic Piano Transcription标题:使用非分层Transformer自动钢琴抄写评分间隔摘要:神经半马尔可夫条件随机场(semi-CRF)框架已经证明了基于事件的钢琴转录的希望。在这个框架中,所有事件(音符或踏板)都被表示为与特定事件类型相关的闭合音程。神经半CRF方法需要一个区间评分矩阵,为每个候选区间分配一个分数。然而,设计一个有效的和有表现力的架构的评分间隔是不平凡的。在本文中,我们介绍了一种简单的方法,评分区间使用缩放的内积操作,类似于如何注意力评分在Transformers。我们从理论上证明,由于编码非重叠区间的特殊结构,在温和的条件下,内积运算的表达能力足以代表一个理想的评分矩阵,可以产生正确的转录结果。然后,我们证明了一个编码器,只有非层次的Transformer骨干,仅在低时间分辨率的功能地图上运行,是能够转录钢琴音符和踏板的高精度和时间精度。实验表明,我们的方法在Maestro数据集上的F1度量方面,在所有子任务上实现了新的最先进的性能。摘要:The neural semi-Markov Conditional Random Field (semi-CRF) framework has demonstrated promise for event-based piano transcription. In this framework, all events (notes or pedals) are represented as closed intervals tied to specific event types. The neural semi-CRF approach requires an interval scoring matrix that assigns a score for every candidate interval. However, designing an efficient and expressive architecture for scoring intervals is not trivial. In this paper, we introduce a simple method for scoring intervals using scaled inner product operations that resemble how attention scoring is done in transformers. We show theoretically that, due to the special structure from encoding the non-overlapping intervals, under a mild condition, the inner product operations are expressive enough to represent an ideal scoring matrix that can yield the correct transcription result. We then demonstrate that an encoder-only non-hierarchical transformer backbone, operating only on a low-time-resolution feature map, is capable of transcribing piano notes and pedals with high accuracy and time precision. The experiment shows that our approach achieves the new state-of-the-art performance across all subtasks in terms of the F1 measure on the Maestro dataset.【3】 Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan标题:多语言环境中的面部声音协会(FAME)挑战2024年评估计划作者:Muhammad Saad Saeed,Shah Nawaz,Muhammad Salman Tahir,Rohan Kumar Das,Muhammad Zaigham Zaheer,Marta Moscati,Markus Schedl,Muhammad Haris Khan,Karthik Nandakumar,Muhammad Haroon Yousaf备注:ACM Multimedia Conference - Grand Challenge摘要:技术的进步导致了多模态系统在各种实际应用中的使用。其中,视听系统是应用最为广泛的多模态系统之一。近年来,由于人的面部和声音之间存在独特的相关性,将它们关联起来已经得到了关注。2024年多语言环境下的人脸语音关联(FAME)挑战赛专注于探索多语言场景下的人脸语音关联。这种情况的灵感来自于这样一个事实,即世界上一半的人口是双语的,人们通常在多语言的情况下进行交流。该挑战使用数据集,即多语言视听(MAV-Celeb),用于探索多语言环境中的面部-语音关联。此报告提供FAME挑战的挑战、数据集、基线和任务详细信息的详细信息。摘要:The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associating face and voice of a person has gained attention due to presence of unique correlation between them. The Face-voice Association in Multilingual Environments (FAME) Challenge 2024 focuses on exploring face-voice association under a unique condition of multilingual scenario. This condition is inspired from the fact that half of the world's population is bilingual and most often people communicate under multilingual scenario. The challenge uses a dataset namely, Multilingual Audio-Visual (MAV-Celeb) for exploring face-voice association in multilingual environments. This report provides the details of the challenge, dataset, baselines and task details for the FAME Challenge.【4】 Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling标题:用于并行化TTC前端建模的先验不可知多尺度对比文本-音频预训练作者:Quanxiu Wang,Hui Huang,Mingjie Wang,Yong Dai,Jinzuomu Zhong,Benlai Tang摘要:在过去的十年中,一系列不懈的努力一直致力于开发高表达和可控的文本到语音(TTS)系统。一般来说,整体TTS包括两个互连的组件:前端模块和后端模块。前端擅长从原始文本输入中捕获语言表示,而后端模块将语言提示转换为语音。研究界对前端组件的研究表现出越来越大的兴趣,认识到其在文本到语音系统中的关键作用,包括文本规范化(TN),韵律边界预测(PBP)和多音字消歧(PD)。尽管如此,注释文本数据不足以及对同质文本信号的依赖所带来的局限性大大削弱了其监督学习的有效性。为了避免这一障碍,本文提出了一种新的两级TTS前端预测流水线,命名为TAP-FM。具体来说,在第一个学习阶段,我们提出了一个多尺度对比文本-音频预训练协议(MC-TAP),它致力于以无监督的方式通过多粒度对比预训练获得更丰富的见解。我们的框架没有在之前的预训练方法中挖掘同质特征,而是展示了深入研究全局和局部文本音频语义和声学表示的能力。此外,一个并行TTS前端模型被精心设计,以执行TN,PD和PBP预测任务,分别在第二阶段。最后,大量的实验说明了我们提出的方法的优越性,实现国家的最先进的性能。摘要:Over the past decade, a series of unflagging efforts have been dedicated to developing highly expressive and controllable text-to-speech (TTS) systems. In general, the holistic TTS comprises two interconnected components: the frontend module and the backend module. The frontend excels in capturing linguistic representations from the raw text input, while the backend module converts linguistic cues to speech. The research community has shown growing interest in the study of the frontend component, recognizing its pivotal role in text-to-speech systems, including Text Normalization (TN), Prosody Boundary Prediction (PBP), and Polyphone Disambiguation (PD). Nonetheless, the limitations posed by insufficient annotated textual data and the reliance on homogeneous text signals significantly undermine the effectiveness of its supervised learning. To evade this obstacle, a novel two-stage TTS frontend prediction pipeline, named TAP-FM, is proposed in this paper. Specifically, during the first learning phase, we present a Multi-scale Contrastive Text-audio Pre-training protocol (MC-TAP), which hammers at acquiring richer insights via multi-granularity contrastive pre-training in an unsupervised manner. Instead of mining homogeneous features in prior pre-training approaches, our framework demonstrates the ability to delve deep into both global and local text-audio semantic and acoustic representations. Furthermore, a parallelized TTS frontend model is delicately devised to execute TN, PD, and PBP prediction tasks, respectively in the second stage. Finally, extensive experiments illustrate the superiority of our proposed method, achieving state-of-the-art performance.【5】 An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging作者:Gabriel Meseguer-Brocal,Dorian Desblancs,Romain Hennequin摘要:自监督学习已经成为一种强大的方法,可以在大量未标记的数据上预训练可推广的机器学习模型。它在音乐领域尤其引人注目,因为获得标记数据非常耗时,容易出错,而且模棱两可。在自我监督过程中,模型在借口任务上进行训练,主要目标是获取鲁棒和信息丰富的特征,这些特征可以在以后针对特定的下游任务进行微调。借口任务的选择是至关重要的,因为它引导模型用有意义的信息编码约束来塑造特征空间。在音乐方面,大多数作品都依赖于对比学习或掩蔽技术。在这项研究中,我们通过调查和比较新的自我监督方法的音乐标记的性能,扩大了适用于音乐的借口任务的范围。我们开源了一个简单的ResNet模型,该模型在数百万曲目的多样化目录上训练。我们的研究结果表明,尽管大多数这些预训练方法都会产生类似的下游结果,但与其他自监督预训练方法相比,对比学习始终会产生更好的下游性能。这在有限数据的下游上下文中是正确的。摘要:Self-supervised learning has emerged as a powerful way to pre-train generalizable machine learning models on large amounts of unlabeled data. It is particularly compelling in the music domain, where obtaining labeled data is time-consuming, error-prone, and ambiguous. During the self-supervised process, models are trained on pretext tasks, with the primary objective of acquiring robust and informative features that can later be fine-tuned for specific downstream tasks. The choice of the pretext task is critical as it guides the model to shape the feature space with meaningful constraints for information encoding. In the context of music, most works have relied on contrastive learning or masking techniques. In this study, we expand the scope of pretext tasks applied to music by investigating and comparing the performance of new self-supervised methods for music tagging. We open-source a simple ResNet model trained on a diverse catalog of millions of tracks. Our results demonstrate that, although most of these pre-training methods result in similar downstream results, contrastive learning consistently results in better downstream performance compared to other self-supervised pre-training methods. This holds true in a limited-data downstream context.【6】 Voice Attribute Editing with Text Prompt作者:Zhengyan Sheng,Yang Ai,Li-Juan Liu,Jia Pan,Zhen-Hua Ling摘要:尽管最近在利用文本提示提供对语音风格的控制的语音生成方面取得了进步,但是合成语音中的语音属性仍然难以控制并且具有挑战性。本文介绍了一种新颖的任务:带文本提示的语音属性编辑,其目标是根据文本提示中描述的动作对语音属性进行相应的修改。为了解决这个问题,VoxEditor,一个端到端的生成模型,提出。在VoxEditor中,针对文本提示的不足,设计了一个剩余记忆(ResMem)块,有效地将语音属性和这些描述符映射到共享特征空间。此外,ResMem模块通过语音属性程度预测(VADP)模块进行增强,以将语音属性与相应的描述符对齐,解决了由于语音属性的非定量描述而导致的文本提示的不精确性。我们还建立了开源的VCTK-RVA数据集,该数据集在手动注释方面处于领先地位,详细说明了不同说话者之间的语音特征差异。大量的实验表明,我们所提出的方法的有效性和通用性的客观和主观的指标。数据集和音频样本可在网站上查阅。摘要:Despite recent advancements in speech generation with text prompt providing control over speech style, voice attributes in synthesized speech remain elusive and challenging to control. This paper introduces a novel task: voice attribute editing with text prompt, with the goal of making relative modifications to voice attributes according to the actions described in the text prompt. To solve this task, VoxEditor, an end-to-end generative model, is proposed. In VoxEditor, addressing the insufficiency of text prompt, a Residual Memory (ResMem) block is designed, that efficiently maps voice attributes and these descriptors into the shared feature space. Additionally, the ResMem block is enhanced with a voice attribute degree prediction (VADP) block to align voice attributes with corresponding descriptors, addressing the imprecision of text prompt caused by non-quantitative descriptions of voice attributes. We also establish the open-source VCTK-RVA dataset, which leads the way in manual annotations detailing voice characteristic differences among different speakers. Extensive experiments demonstrate the effectiveness and generalizability of our proposed method in terms of both objective and subjective metrics. The dataset and audio samples are available on the website.
【7】 Interactive Sonification for Health and Energy using ChucK and Unity标题:使用ChucK和Unity实现健康和能源的互动Sonification作者:Yichun Zhao,George Tzanetakis摘要:声化可以提供关于数据的有价值的见解,但大多数现有方法都不是设计成由用户以交互方式控制的。交互使发声的设计者能够更快速地试验声音设计,并允许通过与各种控制参数交互来实时修改发声。在本文中,我们描述了两个案例研究的交互式声化,利用公开的数据集,最近在国际会议上的听觉显示(ICAD)。它们来自健康和能源领域:脑电图(EEG)α波数据和由二氧化氮,二氧化硫,一氧化碳和臭氧组成的空气污染物数据。我们展示了如何可以重新创建这些sonfications,以支持使用ChucK,Unity和Chunity构建的通用交互式sonification框架进行交互。除了支持典型的发音方法,是常见的现有的发音工具包,我们的框架引入了新的方法,如支持离散事件,交错播放多个数据流进行比较,并使用调频(FM)合成方面的一个数据属性调制另一个。我们还描述了如何使用这些新功能来改善我们所研究的两个数据集的发音体验。摘要:Sonification can provide valuable insights about data but most existing approaches are not designed to be controlled by the user in an interactive fashion. Interactions enable the designer of the sonification to more rapidly experiment with sound design and allow the sonification to be modified in real-time by interacting with various control parameters. In this paper, we describe two case studies of interactive sonification that utilize publicly available datasets that have been described recently in the International Conference on Auditory Display (ICAD). They are from the health and energy domains: electroencephalogram (EEG) alpha wave data and air pollutant data consisting of nitrogen dioxide, sulfur dioxide, carbon monoxide, and ozone. We show how these sonfications can be recreated to support interaction utilizing a general interactive sonification framework built using ChucK, Unity, and Chunity. In addition to supporting typical sonification methods that are common in existing sonification toolkits, our framework introduces novel methods such as supporting discrete events, interleaved playback of multiple data streams for comparison, and using frequency modulation (FM) synthesis in terms of one data attribute modulating another. We also describe how these new functionalities can be used to improve the sonification experience of the two datasets we have investigated.【8】 Anatomy of Industrial Scale Multilingual ASR作者:Francis McCann Ramirez,Luka Chkhetiani,Andrew Ehrenberg,Robert McHardy,Rami Botros,Yash Khare,Andrea Vanzo,Taufiquzzaman Peyash,Gabriel Oexle,Michael Liang,Ilya Sklyar,Enver Fakhan,Ahmed Efty,Daniel McCrystal,Sam Flamini,Domenic Donato,Takuya Yoshioka摘要:本文描述了AssemblyAI的工业级自动语音识别(ASR)系统,旨在满足服务于各种应用需求的大规模多语言ASR的要求。我们的系统利用了不同的训练数据集,包括四种语言的无监督(1250万小时),监督(188 k小时)和伪标记(160万小时)数据。我们提供了我们的模型架构的详细描述,包括一个完整的上下文600 M参数的Conformer编码器预训练与BEST-RQ和RNN-T解码器微调联合编码器。我们的广泛评估表明,与更大、计算更昂贵的模型(如Whisper large和Canary-1B)相比,它的字错误率(WER)具有竞争力。此外,我们的架构选择产生了几个关键优势,包括改进的代码切换能力,与优化的Whisper基线相比,推理速度提高了5倍,语音数据的幻觉率降低了30%,与Whisper相比,环境噪声降低了90%,以及显着提高的时间戳准确性。在整个工作中,我们采用以系统为中心的方法来分析成熟的ASR模型的各个方面,以获得对大规模运营的真实服务有用的实际相关见解。摘要:This paper describes AssemblyAI's industrial-scale automatic speech recognition (ASR) system, designed to meet the requirements of large-scale, multilingual ASR serving various application needs. Our system leverages a diverse training dataset comprising unsupervised (12.5M hours), supervised (188k hours), and pseudo-labeled (1.6M hours) data across four languages. We provide a detailed description of our model architecture, consisting of a full-context 600M-parameter Conformer encoder pre-trained with BEST-RQ and an RNN-T decoder fine-tuned jointly with the encoder. Our extensive evaluation demonstrates competitive word error rates (WERs) against larger and more computationally expensive models, such as Whisper large and Canary-1B. Furthermore, our architectural choices yield several key advantages, including an improved code-switching capability, a 5x inference speedup compared to an optimized Whisper baseline, a 30% reduction in hallucination rate on speech data, and a 90% reduction in ambient noise compared to Whisper, along with significantly improved time-stamp accuracy. Throughout this work, we adopt a system-centric approach to analyzing various aspects of fully-fledged ASR models to gain practically relevant insights useful for real-world services operating at scale.
【1】 Anatomy of Industrial Scale Multilingual ASR作者:Francis McCann Ramirez,Luka Chkhetiani,Andrew Ehrenberg,Robert McHardy,Rami Botros,Yash Khare,Andrea Vanzo,Taufiquzzaman Peyash,Gabriel Oexle,Michael Liang,Ilya Sklyar,Enver Fakhan,Ahmed Efty,Daniel McCrystal,Sam Flamini,Domenic Donato,Takuya Yoshioka摘要:本文描述了AssemblyAI的工业级自动语音识别(ASR)系统,旨在满足服务于各种应用需求的大规模多语言ASR的要求。我们的系统利用了不同的训练数据集,包括四种语言的无监督(1250万小时),监督(188 k小时)和伪标记(160万小时)数据。我们提供了我们的模型架构的详细描述,包括一个完整的上下文600 M参数的Conformer编码器预训练与BEST-RQ和RNN-T解码器微调联合编码器。我们的广泛评估表明,与更大、计算更昂贵的模型(如Whisper large和Canary-1B)相比,它的字错误率(WER)具有竞争力。此外,我们的架构选择产生了几个关键优势,包括改进的代码切换能力,与优化的Whisper基线相比,推理速度提高了5倍,语音数据的幻觉率降低了30%,与Whisper相比,环境噪声降低了90%,以及显着提高的时间戳准确性。在整个工作中,我们采用以系统为中心的方法来分析成熟的ASR模型的各个方面,以获得对大规模运营的真实服务有用的实际相关见解。摘要:This paper describes AssemblyAI's industrial-scale automatic speech recognition (ASR) system, designed to meet the requirements of large-scale, multilingual ASR serving various application needs. Our system leverages a diverse training dataset comprising unsupervised (12.5M hours), supervised (188k hours), and pseudo-labeled (1.6M hours) data across four languages. We provide a detailed description of our model architecture, consisting of a full-context 600M-parameter Conformer encoder pre-trained with BEST-RQ and an RNN-T decoder fine-tuned jointly with the encoder. Our extensive evaluation demonstrates competitive word error rates (WERs) against larger and more computationally expensive models, such as Whisper large and Canary-1B. Furthermore, our architectural choices yield several key advantages, including an improved code-switching capability, a 5x inference speedup compared to an optimized Whisper baseline, a 30% reduction in hallucination rate on speech data, and a 90% reduction in ambient noise compared to Whisper, along with significantly improved time-stamp accuracy. Throughout this work, we adopt a system-centric approach to analyzing various aspects of fully-fledged ASR models to gain practically relevant insights useful for real-world services operating at scale.【2】 A Large-Scale Evaluation of Speech Foundation Models作者:Shu-wen Yang,Heng-Jui Chang,Zili Huang,Andy T. Liu,Cheng-I Lai,Haibin Wu,Jiatong Shi,Xuankai Chang,Hsiang-Sheng Tsai,Wen-Chin Huang,Tzu-hsun Feng,Po-Han Chi,Yist Y. Lin,Yung-Sung Chuang,Tzu-Hsien Huang,Wei-Cheng Tseng,Kushal Lakhotia,Shang-Wen Li,Abdelrahman Mohamed,Shinji Watanabe,Hung-yi Lee备注:The extended journal version for SUPERB and SUPERB-SG. Accepted to TASLP. The arxiv version is further refined摘要:基础模型范例利用共享的基础模型为各种任务实现最先进的(SOTA)性能,只需要最少的特定于下游的建模和数据注释。这种方法在自然语言处理(NLP)领域中至关重要。然而,语音处理社区缺乏类似的设置来系统地探索范式。在这项工作中,我们建立了语音处理通用性能基准(SUPERB)研究语音的有效性的范例。我们提出了一个统一的多任务框架,以解决语音处理任务的SUPERB使用冻结的基础模型,然后任务专用的,轻量级的预测头。将我们的结果与社区提交的结果相结合,我们验证了基础模型范式对于语音是有希望的,并且我们的多任务框架是简单而有效的,因为表现最好的基础模型在大多数SUPERB任务中表现出具有竞争力的泛化能力。为了实现可重复性和可扩展性,我们开发了一个长期维护的平台,该平台支持确定性基准测试,允许通过在线排行榜共享结果,并通过社区驱动的基准测试数据库促进协作,以支持新的开发周期。最后,我们进行了一系列分析,以深入了解SUPERB和语音基础模型,包括模型内部任务之间的信息流,加权和基准测试协议的正确性以及基准测试的统计意义和鲁棒性。摘要:The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. In this work, we establish the Speech processing Universal PERformance Benchmark (SUPERB) to study the effectiveness of the paradigm for speech. We propose a unified multi-tasking framework to address speech processing tasks in SUPERB using a frozen foundation model followed by task-specialized, lightweight prediction heads. Combining our results with community submissions, we verify that the foundation model paradigm is promising for speech, and our multi-tasking framework is simple yet effective, as the best-performing foundation model shows competitive generalizability across most SUPERB tasks. For reproducibility and extensibility, we have developed a long-term maintained platform that enables deterministic benchmarking, allows for result sharing via an online leaderboard, and promotes collaboration through a community-driven benchmark database to support new development cycles. Finally, we conduct a series of analyses to offer an in-depth understanding of SUPERB and speech foundation models, including information flows across tasks inside the models, the correctness of the weighted-sum benchmarking protocol and the statistical significance and robustness of the benchmark.【3】 Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment标题:文本转歌:迈向可控音乐一代激发声乐和伴奏作者:Hong Zhiqing,Huang Rongjie,Cheng Xize,Wang Yongqi,Li Ruiqi,You Fuming,Zhao Zhou,Zhang Zhimeng摘要:一首歌是歌声和伴奏的结合。然而,现有的作品集中在歌唱声音合成和音乐生成独立。很少有人注意探索歌曲合成。在这项工作中,我们提出了一个新的任务,称为文本到歌曲合成,其中包括人声和milliments生成。我们开发的Melodist,一个两阶段的文本到歌曲的方法,包括歌声合成(SVS)和人声到伴奏(V2A)的合成。Melodist利用三塔对比预训练来学习更有效的文本表示,以实现可控的V2A合成。为了缓解数据不足的问题,本文建立了一个从音乐网站中挖掘出的中文歌曲数据集。我们的数据集上的评估结果表明,Melodist可以合成具有可比质量和风格一致性的歌曲。音频样本可以在https: text2songMelodist.github.io Sample 上找到。摘要:A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel task called text-to-song synthesis which incorporating both vocals and accompaniments generation. We develop Melodist, a two-stage text-to-song method that consists of singing voice synthesis (SVS) and vocal-to-accompaniment (V2A) synthesis. Melodist leverages tri-tower contrastive pretraining to learn more effective text representation for controllable V2A synthesis. A Chinese song dataset mined from a music website is built up to alleviate data scarcity for our research. The evaluation results on our dataset demonstrate that Melodist can synthesize songs with comparable quality and style consistency. Audio samples can be found in https: text2songMelodist.github.io Sample .【4】 Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization标题:Tango 2:通过直接偏好优化调整基于扩散的文本到音频生成作者:Navonil Majumder,Chia-Yu Hung,Deepanway Ghosal,Wei-Ning Hsu,Rada Mihalcea,Soujanya Poria摘要:生成式多模态内容在大部分内容创作领域越来越普遍,因为它有可能让艺术家和媒体人员通过快速将他们的想法变为现实来创建预制作模型。从文本提示生成音频是音乐和电影行业中这种过程的一个重要方面。最近的许多基于扩散的文本到音频模型专注于在大量的文本-音频对数据集上训练日益复杂的扩散模型。这些模型没有明确地关注概念或事件的存在以及它们在输出音频中相对于输入提示的时间顺序。我们的假设集中在音频生成的这些方面如何在有限数据的情况下提高音频生成性能。因此,在这项工作中,使用现有的文本到音频模型Tango,我们综合创建了一个偏好数据集,其中每个提示都有一个获胜者音频输出和一些失败者音频输出供扩散模型学习。从理论上讲,loser输出中有一些来自提示符的概念丢失或顺序不正确。我们使用我们的偏好数据集上的扩散DPO(直接偏好优化)损失来微调公开可用的Tango文本到音频模型,并表明它在自动和手动评估指标方面都比Tango和AudioLDM 2改进了音频输出。摘要:Generative multimodal content is increasingly prevalent in much of the content creation arena, as it has the potential to allow artists and media personnel to create pre-production mockups by quickly bringing their ideas to life. The generation of audio from text prompts is an important aspect of such processes in the music and film industry. Many of the recent diffusion-based text-to-audio models focus on training increasingly sophisticated diffusion models on a large set of datasets of prompt-audio pairs. These models do not explicitly focus on the presence of concepts or events and their temporal ordering in the output audio with respect to the input prompt. Our hypothesis is focusing on how these aspects of audio generation could improve audio generation performance in the presence of limited data. As such, in this work, using an existing text-to-audio model Tango, we synthetically create a preference dataset where each prompt has a winner audio output and some loser audio outputs for the diffusion model to learn from. The loser outputs, in theory, have some concepts from the prompt missing or in an incorrect order. We fine-tune the publicly available Tango text-to-audio model using diffusion-DPO (direct preference optimization) loss on our preference dataset and show that it leads to improved audio output over Tango and AudioLDM2, in terms of both automatic- and manual-evaluation metrics.
【5】 Scoring Intervals using Non-hierarchical Transformer For Automatic Piano Transcription标题:使用非分层Transformer自动钢琴抄写评分间隔摘要:神经半马尔可夫条件随机场(semi-CRF)框架已经证明了基于事件的钢琴转录的希望。在这个框架中,所有事件(音符或踏板)都被表示为与特定事件类型相关的闭合音程。神经半CRF方法需要一个区间评分矩阵,为每个候选区间分配一个分数。然而,设计一个有效的和有表现力的架构的评分间隔是不平凡的。在本文中,我们介绍了一种简单的方法,评分区间使用缩放的内积操作,类似于如何注意力评分在Transformers。我们从理论上证明,由于编码非重叠区间的特殊结构,在温和的条件下,内积运算的表达能力足以代表一个理想的评分矩阵,可以产生正确的转录结果。然后,我们证明了一个编码器,只有非层次的Transformer骨干,仅在低时间分辨率的功能地图上运行,是能够转录钢琴音符和踏板的高精度和时间精度。实验表明,我们的方法在Maestro数据集上的F1度量方面,在所有子任务上实现了新的最先进的性能。摘要:The neural semi-Markov Conditional Random Field (semi-CRF) framework has demonstrated promise for event-based piano transcription. In this framework, all events (notes or pedals) are represented as closed intervals tied to specific event types. The neural semi-CRF approach requires an interval scoring matrix that assigns a score for every candidate interval. However, designing an efficient and expressive architecture for scoring intervals is not trivial. In this paper, we introduce a simple method for scoring intervals using scaled inner product operations that resemble how attention scoring is done in transformers. We show theoretically that, due to the special structure from encoding the non-overlapping intervals, under a mild condition, the inner product operations are expressive enough to represent an ideal scoring matrix that can yield the correct transcription result. We then demonstrate that an encoder-only non-hierarchical transformer backbone, operating only on a low-time-resolution feature map, is capable of transcribing piano notes and pedals with high accuracy and time precision. The experiment shows that our approach achieves the new state-of-the-art performance across all subtasks in terms of the F1 measure on the Maestro dataset.
【6】 Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan标题:多语言环境中的面部声音协会(FAME)挑战2024年评估计划作者:Muhammad Saad Saeed,Shah Nawaz,Muhammad Salman Tahir,Rohan Kumar Das,Muhammad Zaigham Zaheer,Marta Moscati,Markus Schedl,Muhammad Haris Khan,Karthik Nandakumar,Muhammad Haroon Yousaf备注:ACM Multimedia Conference - Grand Challenge摘要:技术的进步导致了多模态系统在各种实际应用中的使用。其中,视听系统是应用最为广泛的多模态系统之一。近年来,由于人的面部和声音之间存在独特的相关性,将它们关联起来已经得到了关注。2024年多语言环境下的人脸语音关联(FAME)挑战赛专注于探索多语言场景下的人脸语音关联。这种情况的灵感来自于这样一个事实,即世界上一半的人口是双语的,人们通常在多语言的情况下进行交流。该挑战使用数据集,即多语言视听(MAV-Celeb),用于探索多语言环境中的面部-语音关联。此报告提供FAME挑战的挑战、数据集、基线和任务详细信息的详细信息。摘要:The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associating face and voice of a person has gained attention due to presence of unique correlation between them. The Face-voice Association in Multilingual Environments (FAME) Challenge 2024 focuses on exploring face-voice association under a unique condition of multilingual scenario. This condition is inspired from the fact that half of the world's population is bilingual and most often people communicate under multilingual scenario. The challenge uses a dataset namely, Multilingual Audio-Visual (MAV-Celeb) for exploring face-voice association in multilingual environments. This report provides the details of the challenge, dataset, baselines and task details for the FAME Challenge.【7】 Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling标题:用于并行化TTC前端建模的先验不可知多尺度对比文本-音频预训练作者:Quanxiu Wang,Hui Huang,Mingjie Wang,Yong Dai,Jinzuomu Zhong,Benlai Tang摘要:在过去的十年中,一系列不懈的努力一直致力于开发高表达和可控的文本到语音(TTS)系统。一般来说,整体TTS包括两个互连的组件:前端模块和后端模块。前端擅长从原始文本输入中捕获语言表示,而后端模块将语言提示转换为语音。研究界对前端组件的研究表现出越来越大的兴趣,认识到其在文本到语音系统中的关键作用,包括文本规范化(TN),韵律边界预测(PBP)和多音字消歧(PD)。尽管如此,注释文本数据不足以及对同质文本信号的依赖所带来的局限性大大削弱了其监督学习的有效性。为了避免这一障碍,本文提出了一种新的两级TTS前端预测流水线,命名为TAP-FM。具体来说,在第一个学习阶段,我们提出了一个多尺度对比文本-音频预训练协议(MC-TAP),它致力于以无监督的方式通过多粒度对比预训练获得更丰富的见解。我们的框架没有在之前的预训练方法中挖掘同质特征,而是展示了深入研究全局和局部文本音频语义和声学表示的能力。此外,一个并行TTS前端模型被精心设计,以执行TN,PD和PBP预测任务,分别在第二阶段。最后,大量的实验说明了我们提出的方法的优越性,实现国家的最先进的性能。摘要:Over the past decade, a series of unflagging efforts have been dedicated to developing highly expressive and controllable text-to-speech (TTS) systems. In general, the holistic TTS comprises two interconnected components: the frontend module and the backend module. The frontend excels in capturing linguistic representations from the raw text input, while the backend module converts linguistic cues to speech. The research community has shown growing interest in the study of the frontend component, recognizing its pivotal role in text-to-speech systems, including Text Normalization (TN), Prosody Boundary Prediction (PBP), and Polyphone Disambiguation (PD). Nonetheless, the limitations posed by insufficient annotated textual data and the reliance on homogeneous text signals significantly undermine the effectiveness of its supervised learning. To evade this obstacle, a novel two-stage TTS frontend prediction pipeline, named TAP-FM, is proposed in this paper. Specifically, during the first learning phase, we present a Multi-scale Contrastive Text-audio Pre-training protocol (MC-TAP), which hammers at acquiring richer insights via multi-granularity contrastive pre-training in an unsupervised manner. Instead of mining homogeneous features in prior pre-training approaches, our framework demonstrates the ability to delve deep into both global and local text-audio semantic and acoustic representations. Furthermore, a parallelized TTS frontend model is delicately devised to execute TN, PD, and PBP prediction tasks, respectively in the second stage. Finally, extensive experiments illustrate the superiority of our proposed method, achieving state-of-the-art performance.【8】 An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging作者:Gabriel Meseguer-Brocal,Dorian Desblancs,Romain Hennequin摘要:自监督学习已经成为一种强大的方法,可以在大量未标记的数据上预训练可推广的机器学习模型。它在音乐领域尤其引人注目,因为获得标记数据非常耗时,容易出错,而且模棱两可。在自我监督过程中,模型在借口任务上进行训练,主要目标是获取鲁棒和信息丰富的特征,这些特征可以在以后针对特定的下游任务进行微调。借口任务的选择是至关重要的,因为它引导模型用有意义的信息编码约束来塑造特征空间。在音乐方面,大多数作品都依赖于对比学习或掩蔽技术。在这项研究中,我们通过调查和比较新的自我监督方法的音乐标记的性能,扩大了适用于音乐的借口任务的范围。我们开源了一个简单的ResNet模型,该模型在数百万曲目的多样化目录上训练。我们的研究结果表明,尽管大多数这些预训练方法都会产生类似的下游结果,但与其他自监督预训练方法相比,对比学习始终会产生更好的下游性能。这在有限数据的下游上下文中是正确的。摘要:Self-supervised learning has emerged as a powerful way to pre-train generalizable machine learning models on large amounts of unlabeled data. It is particularly compelling in the music domain, where obtaining labeled data is time-consuming, error-prone, and ambiguous. During the self-supervised process, models are trained on pretext tasks, with the primary objective of acquiring robust and informative features that can later be fine-tuned for specific downstream tasks. The choice of the pretext task is critical as it guides the model to shape the feature space with meaningful constraints for information encoding. In the context of music, most works have relied on contrastive learning or masking techniques. In this study, we expand the scope of pretext tasks applied to music by investigating and comparing the performance of new self-supervised methods for music tagging. We open-source a simple ResNet model trained on a diverse catalog of millions of tracks. Our results demonstrate that, although most of these pre-training methods result in similar downstream results, contrastive learning consistently results in better downstream performance compared to other self-supervised pre-training methods. This holds true in a limited-data downstream context.
【9】 Voice Attribute Editing with Text Prompt作者:Zhengyan Sheng,Yang Ai,Li-Juan Liu,Jia Pan,Zhen-Hua Ling摘要:尽管最近在利用文本提示提供对语音风格的控制的语音生成方面取得了进步,但是合成语音中的语音属性仍然难以控制并且具有挑战性。本文介绍了一种新颖的任务:带文本提示的语音属性编辑,其目标是根据文本提示中描述的动作对语音属性进行相应的修改。为了解决这个问题,VoxEditor,一个端到端的生成模型,提出。在VoxEditor中,针对文本提示的不足,设计了一个剩余记忆(ResMem)块,有效地将语音属性和这些描述符映射到共享特征空间。此外,ResMem模块通过语音属性程度预测(VADP)模块进行增强,以将语音属性与相应的描述符对齐,解决了由于语音属性的非定量描述而导致的文本提示的不精确性。我们还建立了开源的VCTK-RVA数据集,该数据集在手动注释方面处于领先地位,详细说明了不同说话者之间的语音特征差异。大量的实验表明,我们所提出的方法的有效性和通用性的客观和主观的指标。数据集和音频样本可在网站上查阅。摘要:Despite recent advancements in speech generation with text prompt providing control over speech style, voice attributes in synthesized speech remain elusive and challenging to control. This paper introduces a novel task: voice attribute editing with text prompt, with the goal of making relative modifications to voice attributes according to the actions described in the text prompt. To solve this task, VoxEditor, an end-to-end generative model, is proposed. In VoxEditor, addressing the insufficiency of text prompt, a Residual Memory (ResMem) block is designed, that efficiently maps voice attributes and these descriptors into the shared feature space. Additionally, the ResMem block is enhanced with a voice attribute degree prediction (VADP) block to align voice attributes with corresponding descriptors, addressing the imprecision of text prompt caused by non-quantitative descriptions of voice attributes. We also establish the open-source VCTK-RVA dataset, which leads the way in manual annotations detailing voice characteristic differences among different speakers. Extensive experiments demonstrate the effectiveness and generalizability of our proposed method in terms of both objective and subjective metrics. The dataset and audio samples are available on the website.
【10】 Interactive Sonification for Health and Energy using ChucK and Unity标题:使用ChucK和Unity实现健康和能源的互动Sonification作者:Yichun Zhao,George Tzanetakis摘要:声化可以提供关于数据的有价值的见解,但大多数现有方法都不是设计成由用户以交互方式控制的。交互使发声的设计者能够更快速地试验声音设计,并允许通过与各种控制参数交互来实时修改发声。在本文中,我们描述了两个案例研究的交互式声化,利用公开的数据集,最近在国际会议上的听觉显示(ICAD)。它们来自健康和能源领域:脑电图(EEG)α波数据和由二氧化氮,二氧化硫,一氧化碳和臭氧组成的空气污染物数据。我们展示了如何可以重新创建这些sonfications,以支持使用ChucK,Unity和Chunity构建的通用交互式sonification框架进行交互。除了支持典型的发音方法,是常见的现有的发音工具包,我们的框架引入了新的方法,如支持离散事件,交错播放多个数据流进行比较,并使用调频(FM)合成方面的一个数据属性调制另一个。我们还描述了如何使用这些新功能来改善我们所研究的两个数据集的发音体验。摘要:Sonification can provide valuable insights about data but most existing approaches are not designed to be controlled by the user in an interactive fashion. Interactions enable the designer of the sonification to more rapidly experiment with sound design and allow the sonification to be modified in real-time by interacting with various control parameters. In this paper, we describe two case studies of interactive sonification that utilize publicly available datasets that have been described recently in the International Conference on Auditory Display (ICAD). They are from the health and energy domains: electroencephalogram (EEG) alpha wave data and air pollutant data consisting of nitrogen dioxide, sulfur dioxide, carbon monoxide, and ozone. We show how these sonfications can be recreated to support interaction utilizing a general interactive sonification framework built using ChucK, Unity, and Chunity. In addition to supporting typical sonification methods that are common in existing sonification toolkits, our framework introduces novel methods such as supporting discrete events, interleaved playback of multiple data streams for comparison, and using frequency modulation (FM) synthesis in terms of one data attribute modulating another. We also describe how these new functionalities can be used to improve the sonification experience of the two datasets we have investigated