本文经arXiv每日学术速递授权转载
标题: 通过多算法分析和用户友好的可视化增强音乐流派分类
作者:Navin Kamuni,Dheerendra Panwar
链接:点击下载PDF文件
摘要:这项研究的目的是教算法如何识别不同类型的音乐。用户将提交歌曲进行分析。由于算法以前从未听过这些歌曲,因此它需要弄清楚是什么让每首歌曲都独一无二。它通过将歌曲分解为不同的部分,并通过监督学习来学习节奏,旋律和音调等内容,因为程序从已经标记的示例中学习。分类音乐时要考虑的一个重要因素是它的类型,这可能相当复杂。为了确保准确性,我们使用了五种不同的算法来分析歌曲,每种算法都独立工作。这有助于我们更全面地了解每首歌的特点。因此,我们的目标是正确识别每首提交歌曲的类型。分析完成后,结果将使用图形工具呈现,便于用户理解和提供反馈。摘要:The aim of this study is to teach an algorithm how to recognize different types of music. Users will submit songs for analysis. Since the algorithm hasn't heard these songs before, it needs to figure out what makes each song unique. It does this by breaking down the songs into different parts and studying things like rhythm, melody, and tone via supervised learning because the program learns from examples that are already labelled. One important thing to consider when classifying music is its genre, which can be quite complex. To ensure accuracy, we use five different algorithms, each working independently, to analyze the songs. This helps us get a more complete understanding of each song's characteristics. Therefore, our goal is to correctly identify the genre of each submitted song. Once the analysis is done, the results are presented using a graphing tool, making it easy for users to understand and provide feedback.
【2】 A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
标题: 使用半监督语音嵌入的PD检测新型融合架构
作者:Tariq Adnan,Abdelrahman Abdelkader,Zipei Liu,Ekram Hossain,Sooyong Park,MD Saiful Islam,Ehsan Hoque
备注:25 pages, 5 figures, and 4 tables
链接:点击下载PDF文件
摘要:我们提出了一个框架,以识别帕金森氏病(PD)通过英语pangram话语语音收集使用Web应用程序从不同的记录设置和环境,包括参与者的家。我们的数据集包括1306名参与者的全球队列,其中包括392名诊断为PD的患者。利用数据集的多样性,跨越各种人口统计学属性(如年龄,性别和种族),我们使用了来自半监督模型(如Wav2Vec 2.0,WavLM和ImageBind)的深度学习嵌入,代表与PD相关的语音动态。我们用于PD分类的新型融合模型将不同的语音嵌入对齐到一个有凝聚力的特征空间中,与标准的基于级联的融合模型和其他基线(包括基于传统声学特征构建的模型)相比,表现出卓越的性能。在随机数据分割配置中,该模型的受试者工作特征曲线下面积(AUROC)为88.94%,准确度为85.65%。严格的统计分析证实,我们的模型在性别、种族和年龄方面的各种人口统计学亚组中表现公平,并且无论疾病持续时间如何都保持稳健。此外,当对从临床环境和PD护理中心收集的两个完全看不见的测试数据集进行测试时,我们的模型分别保持了82.12%和78.44%的AUROC评分。这肯定了该模型的鲁棒性,以及它在现实世界应用中提高可及性和健康公平性的潜力。摘要:We present a framework to recognize Parkinson's disease (PD) through an English pangram utterance speech collected using a web application from diverse recording settings and environments, including participants' homes. Our dataset includes a global cohort of 1306 participants, including 392 diagnosed with PD. Leveraging the diversity of the dataset, spanning various demographic properties (such as age, sex, and ethnicity), we used deep learning embeddings derived from semi-supervised models such as Wav2Vec 2.0, WavLM, and ImageBind representing the speech dynamics associated with PD. Our novel fusion model for PD classification, which aligns different speech embeddings into a cohesive feature space, demonstrated superior performance over standard concatenation-based fusion models and other baselines (including models built on traditional acoustic features). In a randomized data split configuration, the model achieved an Area Under the Receiver Operating Characteristic Curve (AUROC) of 88.94% and an accuracy of 85.65%. Rigorous statistical analysis confirmed that our model performs equitably across various demographic subgroups in terms of sex, ethnicity, and age, and remains robust regardless of disease duration. Furthermore, our model, when tested on two entirely unseen test datasets collected from clinical settings and from a PD care center, maintained AUROC scores of 82.12% and 78.44%, respectively. This affirms the model's robustness and it's potential to enhance accessibility and health equity in real-world applications.
【3】 Sok: Comprehensive Security Overview, Challenges, and Future Directions of Voice-Controlled Systems
标题: Sok:语音控制系统的全面安全概述、挑战和未来方向
作者:Haozhe Xu,Cong Wu,Yangyang Gu,Xingcan Shang,Jing Chen,Kun He,Ruiying Du
链接:点击下载PDF文件
摘要:语音控制系统(VoIP)与智能设备的集成及其在日常生活中的日益增长,突出了其安全性的重要性。目前的研究已经发现了许多漏洞,对用户隐私和安全构成了重大风险。然而,仍然没有对这些脆弱性和相应的解决办法进行连贯和系统的审查。这种缺乏全面分析的情况,给IBM的设计人员在充分理解和缓解这些系统中的安全问题方面带来了挑战。 为了解决这一差距,我们的研究引入了一个层次模型结构的分类,并以系统的方式分析现有的文献提供了一个新的镜头。我们根据攻击的技术原理对攻击进行分类,并彻底评估各种属性,例如攻击方法、目标、向量和行为。此外,我们巩固和评估当前研究中提出的防御机制,为增强网络安全提供可操作的建议。我们的工作通过简化网络安全固有的复杂性,帮助设计人员有效地识别和应对潜在威胁,并为网络安全研究的未来发展奠定基础,做出了重大贡献。摘要:The integration of Voice Control Systems (VCS) into smart devices and their growing presence in daily life accentuate the importance of their security. Current research has uncovered numerous vulnerabilities in VCS, presenting significant risks to user privacy and security. However, a cohesive and systematic examination of these vulnerabilities and the corresponding solutions is still absent. This lack of comprehensive analysis presents a challenge for VCS designers in fully understanding and mitigating the security issues within these systems. Addressing this gap, our study introduces a hierarchical model structure for VCS, providing a novel lens for categorizing and analyzing existing literature in a systematic manner. We classify attacks based on their technical principles and thoroughly evaluate various attributes, such as their methods, targets, vectors, and behaviors. Furthermore, we consolidate and assess the defense mechanisms proposed in current research, offering actionable recommendations for enhancing VCS security. Our work makes a significant contribution by simplifying the complexity inherent in VCS security, aiding designers in effectively identifying and countering potential threats, and setting a foundation for future advancements in VCS security research.
【4】 RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
标题: RSET:基于重命名的情感转移语音合成排序方法
作者:Haoxiang Shi,Jianzong Wang,Xulong Zhang,Ning Cheng,Jun Yu,Jing Xiao
备注:Accepted by the 8th APWeb-WAIM International Joint Conference on Web and Big Data
链接:点击下载PDF文件
摘要:虽然目前的文语转换模型能够产生高质量的语音样本,但在开发情感强度可控的文语转换方面仍然存在挑战。现有的TTS模型大多通过从参考语音中提取强度信息来实现情感强度控制。然而,由于缺乏类内情感强度的建模和模型的信息解耦能力,生成的语音不能实现细粒度的情感强度控制,存在信息泄漏问题。本文提出了一种情感转移TTS模型,该模型定义了一种基于重映射的排序方法来建模类内相对强度信息,结合互信息(MI)来解耦说话人和情感信息,并合成具有可感知强度差异的表达性语音。实验表明,该模型在保持说话人信息的同时,实现了细粒度的情感控制。摘要:Although current Text-To-Speech (TTS) models are able to generate high-quality speech samples, there are still challenges in developing emotion intensity controllable TTS. Most existing TTS models achieve emotion intensity control by extracting intensity information from reference speeches. Unfortunately, limited by the lack of modeling for intra-class emotion intensity and the model's information decoupling capability, the generated speech cannot achieve fine-grained emotion intensity control and suffers from information leakage issues. In this paper, we propose an emotion transfer TTS model, which defines a remapping-based sorting method to model intra-class relative intensity information, combined with Mutual Information (MI) to decouple speaker and emotion information, and synthesizes expressive speeches with perceptible intensity differences. Experiments show that our model achieves fine-grained emotion control while preserving speaker information.
【5】 A Real-Time Voice Activity Detection Based On Lightweight Neural
标题: 基于轻量级神经网络的实时语音活动检测
作者:Jidong Jia,Pei Zhao,Di Wang
链接:点击下载PDF文件
摘要:语音活动检测(VAD)是在音频流中检测语音的任务,由于现实环境中存在大量不可见的噪声和低信噪比,因此具有挑战性。最近,基于神经网络的VAD在一定程度上缓解了性能的下降。然而,大多数现有的研究都采用了过大的模型,并纳入未来的背景下,而忽略了评估模型的运行效率和延迟。在本文中,我们提出了一种称为MagicNet的轻量级实时神经网络,它利用了因果和深度可分离的1-D卷积和GRU。在不依赖于未来的功能作为输入,我们提出的模型进行了比较,与两个国家的最先进的算法合成域内外测试数据集。评估结果表明,MagicNet可以实现更好的性能和鲁棒性与更少的参数成本。摘要:Voice activity detection (VAD) is the task of detecting speech in an audio stream, which is challenging due to numerous unseen noises and low signal-to-noise ratios in real environments. Recently, neural network-based VADs have alleviated the degradation of performance to some extent. However, the majority of existing studies have employed excessively large models and incorporated future context, while neglecting to evaluate the operational efficiency and latency of the models. In this paper, we propose a lightweight and real-time neural network called MagicNet, which utilizes casual and depth separable 1-D convolutions and GRU. Without relying on future features as input, our proposed model is compared with two state-of-the-art algorithms on synthesized in-domain and out-domain test datasets. The evaluation results demonstrate that MagicNet can achieve improved performance and robustness with fewer parameter costs.
【6】 Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
标题: 使用音频到配乐自动转录管道重建Charlie Parker Omnebook
作者:Xavier Riley,Simon Dixon
链接:点击下载PDF文件
摘要:《查理·帕克百科全书》是爵士音乐教育的基石,钢琴家伊森·艾弗森将其描述为“有史以来出版的最重要的爵士乐教育文本”。在这项工作中,我们提出了一个新的转录管道,并探讨在何种程度上的最先进的音乐技术能够重建这些分数直接从音频没有人为干预。我们的管道包括:新训练的萨克斯管的源分离模型、新的独奏萨克斯管的MIDI转录模型以及现有的单声道乐器的MIDI到乐谱方法的适应。 为了评估这个管道,我们还提供了一个增强的数据集查理帕克transmittance作为分数音频对与准确的双对齐和强拍注释。这代表了自动音频到乐谱转录的一个具有挑战性的新基准,我们希望这将推动研究超越单独转录音频到乐谱的领域。 总之,这些形成了另一个步骤,生产乐谱,音乐家可以直接使用,而不需要繁琐的更正或修改。为了方便未来的研究,所有模型检查点和数据都可以与转录管道的代码一起下载。我们的模块化管道的改进可能有一天使复杂的爵士乐独奏的自动转录成为一种常规的可能性,从而丰富了音乐教育和保存的可用资源。摘要:The Charlie Parker Omnibook is a cornerstone of jazz music education, described by pianist Ethan Iverson as "the most important jazz education text ever published". In this work we propose a new transcription pipeline and explore the extent to which state of the art music technology is able to reconstruct these scores directly from the audio without human intervention. Our pipeline includes: a newly trained source separation model for saxophone, a new MIDI transcription model for solo saxophone and an adaptation of an existing MIDI-to-score method for monophonic instruments. To assess this pipeline we also provide an enhanced dataset of Charlie Parker transcriptions as score-audio pairs with accurate MIDI alignments and downbeat annotations. This represents a challenging new benchmark for automatic audio-to-score transcription that we hope will advance research into areas beyond transcribing audio-to-MIDI alone. Together, these form another step towards producing scores that musicians can use directly, without the need for onerous corrections or revisions. To facilitate future research, all model checkpoints and data are made available to download along with code for the transcription pipeline. Improvements in our modular pipeline could one day make the automatic transcription of complex jazz solos a routine possibility, thereby enriching the resources available for music education and preservation.
【7】 C3LLM: Conditional Multimodal Content Generation Using Large Language Models
标题: C3 LLM:使用大型语言模型的条件多模式内容生成
作者:Zixuan Wang,Qinkai Duan,Yu-Wing Tai,Chi-Keung Tang
链接:点击下载PDF文件
摘要:我们介绍了C3 LLM(Conditioned-on-Three-Modalities Large Language Models),这是一个将视频到音频,音频到文本和文本到音频三个任务结合在一起的新框架。C3 LLM采用大型语言模型(LLM)结构作为桥梁,用于对齐不同的模态,合成给定的条件信息,并以离散的方式进行多模态生成。我们的贡献如下。首先,我们使用预先训练的音频码本为音频生成任务调整了分层结构。具体来说,我们训练LLM以从给定条件生成音频语义令牌,并且进一步使用非自回归Transformer来生成分层中的不同级别的声学令牌,以更好地增强所生成的音频的保真度。其次,基于LLM最初是为下一个单词预测方法的离散任务而设计的直觉,我们使用离散表示来生成音频,并将其语义压缩到声学标记中,类似于向LLM添加“声学词汇”。第三,我们的方法将之前的音频理解,视频到音频生成和文本到音频生成的任务结合到一个统一的模型中,以端到端的方式提供更多的通用性。我们的C3 LLM通过各种自动化评估指标实现了改进的结果,与以前的方法相比,提供了更好的语义对齐。摘要:We introduce C3LLM (Conditioned-on-Three-Modalities Large Language Models), a novel framework combining three tasks of video-to-audio, audio-to-text, and text-to-audio together. C3LLM adapts the Large Language Model (LLM) structure as a bridge for aligning different modalities, synthesizing the given conditional information, and making multimodal generation in a discrete manner. Our contributions are as follows. First, we adapt a hierarchical structure for audio generation tasks with pre-trained audio codebooks. Specifically, we train the LLM to generate audio semantic tokens from the given conditions, and further use a non-autoregressive transformer to generate different levels of acoustic tokens in layers to better enhance the fidelity of the generated audio. Second, based on the intuition that LLMs were originally designed for discrete tasks with the next-word prediction method, we use the discrete representation for audio generation and compress their semantic meanings into acoustic tokens, similar to adding "acoustic vocabulary" to LLM. Third, our method combines the previous tasks of audio understanding, video-to-audio generation, and text-to-audio generation together into one unified model, providing more versatility in an end-to-end fashion. Our C3LLM achieves improved results through various automated evaluation metrics, providing better semantic alignment compared to previous methods.
【8】 Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
标题: 使用严格延时神经网络的Carnatic Raga识别系统
作者:Sanjay Natesan,Homayoon Beigi
备注:7 pages, 2 tables, 3 figures
链接:点击下载PDF文件
摘要:基于大规模机器学习的Raga识别仍然是Carnatic音乐背后计算方面的一个重要问题。每一个拉加都有许多独特的内在旋律模式,可以用来很容易地将它们与其他的旋律模式区分开来。这些raga也可以用来聚类同一raga中的歌曲,以及识别其他密切相关的raga中的歌曲。在这种情况下,使用包括使用离散傅里叶变换和使用三角滤波的步骤的组合来分析输入声音,以创建可能音符的定制区间,从特定音符的存在或缺乏中提取特征。使用神经网络的组合,包括一维卷积神经网络(通常称为时延神经网络)和长短期记忆(LSTM),这是一种形式的递归神经网络,可以创建分类策略的骨干来构建模型。此外,为了帮助shruti的变化,将实施一种基于长时间注意力的机制来确定频率的相对变化,而不是绝对差异。这将在训练不同灌木中的音频片段时提供更有意义的数据点。为了评估分类器的准确性,使用了676个记录的数据集。这些歌曲分布在拉格斯的列表中。这个程序的目标是能够有效和高效地标记更广泛的音频片段,包括更多的shrutis,ragas和更多的背景噪音。摘要:Large scale machine learning-based Raga identification continues to be a nontrivial issue in the computational aspects behind Carnatic music. Each raga consists of many unique and intrinsic melodic patterns that can be used to easily identify them from others. These ragas can also then be used to cluster songs within the same raga, as well as identify songs in other closely related ragas. In this case, the input sound is analyzed using a combination of steps including using a Discrete Fourier transformation and using Triangular Filtering to create custom bins of possible notes, extracting features from the presence of particular notes or lack thereof. Using a combination of Neural Networks including 1D Convolutional Neural Networks conventionally known as Time-Delay Neural Networks) and Long Short-Term Memory (LSTM), which are a form of Recurrent Neural Networks, the backbone of the classification strategy to build the model can be created. In addition, to help with variations in shruti, a long-time attention-based mechanism will be implemented to determine the relative changes in frequency rather than the absolute differences. This will provide a much more meaningful data point when training audio clips in different shrutis. To evaluate the accuracy of the classifier, a dataset of 676 recordings is used. The songs are distributed across the list of ragas. The goal of this program is to be able to effectively and efficiently label a much wider range of audio clips in more shrutis, ragas, and with more background noise.
【9】 Quality-aware Masked Diffusion Transformer for Enhanced Music Generation
标题: 用于增强音乐生成的质量感知掩蔽扩散Transformer
作者:Chang Li,Ruoyu Wang,Lijuan Liu,Jun Du,Yixuan Sun,Zilu Guo,Zhenrong Zhang,Yuan Jiang
链接:点击下载PDF文件
摘要:近年来,基于扩散的文本到音乐(TTM)生成已经获得了突出,提供了一种新的方法来合成音乐内容从文本描述。在这一生成过程中实现高准确性和多样性需要大量高质量的数据,而这些数据往往只占可用数据集的一小部分。在开源数据集中,错误标记、弱标记、未标记数据和低质量音乐波形等问题的普遍存在严重阻碍了音乐生成模型的开发。为了克服这些挑战,我们引入了一种新的质量感知掩蔽扩散Transformer(QA-MDT)方法,使生成模型能够在训练过程中识别输入音乐波形的质量。基于音乐信号的独特属性,我们已经适应并实现了用于TTM任务的MDT模型,同时进一步揭示了其独特的质量控制能力。此外,我们解决了低质量的字幕与字幕细化数据处理方法的问题。我们的演示页面显示在https: qa-mdt.github.io 。代码https: github.com ivcylc qa-mdt摘要:In recent years, diffusion-based text-to-music (TTM) generation has gained prominence, offering a novel approach to synthesizing musical content from textual descriptions. Achieving high accuracy and diversity in this generation process requires extensive, high-quality data, which often constitutes only a fraction of available datasets. Within open-source datasets, the prevalence of issues like mislabeling, weak labeling, unlabeled data, and low-quality music waveform significantly hampers the development of music generation models. To overcome these challenges, we introduce a novel quality-aware masked diffusion transformer (QA-MDT) approach that enables generative models to discern the quality of input music waveform during training. Building on the unique properties of musical signals, we have adapted and implemented a MDT model for TTM task, while further unveiling its distinct capacity for quality control. Moreover, we address the issue of low-quality captions with a caption refinement data processing approach. Our demo page is shown in https: qa-mdt.github.io . Code on https: github.com ivcylc qa-mdt
【10】 Crossmodal ASR Error Correction with Discrete Speech Units
标题: 使用离散语音单元的交叉模式ASB误差纠正
作者:Yuanchao Li,Pinzhen Chen,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:在说话风格与用于训练ASR系统的说话风格不同的情况下,ASR仍然不能令人满意,从而导致错误的转录。为了解决这个问题,需要ASR纠错(AEC),一种ASR后处理方法。在这项工作中,我们解决了一个未充分研究的问题:低资源域外(LROOD)问题,通过调查跨模态AEC非常有限的下游数据与1-最好的假设转录。我们探索了预训练和微调策略,并发现了ASR域差异现象,为LROOD数据提供了适当的训练方案。此外,我们建议将离散的语音单元对齐,并加强字嵌入,以提高AEC质量。多个语料库和多个评价指标的结果表明,我们提出的AEC方法的LROOD数据的可行性和有效性,以及它的普遍性和大规模数据的优越性。最后,语音情感识别的研究证实,我们的模型产生ASR错误鲁棒的成绩单适合下游应用。摘要:ASR remains unsatisfactory in scenarios where the speaking style diverges from that used to train ASR systems, resulting in erroneous transcripts. To address this, ASR Error Correction (AEC), a post-ASR processing approach, is required. In this work, we tackle an understudied issue: the Low-Resource Out-of-Domain (LROOD) problem, by investigating crossmodal AEC on very limited downstream data with 1-best hypothesis transcription. We explore pre-training and fine-tuning strategies and uncover an ASR domain discrepancy phenomenon, shedding light on appropriate training schemes for LROOD data. Moreover, we propose the incorporation of discrete speech units to align with and enhance the word embeddings for improving AEC quality. Results from multiple corpora and several evaluation metrics demonstrate the feasibility and efficacy of our proposed AEC approach on LROOD data, as well as its generalizability and superiority on large-scale data. Finally, a study on speech emotion recognition confirms that our model produces ASR error-robust transcripts suitable for downstream applications.
【11】 Spiketrum: An FPGA-based Implementation of a Neuromorphic Cochlea
标题: Spiketrum:基于FPGA的神经形态可卡因实现
作者:MHD Anas Alsakkal,Jayawan Wijekoon
备注:Submitted to "IEEE Transactions on Circuits and Systems"
链接:点击下载PDF文件
摘要:本文提出了一种新的基于FPGA的神经形态耳蜗,利用通用的尖峰编码算法,Spiketrum。本研究的重点是这个耳蜗模型的发展和表征,它擅长将音频振动转化为生物学上逼真的听觉尖峰序列。这些尖峰序列旨在承受神经波动和尖峰损失,同时准确地封装音频的空间和精确的时间特征,以及传入振动的强度。值得注意的功能包括以最小的信息损失生成实时尖峰序列的能力和重建原始信号的能力。这种微调功能允许用户优化尖峰速率,从而在输出质量和功耗之间实现最佳平衡。此外,将反馈系统集成到Spiketrum中可以选择性地放大特定功能,同时衰减其他功能,从而促进基于应用要求的自适应功耗。硬件实现支持基于尖峰和基于非尖峰的处理器,使其通用于各种计算系统。耳蜗的能力,编码不同的感官信息,延伸到声音波形,定位它作为一个有前途的感官输入,为当前和未来的尖峰为基础的智能计算系统,提供紧凑和实时尖峰训练生成。摘要:This paper presents a novel FPGA-based neuromorphic cochlea, leveraging the general-purpose spike-coding algorithm, Spiketrum. The focus of this study is on the development and characterization of this cochlea model, which excels in transforming audio vibrations into biologically realistic auditory spike trains. These spike trains are designed to withstand neural fluctuations and spike losses while accurately encapsulating the spatial and precise temporal characteristics of audio, along with the intensity of incoming vibrations. Noteworthy features include the ability to generate real-time spike trains with minimal information loss and the capacity to reconstruct original signals. This fine-tuning capability allows users to optimize spike rates, achieving an optimal balance between output quality and power consumption. Furthermore, the integration of a feedback system into Spiketrum enables selective amplification of specific features while attenuating others, facilitating adaptive power consumption based on application requirements. The hardware implementation supports both spike-based and non-spike-based processors, making it versatile for various computing systems. The cochlea's ability to encode diverse sensory information, extending beyond sound waveforms, positions it as a promising sensory input for current and future spike-based intelligent computing systems, offering compact and real-time spike train generation.
eess.AS音频处理
【1】 Speech Loudness in Broadcasting and Streaming标题: 广播和流媒体中的语音响度
作者:Matteo Torcoli,Mhd Modar Halimeh,Thomas Leitz,Yannik Grewe,Michael Kratschmer,Bernhard Neugebauer,Adrian Murtaza,Harald Fuchs,Emanuël A. P. Habets
备注:Accepted for presentation at the Audio Engineering Society (AES) 156th Convention, June 2024, Madrid, Spain
链接:点击下载PDF文件
摘要:在广播和流媒体中引入和调节响度给观众带来了明显的好处,例如,节目和频道之间的一致性水平。然而,在某些段落中,经常报告言语响度太低,这可能会妨碍对电影和电视节目的充分理解和欣赏。本文建议扩大该行业通常使用的基于质量的措施集。我们专注于语音响度,我们表明,当干净的语音不可用时,深度神经网络(DNN)可用于隔离语音信号,从而准确估计语音响度,与语音门控响度相比,提供更精确的估计。此外,我们定义了关键通道,即,演讲可能难以理解的段落。关键段落基于局部语音响度偏差(SLD)和局部语音与背景响度差(SBLD)来定义,因为SLD和SBLD显著地有助于可懂度和收听努力。与其他更全面的可懂度和听力测量相比,SLD和SBLD可以直接测量,直观,最重要的是,可以通过调整混音中的语音电平或在用户端实现个性化来轻松控制。最后,提供的例子表明,关键段落的检测可以支持内容制作期间和之后的语音信号的评估和控制。摘要:The introduction and regulation of loudness in broadcasting and streaming brought clear benefits to the audience, e.g., a level of uniformity across programs and channels. Yet, speech loudness is frequently reported as being too low in certain passages, which can hinder the full understanding and enjoyment of movies and TV programs. This paper proposes expanding the set of loudness-based measures typically used in the industry. We focus on speech loudness, and we show that, when clean speech is not available, Deep Neural Networks (DNNs) can be used to isolate the speech signal and so to accurately estimate speech loudness, providing a more precise estimate compared to speech-gated loudness. Moreover, we define critical passages, i.e., passages in which speech is likely to be hard to understand. Critical passages are defined based on the local Speech Loudness Deviation (SLD) and the local Speech-to-Background Loudness Difference (SBLD), as SLD and SBLD significantly contribute to intelligibility and listening effort. In contrast to other more comprehensive measures of intelligibility and listening effort, SLD and SBLD can be straightforwardly measured, are intuitive, and, most importantly, can be easily controlled by adjusting the speech level in the mix or by enabling personalization at the user's end. Finally, examples are provided that show how the detection of critical passages can support the evaluation and control of the speech signal during and after content production.
【2】 A Variance-Preserving Interpolation Approach for Diffusion Models with Applications to Single Channel Speech Enhancement and Recognition
标题: 扩散模型的保方差插值方法及其在单通道语音增强和识别中的应用
作者:Zilu Guo,Qing Wang,Jun Du,Jia Pan,Qing-Feng Liu,Chin-Hui
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个方差保持插值框架,以改善单通道语音增强(SE)和自动语音识别(ASR)的扩散模型。这种新的方差保持插值扩散模型(VPIDM)的方法只需要25个迭代步骤,并避免了需要一个校正器,在现有的方差爆炸插值扩散模型(VEIDM)的一个重要元素。VPIDM和VEIDM之间的两个显着区别是状态变量的平均值的标度函数和相对于平均值的标度对方差施加的约束。我们对VPIDM的理论机制进行了系统的探索,并以VPIDM为前端,深入研究了VPIDM在SE和ASR中的应用。我们提出的方法,在两个不同的数据集上进行评估,证明了VPIDM的优越性能比传统的歧视SE算法。此外,我们评估所提出的模型在不同的信噪比(SNR)水平下的性能。研究表明,VPIDM的目标噪声消除的鲁棒性相比,VEIDM。此外,利用VPIDM和VEIDM的中间输出,提高了ASR的准确性,从而突出了我们提出的方法的实际效果。摘要:In this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpolation diffusion model (VPIDM) approach requires only 25 iterative steps and obviates the need for a corrector, an essential element in the existing variance-exploding interpolation diffusion model (VEIDM). Two notable distinctions between VPIDM and VEIDM are the scaling function of the mean of state variables and the constraint imposed on the variance relative to the mean's scale. We conduct a systematic exploration of the theoretical mechanism underlying VPIDM and develop insights regarding VPIDM's applications in SE and ASR using VPIDM as a frontend. Our proposed approach, evaluated on two distinct data sets, demonstrates VPIDM's superior performances over conventional discriminative SE algorithms. Furthermore, we assess the performance of the proposed model under varying signal-to-noise ratio (SNR) levels. The investigation reveals VPIDM's improved robustness in target noise elimination when compared to VEIDM. Furthermore, utilizing the mid-outputs of both VPIDM and VEIDM results in enhanced ASR accuracies, thereby highlighting the practical efficacy of our proposed approach.
【3】 Speech enhancement deep-learning architecture for efficient edge processing
标题: 语音增强深度学习架构可实现高效边缘处理
作者:Monisankha Pal,Arvind Ramanathan,Ted Wada,Ashutosh Pandey
链接:点击下载PDF文件
摘要:深度学习已经成为语音增强任务的实际选择方法,语音质量得到了显着改善。然而,针对低功率边缘设备的具有减小的尺寸和计算的实时处理显著地降低了语音质量。最近,基于transformer的体系结构大大降低了内存需求,并提供了通过局部和全局上下文来提高模型性能的方法。然而,Transformer操作在计算上仍然是繁重的。在这项工作中,我们引入了基于WaveUNet挤压激励Res 2(WSR)的度量生成对抗网络(WSR-MGAN)架构,该架构可以在低功耗边缘设备上有效地实现噪声抑制任务,同时保持语音质量。我们使用Res 2Net块来利用多尺度特征,这些块可以与语音处理任务中使用的频谱内容相关。在生成器中,我们集成了挤压激励块(SEB)与多尺度功能,用于维护局部和全局上下文以及门控递归单元(GRU)。所提出的方法是通过一个组合的损失函数计算的原始波形,多分辨率的幅度谱图,和客观的指标,使用一个度量矩阵进行优化。在VoiceBank+DEMAND和DNS-2020挑战数据集上的各种客观指标的实验结果表明,所提出的语音增强(SE)方法优于基线,并在时域上实现了最先进的(SOTA)性能。摘要:Deep learning has become a de facto method of choice for speech enhancement tasks with significant improvements in speech quality. However, real-time processing with reduced size and computations for low-power edge devices drastically degrades speech quality. Recently, transformer-based architectures have greatly reduced the memory requirements and provided ways to improve the model performance through local and global contexts. However, the transformer operations remain computationally heavy. In this work, we introduce WaveUNet squeeze-excitation Res2 (WSR)-based metric generative adversarial network (WSR-MGAN) architecture that can be efficiently implemented on low-power edge devices for noise suppression tasks while maintaining speech quality. We utilize multi-scale features using Res2Net blocks that can be related to spectral content used in speech-processing tasks. In the generator, we integrate squeeze-excitation blocks (SEB) with multi-scale features for maintaining local and global contexts along with gated recurrent units (GRUs). The proposed approach is optimized through a combined loss function calculated over raw waveform, multi-resolution magnitude spectrogram, and objective metrics using a metric discriminator. Experimental results in terms of various objective metrics on VoiceBank+DEMAND and DNS-2020 challenge datasets demonstrate that the proposed speech enhancement (SE) approach outperforms the baselines and achieves state-of-the-art (SOTA) performance in the time domain.
【4】 Crossmodal ASR Error Correction with Discrete Speech Units
标题: 使用离散语音单元的交叉模式ASB误差纠正
作者:Yuanchao Li,Pinzhen Chen,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:在说话风格与用于训练ASR系统的说话风格不同的情况下,ASR仍然不能令人满意,从而导致错误的转录。为了解决这个问题,需要ASR纠错(AEC),一种ASR后处理方法。在这项工作中,我们解决了一个未充分研究的问题:低资源域外(LROOD)问题,通过调查跨模态AEC非常有限的下游数据与1-最好的假设转录。我们探索了预训练和微调策略,并发现了ASR域差异现象,为LROOD数据提供了适当的训练方案。此外,我们建议将离散的语音单元对齐,并加强字嵌入,以提高AEC质量。多个语料库和多个评价指标的结果表明,我们提出的AEC方法的LROOD数据的可行性和有效性,以及它的普遍性和大规模数据的优越性。最后,语音情感识别的研究证实,我们的模型产生ASR错误鲁棒的成绩单适合下游应用。摘要:ASR remains unsatisfactory in scenarios where the speaking style diverges from that used to train ASR systems, resulting in erroneous transcripts. To address this, ASR Error Correction (AEC), a post-ASR processing approach, is required. In this work, we tackle an understudied issue: the Low-Resource Out-of-Domain (LROOD) problem, by investigating crossmodal AEC on very limited downstream data with 1-best hypothesis transcription. We explore pre-training and fine-tuning strategies and uncover an ASR domain discrepancy phenomenon, shedding light on appropriate training schemes for LROOD data. Moreover, we propose the incorporation of discrete speech units to align with and enhance the word embeddings for improving AEC quality. Results from multiple corpora and several evaluation metrics demonstrate the feasibility and efficacy of our proposed AEC approach on LROOD data, as well as its generalizability and superiority on large-scale data. Finally, a study on speech emotion recognition confirms that our model produces ASR error-robust transcripts suitable for downstream applications.
【5】 Spiketrum: An FPGA-based Implementation of a Neuromorphic Cochlea
标题: Spiketrum:基于FPGA的神经形态可卡因实现
作者:MHD Anas Alsakkal,Jayawan Wijekoon
备注:Submitted to "IEEE Transactions on Circuits and Systems"
链接:点击下载PDF文件
摘要:本文提出了一种新的基于FPGA的神经形态耳蜗,利用通用的尖峰编码算法,Spiketrum。本研究的重点是这个耳蜗模型的发展和表征,它擅长将音频振动转化为生物学上逼真的听觉尖峰序列。这些尖峰序列旨在承受神经波动和尖峰损失,同时准确地封装音频的空间和精确的时间特征,以及传入振动的强度。值得注意的功能包括以最小的信息损失生成实时尖峰序列的能力和重建原始信号的能力。这种微调功能允许用户优化尖峰速率,从而在输出质量和功耗之间实现最佳平衡。此外,将反馈系统集成到Spiketrum中可以选择性地放大特定功能,同时衰减其他功能,从而促进基于应用要求的自适应功耗。硬件实现支持基于尖峰和基于非尖峰的处理器,使其通用于各种计算系统。耳蜗的能力,编码不同的感官信息,延伸到声音波形,定位它作为一个有前途的感官输入,为当前和未来的尖峰为基础的智能计算系统,提供紧凑和实时尖峰训练生成。摘要:This paper presents a novel FPGA-based neuromorphic cochlea, leveraging the general-purpose spike-coding algorithm, Spiketrum. The focus of this study is on the development and characterization of this cochlea model, which excels in transforming audio vibrations into biologically realistic auditory spike trains. These spike trains are designed to withstand neural fluctuations and spike losses while accurately encapsulating the spatial and precise temporal characteristics of audio, along with the intensity of incoming vibrations. Noteworthy features include the ability to generate real-time spike trains with minimal information loss and the capacity to reconstruct original signals. This fine-tuning capability allows users to optimize spike rates, achieving an optimal balance between output quality and power consumption. Furthermore, the integration of a feedback system into Spiketrum enables selective amplification of specific features while attenuating others, facilitating adaptive power consumption based on application requirements. The hardware implementation supports both spike-based and non-spike-based processors, making it versatile for various computing systems. The cochlea's ability to encode diverse sensory information, extending beyond sound waveforms, positions it as a promising sensory input for current and future spike-based intelligent computing systems, offering compact and real-time spike train generation.
【6】 Enhancing Music Genre Classification through Multi-Algorithm Analysis and User-Friendly Visualization
标题: 通过多算法分析和用户友好的可视化增强音乐流派分类
作者:Navin Kamuni,Dheerendra Panwar
链接:点击下载PDF文件
摘要:这项研究的目的是教算法如何识别不同类型的音乐。用户将提交歌曲进行分析。由于算法以前从未听过这些歌曲,因此它需要弄清楚是什么让每首歌曲都独一无二。它通过将歌曲分解为不同的部分,并通过监督学习来学习节奏,旋律和音调等内容,因为程序从已经标记的示例中学习。分类音乐时要考虑的一个重要因素是它的类型,这可能相当复杂。为了确保准确性,我们使用了五种不同的算法来分析歌曲,每种算法都独立工作。这有助于我们更全面地了解每首歌的特点。因此,我们的目标是正确识别每首提交歌曲的类型。分析完成后,结果将使用图形工具呈现,便于用户理解和提供反馈。摘要:The aim of this study is to teach an algorithm how to recognize different types of music. Users will submit songs for analysis. Since the algorithm hasn't heard these songs before, it needs to figure out what makes each song unique. It does this by breaking down the songs into different parts and studying things like rhythm, melody, and tone via supervised learning because the program learns from examples that are already labelled. One important thing to consider when classifying music is its genre, which can be quite complex. To ensure accuracy, we use five different algorithms, each working independently, to analyze the songs. This helps us get a more complete understanding of each song's characteristics. Therefore, our goal is to correctly identify the genre of each submitted song. Once the analysis is done, the results are presented using a graphing tool, making it easy for users to understand and provide feedback.
【7】 Sok: Comprehensive Security Overview, Challenges, and Future Directions of Voice-Controlled Systems
标题: Sok:语音控制系统的全面安全概述、挑战和未来方向
作者:Haozhe Xu,Cong Wu,Yangyang Gu,Xingcan Shang,Jing Chen,Kun He,Ruiying Du
链接:点击下载PDF文件
摘要:语音控制系统(VoIP)与智能设备的集成及其在日常生活中的日益增长,突出了其安全性的重要性。目前的研究已经发现了许多漏洞,对用户隐私和安全构成了重大风险。然而,仍然没有对这些脆弱性和相应的解决办法进行连贯和系统的审查。这种缺乏全面分析的情况,给IBM的设计人员在充分理解和缓解这些系统中的安全问题方面带来了挑战。 为了解决这一差距,我们的研究引入了一个层次模型结构的分类,并以系统的方式分析现有的文献提供了一个新的镜头。我们根据攻击的技术原理对攻击进行分类,并彻底评估各种属性,例如攻击方法、目标、向量和行为。此外,我们巩固和评估当前研究中提出的防御机制,为增强网络安全提供可操作的建议。我们的工作通过简化网络安全固有的复杂性,帮助设计人员有效地识别和应对潜在威胁,并为网络安全研究的未来发展奠定基础,做出了重大贡献。摘要:The integration of Voice Control Systems (VCS) into smart devices and their growing presence in daily life accentuate the importance of their security. Current research has uncovered numerous vulnerabilities in VCS, presenting significant risks to user privacy and security. However, a cohesive and systematic examination of these vulnerabilities and the corresponding solutions is still absent. This lack of comprehensive analysis presents a challenge for VCS designers in fully understanding and mitigating the security issues within these systems. Addressing this gap, our study introduces a hierarchical model structure for VCS, providing a novel lens for categorizing and analyzing existing literature in a systematic manner. We classify attacks based on their technical principles and thoroughly evaluate various attributes, such as their methods, targets, vectors, and behaviors. Furthermore, we consolidate and assess the defense mechanisms proposed in current research, offering actionable recommendations for enhancing VCS security. Our work makes a significant contribution by simplifying the complexity inherent in VCS security, aiding designers in effectively identifying and countering potential threats, and setting a foundation for future advancements in VCS security research.
【8】 RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
标题: RSET:基于重命名的情感转移语音合成排序方法
作者:Haoxiang Shi,Jianzong Wang,Xulong Zhang,Ning Cheng,Jun Yu,Jing Xiao
备注:Accepted by the 8th APWeb-WAIM International Joint Conference on Web and Big Data
链接:点击下载PDF文件
摘要:虽然目前的文语转换模型能够产生高质量的语音样本,但在开发情感强度可控的文语转换方面仍然存在挑战。现有的TTS模型大多通过从参考语音中提取强度信息来实现情感强度控制。然而,由于缺乏类内情感强度的建模和模型的信息解耦能力,生成的语音不能实现细粒度的情感强度控制,存在信息泄漏问题。本文提出了一种情感转移TTS模型,该模型定义了一种基于重映射的排序方法来建模类内相对强度信息,结合互信息(MI)来解耦说话人和情感信息,并合成具有可感知强度差异的表达性语音。实验表明,该模型在保持说话人信息的同时,实现了细粒度的情感控制。摘要:Although current Text-To-Speech (TTS) models are able to generate high-quality speech samples, there are still challenges in developing emotion intensity controllable TTS. Most existing TTS models achieve emotion intensity control by extracting intensity information from reference speeches. Unfortunately, limited by the lack of modeling for intra-class emotion intensity and the model's information decoupling capability, the generated speech cannot achieve fine-grained emotion intensity control and suffers from information leakage issues. In this paper, we propose an emotion transfer TTS model, which defines a remapping-based sorting method to model intra-class relative intensity information, combined with Mutual Information (MI) to decouple speaker and emotion information, and synthesizes expressive speeches with perceptible intensity differences. Experiments show that our model achieves fine-grained emotion control while preserving speaker information.
【9】 A Real-Time Voice Activity Detection Based On Lightweight Neural
标题: 基于轻量级神经网络的实时语音活动检测
作者:Jidong Jia,Pei Zhao,Di Wang
链接:点击下载PDF文件
摘要:语音活动检测(VAD)是在音频流中检测语音的任务,由于现实环境中存在大量不可见的噪声和低信噪比,因此具有挑战性。最近,基于神经网络的VAD在一定程度上缓解了性能的下降。然而,大多数现有的研究都采用了过大的模型,并纳入未来的背景下,而忽略了评估模型的运行效率和延迟。在本文中,我们提出了一种称为MagicNet的轻量级实时神经网络,它利用了因果和深度可分离的1-D卷积和GRU。在不依赖于未来的功能作为输入,我们提出的模型进行了比较,与两个国家的最先进的算法合成域内外测试数据集。评估结果表明,MagicNet可以实现更好的性能和鲁棒性与更少的参数成本。摘要:Voice activity detection (VAD) is the task of detecting speech in an audio stream, which is challenging due to numerous unseen noises and low signal-to-noise ratios in real environments. Recently, neural network-based VADs have alleviated the degradation of performance to some extent. However, the majority of existing studies have employed excessively large models and incorporated future context, while neglecting to evaluate the operational efficiency and latency of the models. In this paper, we propose a lightweight and real-time neural network called MagicNet, which utilizes casual and depth separable 1-D convolutions and GRU. Without relying on future features as input, our proposed model is compared with two state-of-the-art algorithms on synthesized in-domain and out-domain test datasets. The evaluation results demonstrate that MagicNet can achieve improved performance and robustness with fewer parameter costs.
【10】 Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
标题: 使用音频到配乐自动转录管道重建Charlie Parker Omnebook
作者:Xavier Riley,Simon Dixon
链接:点击下载PDF文件
摘要:《查理·帕克百科全书》是爵士音乐教育的基石,钢琴家伊森·艾弗森将其描述为“有史以来出版的最重要的爵士乐教育文本”。在这项工作中,我们提出了一个新的转录管道,并探讨在何种程度上的最先进的音乐技术能够重建这些分数直接从音频没有人为干预。我们的管道包括:新训练的萨克斯管的源分离模型、新的独奏萨克斯管的MIDI转录模型以及现有的单声道乐器的MIDI到乐谱方法的适应。 为了评估这个管道,我们还提供了一个增强的数据集查理帕克transmittance作为分数音频对与准确的双对齐和强拍注释。这代表了自动音频到乐谱转录的一个具有挑战性的新基准,我们希望这将推动研究超越单独转录音频到乐谱的领域。 总之,这些形成了另一个步骤,生产乐谱,音乐家可以直接使用,而不需要繁琐的更正或修改。为了方便未来的研究,所有模型检查点和数据都可以与转录管道的代码一起下载。我们的模块化管道的改进可能有一天使复杂的爵士乐独奏的自动转录成为一种常规的可能性,从而丰富了音乐教育和保存的可用资源。摘要:The Charlie Parker Omnibook is a cornerstone of jazz music education, described by pianist Ethan Iverson as "the most important jazz education text ever published". In this work we propose a new transcription pipeline and explore the extent to which state of the art music technology is able to reconstruct these scores directly from the audio without human intervention. Our pipeline includes: a newly trained source separation model for saxophone, a new MIDI transcription model for solo saxophone and an adaptation of an existing MIDI-to-score method for monophonic instruments. To assess this pipeline we also provide an enhanced dataset of Charlie Parker transcriptions as score-audio pairs with accurate MIDI alignments and downbeat annotations. This represents a challenging new benchmark for automatic audio-to-score transcription that we hope will advance research into areas beyond transcribing audio-to-MIDI alone. Together, these form another step towards producing scores that musicians can use directly, without the need for onerous corrections or revisions. To facilitate future research, all model checkpoints and data are made available to download along with code for the transcription pipeline. Improvements in our modular pipeline could one day make the automatic transcription of complex jazz solos a routine possibility, thereby enriching the resources available for music education and preservation.
【11】 C3LLM: Conditional Multimodal Content Generation Using Large Language Models
标题: C3 LLM:使用大型语言模型的条件多模式内容生成
作者:Zixuan Wang,Qinkai Duan,Yu-Wing Tai,Chi-Keung Tang
链接:点击下载PDF文件
摘要:我们介绍了C3 LLM(Conditioned-on-Three-Modalities Large Language Models),这是一个将视频到音频,音频到文本和文本到音频三个任务结合在一起的新框架。C3 LLM采用大型语言模型(LLM)结构作为桥梁,用于对齐不同的模态,合成给定的条件信息,并以离散的方式进行多模态生成。我们的贡献如下。首先,我们使用预先训练的音频码本为音频生成任务调整了分层结构。具体来说,我们训练LLM以从给定条件生成音频语义令牌,并且进一步使用非自回归Transformer来生成分层中的不同级别的声学令牌,以更好地增强所生成的音频的保真度。其次,基于LLM最初是为下一个单词预测方法的离散任务而设计的直觉,我们使用离散表示来生成音频,并将其语义压缩到声学标记中,类似于向LLM添加“声学词汇”。第三,我们的方法将之前的音频理解,视频到音频生成和文本到音频生成的任务结合到一个统一的模型中,以端到端的方式提供更多的通用性。我们的C3 LLM通过各种自动化评估指标实现了改进的结果,与以前的方法相比,提供了更好的语义对齐。摘要:We introduce C3LLM (Conditioned-on-Three-Modalities Large Language Models), a novel framework combining three tasks of video-to-audio, audio-to-text, and text-to-audio together. C3LLM adapts the Large Language Model (LLM) structure as a bridge for aligning different modalities, synthesizing the given conditional information, and making multimodal generation in a discrete manner. Our contributions are as follows. First, we adapt a hierarchical structure for audio generation tasks with pre-trained audio codebooks. Specifically, we train the LLM to generate audio semantic tokens from the given conditions, and further use a non-autoregressive transformer to generate different levels of acoustic tokens in layers to better enhance the fidelity of the generated audio. Second, based on the intuition that LLMs were originally designed for discrete tasks with the next-word prediction method, we use the discrete representation for audio generation and compress their semantic meanings into acoustic tokens, similar to adding "acoustic vocabulary" to LLM. Third, our method combines the previous tasks of audio understanding, video-to-audio generation, and text-to-audio generation together into one unified model, providing more versatility in an end-to-end fashion. Our C3LLM achieves improved results through various automated evaluation metrics, providing better semantic alignment compared to previous methods.
【12】 Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
标题: 使用严格延时神经网络的Carnatic Raga识别系统
作者:Sanjay Natesan,Homayoon Beigi
备注:7 pages, 2 tables, 3 figures
链接:点击下载PDF文件
摘要:基于大规模机器学习的Raga识别仍然是Carnatic音乐背后计算方面的一个重要问题。每一个拉加都有许多独特的内在旋律模式,可以用来很容易地将它们与其他的旋律模式区分开来。这些raga也可以用来聚类同一raga中的歌曲,以及识别其他密切相关的raga中的歌曲。在这种情况下,使用包括使用离散傅里叶变换和使用三角滤波的步骤的组合来分析输入声音,以创建可能音符的定制区间,从特定音符的存在或缺乏中提取特征。使用神经网络的组合,包括一维卷积神经网络(通常称为时延神经网络)和长短期记忆(LSTM),这是一种形式的递归神经网络,可以创建分类策略的骨干来构建模型。此外,为了帮助shruti的变化,将实施一种基于长时间注意力的机制来确定频率的相对变化,而不是绝对差异。这将在训练不同灌木中的音频片段时提供更有意义的数据点。为了评估分类器的准确性,使用了676个记录的数据集。这些歌曲分布在拉格斯的列表中。这个程序的目标是能够有效和高效地标记更广泛的音频片段,包括更多的shrutis,ragas和更多的背景噪音。摘要:Large scale machine learning-based Raga identification continues to be a nontrivial issue in the computational aspects behind Carnatic music. Each raga consists of many unique and intrinsic melodic patterns that can be used to easily identify them from others. These ragas can also then be used to cluster songs within the same raga, as well as identify songs in other closely related ragas. In this case, the input sound is analyzed using a combination of steps including using a Discrete Fourier transformation and using Triangular Filtering to create custom bins of possible notes, extracting features from the presence of particular notes or lack thereof. Using a combination of Neural Networks including 1D Convolutional Neural Networks conventionally known as Time-Delay Neural Networks) and Long Short-Term Memory (LSTM), which are a form of Recurrent Neural Networks, the backbone of the classification strategy to build the model can be created. In addition, to help with variations in shruti, a long-time attention-based mechanism will be implemented to determine the relative changes in frequency rather than the absolute differences. This will provide a much more meaningful data point when training audio clips in different shrutis. To evaluate the accuracy of the classifier, a dataset of 676 recordings is used. The songs are distributed across the list of ragas. The goal of this program is to be able to effectively and efficiently label a much wider range of audio clips in more shrutis, ragas, and with more background noise.
【13】 Quality-aware Masked Diffusion Transformer for Enhanced Music Generation
标题: 用于增强音乐生成的质量感知掩蔽扩散Transformer
作者:Chang Li,Ruoyu Wang,Lijuan Liu,Jun Du,Yixuan Sun,Zilu Guo,Zhenrong Zhang,Yuan Jiang
链接:点击下载PDF文件
摘要:近年来,基于扩散的文本到音乐(TTM)生成已经获得了突出,提供了一种新的方法来合成音乐内容从文本描述。在这一生成过程中实现高准确性和多样性需要大量高质量的数据,而这些数据往往只占可用数据集的一小部分。在开源数据集中,错误标记、弱标记、未标记数据和低质量音乐波形等问题的普遍存在严重阻碍了音乐生成模型的开发。为了克服这些挑战,我们引入了一种新的质量感知掩蔽扩散Transformer(QA-MDT)方法,使生成模型能够在训练过程中识别输入音乐波形的质量。基于音乐信号的独特属性,我们已经适应并实现了用于TTM任务的MDT模型,同时进一步揭示了其独特的质量控制能力。此外,我们解决了低质量的字幕与字幕细化数据处理方法的问题。我们的演示页面显示在https: qa-mdt.github.io 。代码https: github.com ivcylc qa-mdt摘要:In recent years, diffusion-based text-to-music (TTM) generation has gained prominence, offering a novel approach to synthesizing musical content from textual descriptions. Achieving high accuracy and diversity in this generation process requires extensive, high-quality data, which often constitutes only a fraction of available datasets. Within open-source datasets, the prevalence of issues like mislabeling, weak labeling, unlabeled data, and low-quality music waveform significantly hampers the development of music generation models. To overcome these challenges, we introduce a novel quality-aware masked diffusion transformer (QA-MDT) approach that enables generative models to discern the quality of input music waveform during training. Building on the unique properties of musical signals, we have adapted and implemented a MDT model for TTM task, while further unveiling its distinct capacity for quality control. Moreover, we address the issue of low-quality captions with a caption refinement data processing approach. Our demo page is shown in https: qa-mdt.github.io . Code on https: github.com ivcylc qa-mdt
机器翻译,仅供参考
