今日论文合集:cs.SD语音11篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音

【1】 Towards Speaker Identification with Minimal Dataset and Constrained  Resources using 1D-Convolution Neural Network

标题:使用1D卷积神经网络实现具有最少数据集和受约束资源的说话人识别
链接:https://arxiv.org/abs/2411.15082
作者:Irfan Nafiz Shahan,  Pulok Ahmed Auvi
摘要:语音识别和说话人识别对于安全和个人助理应用至关重要。本文提出了一种轻量级的一维卷积神经网络(1D-CNN),旨在对最小数据集进行说话人识别。我们的方法实现了97.87%的验证准确率,利用数据增强技术来处理背景噪声和有限的训练样本。未来的改进包括在更大的数据集上进行测试,并整合迁移学习方法以增强泛化能力。我们提供所有代码,自定义数据集和训练模型,以促进可重复性。这些资源可以在我们的GitHub存储库中找到:https://github.com/IrfanNafiz/RecMe。
摘要:Voice recognition and speaker identification are vital for applications insecurity and personal assistants. This paper presents a lightweight1D-Convolutional Neural Network (1D-CNN) designed to perform speakeridentification on minimal datasets. Our approach achieves a validation accuracyof 97.87%, leveraging data augmentation techniques to handle background noiseand limited training samples. Future improvements include testing on largerdatasets and integrating transfer learning methods to enhance generalizability.We provide all code, the custom dataset, and the trained models to facilitatereproducibility. These resources are available on our GitHub repository:https://github.com/IrfanNafiz/RecMe.

【2】 DAIRHuM: A Platform for Directly Aligning AI Representations with Human  Musical Judgments applied to Carnatic Music
标题:DAIRHuM:一个将人工智能表示与人类音乐判断直接匹配的平台,应用于狂欢音乐
链接:https://arxiv.org/abs/2411.14907
作者:Prashanth Thattai Ravikumar
备注:4 Pages, ICASSP workshop submission
摘要:将音乐AI模型表示与人类行为量化和对齐是MIR领域的一个重要挑战。本文提出了一个平台,用于探索人工智能音乐模型表示与人类音乐判断(DAIRHuM)之间的直接对齐。它旨在使音乐家和实验者能够标记音乐录音数据集中的相似性,并使用定量分数和视觉图来检查预训练模型与其标签的对齐。DAIRHuM被应用于分析NSynth表示之间的对齐,以及Carnatic四重奏合奏中两个演奏家之间的节奏二重奏,这是一个流派的例子,其中注释数据很少,评估对齐是不平凡的。结果表明,显着的研究结果与人类的判断,节奏和谐的模型对齐,同时突出的节奏感知和音乐相似性的判断,具体到卡纳迪音乐的关键差异。这项工作是使用户能够在卡纳蒂克音乐中探索人类-人工智能模型对齐并推进印度音乐中的MIR研究,同时处理数据稀缺性和文化特殊性的首批努力之一。该平台的开发为代表性不足的音乐类型提供了更好的音乐AI工具。
摘要:Quantifying and aligning music AI model representations with human behavioris an important challenge in the field of MIR. This paper presents a platformfor exploring the Direct alignment between AI music model Representations andHuman Musical judgments (DAIRHuM). It is designed to enable musicians andexperimentalists to label similarities in a dataset of music recordings, andexamine a pre-trained model's alignment with their labels using quantitativescores and visual plots. DAIRHuM is applied to analyze alignment between NSynthrepresentations, and a rhythmic duet between two percussionists in a Carnaticquartet ensemble, an example of a genre where annotated data is scarce andassessing alignment is non-trivial. The results demonstrate significantfindings on model alignment with human judgments of rhythmic harmony, whilehighlighting key differences in rhythm perception and music similarityjudgments specific to Carnatic music. This work is among the first efforts toenable users to explore human-AI model alignment in Carnatic music and advanceMIR research in Indian music while dealing with data scarcity and culturalspecificity. The development of this platform provides greater accessibility tomusic AI tools for under-represented genres.

【3】 Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large  Language Models
标题:谁能抵御聊天音频攻击?大型语言模型的评估基准
链接:https://arxiv.org/abs/2411.14842
作者:Wanqi Yang,  Yanda Li,  Meng Fang,  Yunchao Wei,  Tianyi Zhou,  Ling Chen
摘要:对抗性音频攻击对基于语音的人机交互中越来越多地使用大型语言模型(LLM)构成了重大威胁。虽然现有的研究主要集中在特定于模型的对抗方法上,但现实世界的应用需要一种更普遍和通用的音频对抗攻击方法。在本文中,我们介绍了聊天音频攻击(CAA)基准,包括四种不同类型的音频攻击,其目的是探索LLM在会话场景中对这些音频攻击的脆弱性。为了评估LLM的鲁棒性,我们提出了三种评估策略:标准评估,利用传统指标来量化模型在攻击下的性能;基于GPT-4 o的评估,模拟真实世界的会话复杂性;以及人类评估,提供对用户感知和信任的见解。我们评估了六个最先进的LLM与语音交互功能,包括双子座-1.5-Pro,GPT-4 o,和其他人,使用三种不同的评估方法在CAA基准。我们的综合分析揭示了四种类型的音频攻击对这些模型性能的影响,表明GPT-4 o具有最高水平的弹性。
摘要:Adversarial audio attacks pose a significant threat to the growing use oflarge language models (LLMs) in voice-based human-machine interactions. Whileexisting research has primarily focused on model-specific adversarial methods,real-world applications demand a more generalizable and universal approach toaudio adversarial attacks. In this paper, we introduce the Chat-Audio Attacks(CAA) benchmark including four distinct types of audio attacks, which aims toexplore the the vulnerabilities of LLMs to these audio attacks inconversational scenarios. To evaluate the robustness of LLMs, we propose threeevaluation strategies: Standard Evaluation, utilizing traditional metrics toquantify model performance under attacks; GPT-4o-Based Evaluation, whichsimulates real-world conversational complexities; and Human Evaluation,offering insights into user perception and trust. We evaluate sixstate-of-the-art LLMs with voice interaction capabilities, includingGemini-1.5-Pro, GPT-4o, and others, using three distinct evaluation methods onthe CAA benchmark. Our comprehensive analysis reveals the impact of four typesof audio attacks on the performance of these models, demonstrating that GPT-4oexhibits the highest level of resilience.

【4】 Mode-conditioned music learning and composition: a spiking neural  network inspired by neuroscience and psychology
标题:模式条件音乐学习和作曲:受神经科学和心理学启发的尖峰神经网络
链接:https://arxiv.org/abs/2411.14773
作者:Qian Liang,  Yi Zeng,  Menghaoran Tang
备注:18 pages, 8 figures
摘要:调式是建立音高组织框架和决定和声关系的重要因素之一。以往的研究往往采用简单、死板的对齐方法,忽视了模式的多样性。然而,与人工智能模型相比,人类拥有感知各种模式和键的认知机制。在本文中,我们提出了一种受大脑机制和心理学理论启发的尖峰神经网络来表示音乐模式和键,最终生成包含音调特征的音乐作品。具体而言,贡献如下:1)该模型设计了多个协作子系统,灵感来自相应的大脑区域的结构和功能; 2)我们引入了神经回路进化学习机制,使网络能够学习和生成音乐中的模式相关特征,反映了人类音乐感知中涉及的认知过程。3)结果表明,该模型与音乐心理学领域最重要的关键感知模型之一Krumhansl-Schmuckler模型具有相似的连接框架。4)实验表明,该模型能够生成具有给定调式和调特征的音乐作品。此外,生成的作品的定量评估表明,生成的音乐作品具有调性特征和旋律的适应性,需要生成多样化的音乐内容。通过将神经科学,心理学和音乐理论的见解与先进的神经网络架构相结合,我们的研究旨在创建一个系统,不仅学习和生成音乐,而且还弥合了人类认知和人工智能之间的差距。
摘要:Musical mode is one of the most critical element that establishes theframework of pitch organization and determines the harmonic relationships.Previous works often use the simplistic and rigid alignment method, andoverlook the diversity of modes. However, in contrast to AI models, humanspossess cognitive mechanisms for perceiving the various modes and keys. In thispaper, we propose a spiking neural network inspired by brain mechanisms andpsychological theories to represent musical modes and keys, ultimatelygenerating musical pieces that incorporate tonality features. Specifically, thecontributions are detailed as follows: 1) The model is designed with multiplecollaborated subsystems inspired by the structures and functions ofcorresponding brain regions; 2)We incorporate mechanisms for neural circuitevolutionary learning that enable the network to learn and generatemode-related features in music, reflecting the cognitive processes involved inhuman music perception. 3)The results demonstrate that the proposed model showsa connection framework closely similar to the Krumhansl-Schmuckler model, whichis one of the most significant key perception models in the music psychologydomain. 4) Experiments show that the model can generate music pieces withcharacteristics of the given modes and keys. Additionally, the quantitativeassessments of generated pieces reveals that the generating music pieces haveboth tonality characteristics and the melodic adaptability needed to generatediverse and musical content. By combining insights from neuroscience,psychology, and music theory with advanced neural network architectures, ourresearch aims to create a system that not only learns and generates music butalso bridges the gap between human cognition and artificial intelligence.

【5】 Generative AI for Music and Audio
标题:音乐和音频的生成人工智能
链接:https://arxiv.org/abs/2411.14627
作者:Hao-Wen Dong
备注:PhD Dissertation
摘要:生成式人工智能一直在改变我们与技术互动和消费内容的方式。在未来十年,人工智能技术将重塑我们在各种媒体中创建音频内容的方式,包括音乐,戏剧,电影,游戏,播客和短视频。在这篇论文中,我介绍了我围绕音乐和音频生成AI的三个主要研究方向:1)多轨音乐生成,2)辅助音乐创作工具,3)音频和音乐的多模态学习。通过我的研究,我的目标是回答以下两个基本问题:1)人工智能如何帮助专业人士或业余爱好者创作音乐和音频内容?2)人工智能能否以类似于人类学习音乐的方式学习创作音乐?我的长期目标是降低音乐创作的准入门槛,使音频内容创作民主化
摘要:Generative AI has been transforming the way we interact with technology andconsume content. In the next decade, AI technology will reshape how we createaudio content in various media, including music, theater, films, games,podcasts, and short videos. In this dissertation, I introduce the three maindirections of my research centered around generative AI for music and audio: 1)multitrack music generation, 2) assistive music creation tools, and 3)multimodal learning for audio and music. Through my research, I aim to answerthe following two fundamental questions: 1) How can AI help professionals oramateurs create music and audio content? 2) Can AI learn to create music in away similar to how humans learn music? My long-term goal is to lower thebarrier of entry for music composition and democratize audio content creation

【6】 Listening for Expert Identified Linguistic Features: Assessment of Audio  Deepfake Discernment among Undergraduate Students
标题:聆听专家识别的语言特征:本科生音频Deepfake辨别力评估
链接:https://arxiv.org/abs/2411.14586
作者:Noshaba N. Bhalli,  Nehal Naqvi,  Chloe Evered,  Christine Mallinson,  Vandana P. Janeja
摘要:本文评估了培训本科生通过聆听专家定义的语言特征来提高他们的音频deepfake识别能力的影响。这些功能已被证明可以提高AI算法的性能;在这里,我们确定AI算法的这种改进是否也可以转化为听众的感知意识和辨别能力的提高。由于人类是任何网络安全解决方案中最薄弱的环节,因此我们建议听众识别是提高音频内容可信度的关键因素。在这项研究中,我们确定是否培训,使听众熟悉英语语言的变化,可以提高他们的能力,以辨别音频deepfakes。我们专注于本科生,因为这个人口群体经常接触社交媒体以及在线欺骗和错误信息的可能性。据我们所知,我们的工作是第一个通过这种技术独特地解决英语音频deepfake识别的研究。我们的研究超越了信息训练,通过训练模块向听众引入有针对性的语言线索作为深度伪造识别机制。在实验前/实验后设计中,我们评估了264名学生的培训效果,这些学生是马里兰大学巴尔的摩县所有学生的代表性横截面,以及实验和控制部分。研究结果表明,实验组在评估音频片段时,他们的不满意度在统计学上显著降低,并且他们正确识别最初不确定的片段的能力有所提高。虽然结果很有希望,但未来的研究将探索更强大和全面的培训,以产生更大的影响。
摘要:This paper evaluates the impact of training undergraduate students to improvetheir audio deepfake discernment ability by listening for expert-definedlinguistic features. Such features have been shown to improve performance of AIalgorithms; here, we ascertain whether this improvement in AI algorithms alsotranslates to improvement of the perceptual awareness and discernment abilityof listeners. With humans as the weakest link in any cybersecurity solution, wepropose that listener discernment is a key factor for improving trustworthinessof audio content. In this study we determine whether training that familiarizeslisteners with English language variation can improve their abilities todiscern audio deepfakes. We focus on undergraduate students, as thisdemographic group is constantly exposed to social media and the potential fordeception and misinformation online. To the best of our knowledge, our work isthe first study to uniquely address English audio deepfake discernment throughsuch techniques. Our research goes beyond informational training by introducingtargeted linguistic cues to listeners as a deepfake discernment mechanism, viaa training module. In a pre-/post- experimental design, we evaluated the impactof the training across 264 students as a representative cross section of allstudents at the University of Maryland, Baltimore County, and acrossexperimental and control sections. Findings show that the experimental groupshowed a statistically significant decrease in their unsurety when evaluatingaudio clips and an improvement in their ability to correctly identify clipsthey were initially unsure about. While results are promising, future researchwill explore more robust and comprehensive trainings for greater impact.

【7】 From Statistical Methods to Pre-Trained Models; A Survey on Automatic  Speech Recognition for Resource Scarce Urdu Language
标题:从统计方法到预训练模型资源稀缺乌尔都语自动语音识别研究
链接:https://arxiv.org/abs/2411.14493
作者:Muhammad Sharif,  Zeeshan Abbas,  Jiangyan Yi,  Chenglin Liu
备注:Submitted to SN Computer Science
摘要:自动语音识别(ASR)技术近年来取得了重大进展,彻底改变了人机交互。虽然主要语言从这些发展中受益,但乌尔都语等资源较少的语言面临着独特的挑战。本文对ASR研究的动态景观进行了广泛的探索,特别关注资源受限的乌尔都语,这是南亚国家广泛使用的语言。它概述了目前的研究趋势,技术进步和未来的研究在乌尔都语ASR的潜在方向,旨在为未来的研究人员在这一领域感兴趣铺平道路。通过利用当代技术,分析现有数据集,并评估有效的算法和工具,本文旨在阐明与乌尔都语处理及其融入更广泛的语音研究领域相关的独特挑战和机遇。
摘要:Automatic Speech Recognition (ASR) technology has witnessed significantadvancements in recent years, revolutionizing human-computer interactions.While major languages have benefited from these developments, lesser-resourcedlanguages like Urdu face unique challenges. This paper provides an extensiveexploration of the dynamic landscape of ASR research, focusing particularly onthe resource-constrained Urdu language, which is widely spoken across SouthAsian nations. It outlines current research trends, technological advancements,and potential directions for future studies in Urdu ASR, aiming to pave the wayfor forthcoming researchers interested in this domain. By leveragingcontemporary technologies, analyzing existing datasets, and evaluatingeffective algorithms and tools, the paper seeks to shed light on the uniquechallenges and opportunities associated with Urdu language processing and itsintegration into the broader field of speech research.

【8】 GhostRNN: Reducing State Redundancy in RNN with Cheap Operations
标题:GhostRNN:通过廉价运营减少RNN中的状态冗余
链接:https://arxiv.org/abs/2411.14489
作者:Hang Zhou,  Xiaoxu Zheng,  Yunhe Wang,  Michael Bi Mi,  Deyi Xiong,  Kai Han
备注:None
摘要:能够对长距离依赖关系进行建模的递归神经网络(RNN)被广泛用于各种语音任务,例如,关键字识别(KWS)和语音增强(SE)。由于低资源设备的功耗和内存限制,现实世界的应用迫切需要高效的RNN模型。在本文中,我们提出了一种高效的RNN架构GhostRNN,它通过廉价的操作减少了隐藏状态冗余。特别是,我们观察到隐藏状态的部分维度与训练的RNN模型中的其他维度相似,这表明在特定的RNN中存在冗余。为了减少冗余,从而减少计算成本,我们建议首先生成一些本征态,然后应用廉价的操作来产生鬼状态的基础上的本征态。在KWS和SE任务上的实验表明,所提出的GhostRNN显着减少了内存使用(约40%)和计算成本,同时保持性能相似。
摘要:Recurrent neural network (RNNs) that are capable of modeling long-distancedependencies are widely used in various speech tasks, eg., keyword spotting(KWS) and speech enhancement (SE). Due to the limitation of power and memory inlow-resource devices, efficient RNN models are urgently required for real-worldapplications. In this paper, we propose an efficient RNN architecture,GhostRNN, which reduces hidden state redundancy with cheap operations. Inparticular, we observe that partial dimensions of hidden states are similar tothe others in trained RNN models, suggesting that redundancy exists in specificRNNs. To reduce the redundancy and hence computational cost, we propose tofirst generate a few intrinsic states, and then apply cheap operations toproduce ghost states based on the intrinsic states. Experiments on KWS and SEtasks demonstrate that the proposed GhostRNN significantly reduces the memoryusage (~40%) and computation cost while keeping performance similar.

【9】 Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre  Classification
标题:利用CNN进行注意力引导频谱图序列建模用于音乐流派分类
链接:https://arxiv.org/abs/2411.14474
作者:Aditya Sridhar
备注:6 pages, 7 figures, 17 References
摘要:音乐流派分类是音乐推荐系统、生成算法和文化分析的重要组成部分。在这项工作中,我们提出了一个创新的模型,使用基于注意力的时间签名建模的音乐流派分类。通过卷积神经网络(CNN)和多头注意力层处理频谱图序列,我们的方法捕捉了每件作品中最具时间意义的时刻,为流派识别制作了一个独特的“签名”。这种时间焦点不仅提高了分类精度,而且还揭示了可以直观地映射到听众感知的特定类型特征的见解。我们的研究结果通过突出跨流派的相似性和独特性,与人类的音乐直觉紧密结合,在个性化音乐推荐系统中提供了潜在的应用。这项工作弥合了技术分类任务和细致入微的人类体裁体验之间的差距。
摘要:Music genre classification is a critical component of music recommendationsystems, generation algorithms, and cultural analytics. In this work, wepresent an innovative model for classifying music genres using attention-basedtemporal signature modeling. By processing spectrogram sequences throughConvolutional Neural Networks (CNNs) and multi-head attention layers, ourapproach captures the most temporally significant moments within each piece,crafting a unique "signature" for genre identification. This temporal focus notonly enhances classification accuracy but also reveals insights intogenre-specific characteristics that can be intuitively mapped to listenerperceptions. Our findings offer potential applications in personalized musicrecommendation systems by highlighting cross-genre similarities anddistinctiveness, aligning closely with human musical intuition. This workbridges the gap between technical classification tasks and the nuanced, humanexperience of genre.

【10】 Direct Speech-to-Speech Neural Machine Translation: A Survey
标题:直接语音到语音神经机器翻译:调查
链接:https://arxiv.org/abs/2411.14453
作者:Mahendra Gupta,  Maitreyee Dutta,  Chandresh Kumar Maurya
摘要:语音到语音翻译(S2ST)模型将语音从一种语言转换为具有相同语言信息的另一种目标语言。S2ST对于弥合社区之间的通信差距非常重要,并且具有多种应用。近年来,研究人员引入了直接S2ST模型,该模型具有在不依赖中间文本生成的情况下翻译语音的潜力,具有更好的解码延迟,并且能够保留非语言和非语言特征。然而,直接S2ST尚未实现无缝通信的高质量性能,并且在性能方面仍然落后于级联模型,特别是在现实世界的翻译中。据我们所知,在直接S2ST系统上没有全面的调查,初学者和高级研究人员可以快速调查。目前的工作提供了一个全面的审查直接S2ST模型,数据和应用程序的问题,性能指标。我们批判性地分析了模型在基准数据集上的性能,并提供了研究挑战和未来的方向。
摘要:Speech-to-Speech Translation (S2ST) models transform speech from one languageto another target language with the same linguistic information. S2ST isimportant for bridging the communication gap among communities and has diverseapplications. In recent years, researchers have introduced direct S2ST models,which have the potential to translate speech without relying on intermediatetext generation, have better decoding latency, and the ability to preserveparalinguistic and non-linguistic features. However, direct S2ST has yet toachieve quality performance for seamless communication and still lags behindthe cascade models in terms of performance, especially in real-worldtranslation. To the best of our knowledge, no comprehensive survey is availableon the direct S2ST system, which beginners and advanced researchers can lookupon for a quick survey. The present work provides a comprehensive review ofdirect S2ST models, data and application issues, and performance metrics. Wecritically analyze the models' performance over the benchmark datasets andprovide research challenges and future directions.

【11】 Open-Amp: Synthetic Data Framework for Audio Effect Foundation Models
标题:开放收件箱:音效基础模型的合成数据框架
链接:https://arxiv.org/abs/2411.14972
作者:Alec Wright,  Alistair Carson,  Lauri Juvela
摘要:本文介绍了一个用于生成大规模和多样化音频效果数据的合成数据框架Open-EQUIPMENT。音频效果与许多音乐音频处理和音乐信息检索(MIR)任务相关,例如模拟音频效果的建模,自动混合,音调匹配和转录。现有的音频效果数据集在范围上是有限的,通常包括相对较少的音频效果处理器和有限数量的输入音频信号。我们提出的框架克服了这些问题,通过众包吉他放大器和效果的神经网络仿真,由开源音频效果仿真软件的用户创建。这使得Open-Side的用户能够完全控制效果模型要处理的输入信号,并提供数百种设备的高质量仿真。Open-Escape可以在训练期间在线渲染音频,从而在数据增强方面具有很大的灵活性。我们的实验表明,使用开放式训练训练吉他效果编码器实现了新的国家的最先进的结果,对多个吉他效果分类任务。此外,我们训练一个一对多的吉他效果模型使用Open-EQUIPMENT,并使用它来模拟看不见的模拟效果,通过操纵其学习的潜在空间,表示可转移到模拟吉他效果数据。
摘要:This paper introduces Open-Amp, a synthetic data framework for generatinglarge-scale and diverse audio effects data. Audio effects are relevant to manymusical audio processing and Music Information Retrieval (MIR) tasks, such asmodelling of analog audio effects, automatic mixing, tone matching andtranscription. Existing audio effects datasets are limited in scope, usuallyincluding relatively few audio effects processors and a limited amount of inputaudio signals. Our proposed framework overcomes these issues, by crowdsourcingneural network emulations of guitar amplifiers and effects, created by users ofopen-source audio effects emulation software. This allows users of Open-Ampcomplete control over the input signals to be processed by the effects models,as well as providing high-quality emulations of hundreds of devices. Open-Ampcan render audio online during training, allowing great flexibility in dataaugmentation. Our experiments show that using Open-Amp to train a guitareffects encoder achieves new state-of-the-art results on multiple guitareffects classification tasks. Furthermore, we train a one-to-many guitareffects model using Open-Amp, and use it to emulate unseen analog effects viamanipulation of its learned latent space, indicating transferability to analogguitar effects data.

eess.AS音频处理

【1】 Open-Amp: Synthetic Data Framework for Audio Effect Foundation Models
标题:开放收件箱:音效基础模型的合成数据框架
链接:https://arxiv.org/abs/2411.14972
作者:Alec Wright,  Alistair Carson,  Lauri Juvela
摘要:本文介绍了一个用于生成大规模和多样化音频效果数据的合成数据框架Open-EQUIPMENT。音频效果与许多音乐音频处理和音乐信息检索(MIR)任务相关,例如模拟音频效果的建模,自动混合,音调匹配和转录。现有的音频效果数据集在范围上是有限的,通常包括相对较少的音频效果处理器和有限数量的输入音频信号。我们提出的框架克服了这些问题,通过众包吉他放大器和效果的神经网络仿真,由开源音频效果仿真软件的用户创建。这使得Open-Side的用户能够完全控制效果模型要处理的输入信号,并提供数百种设备的高质量仿真。Open-Escape可以在训练期间在线渲染音频,从而在数据增强方面具有很大的灵活性。我们的实验表明,使用开放式训练训练吉他效果编码器实现了新的国家的最先进的结果,对多个吉他效果分类任务。此外,我们训练一个一对多的吉他效果模型使用Open-EQUIPMENT,并使用它来模拟看不见的模拟效果,通过操纵其学习的潜在空间,表示可转移到模拟吉他效果数据。
摘要:This paper introduces Open-Amp, a synthetic data framework for generatinglarge-scale and diverse audio effects data. Audio effects are relevant to manymusical audio processing and Music Information Retrieval (MIR) tasks, such asmodelling of analog audio effects, automatic mixing, tone matching andtranscription. Existing audio effects datasets are limited in scope, usuallyincluding relatively few audio effects processors and a limited amount of inputaudio signals. Our proposed framework overcomes these issues, by crowdsourcingneural network emulations of guitar amplifiers and effects, created by users ofopen-source audio effects emulation software. This allows users of Open-Ampcomplete control over the input signals to be processed by the effects models,as well as providing high-quality emulations of hundreds of devices. Open-Ampcan render audio online during training, allowing great flexibility in dataaugmentation. Our experiments show that using Open-Amp to train a guitareffects encoder achieves new state-of-the-art results on multiple guitareffects classification tasks. Furthermore, we train a one-to-many guitareffects model using Open-Amp, and use it to emulate unseen analog effects viamanipulation of its learned latent space, indicating transferability to analogguitar effects data.

【2】 Towards Speaker Identification with Minimal Dataset and Constrained  Resources using 1D-Convolution Neural Network
标题:使用1D卷积神经网络实现具有最少数据集和受约束资源的说话人识别
链接:https://arxiv.org/abs/2411.15082
作者:Irfan Nafiz Shahan,  Pulok Ahmed Auvi
摘要:语音识别和说话人识别对于安全和个人助理应用至关重要。本文提出了一种轻量级的一维卷积神经网络(1D-CNN),旨在对最小数据集进行说话人识别。我们的方法实现了97.87%的验证准确率,利用数据增强技术来处理背景噪声和有限的训练样本。未来的改进包括在更大的数据集上进行测试,并整合迁移学习方法以增强泛化能力。我们提供所有代码,自定义数据集和训练模型,以促进可重复性。这些资源可以在我们的GitHub存储库中找到:https://github.com/IrfanNafiz/RecMe。
摘要:Voice recognition and speaker identification are vital for applications insecurity and personal assistants. This paper presents a lightweight1D-Convolutional Neural Network (1D-CNN) designed to perform speakeridentification on minimal datasets. Our approach achieves a validation accuracyof 97.87%, leveraging data augmentation techniques to handle background noiseand limited training samples. Future improvements include testing on largerdatasets and integrating transfer learning methods to enhance generalizability.We provide all code, the custom dataset, and the trained models to facilitatereproducibility. These resources are available on our GitHub repository:https://github.com/IrfanNafiz/RecMe.

【3】 DAIRHuM: A Platform for Directly Aligning AI Representations with Human  Musical Judgments applied to Carnatic Music
标题:DAIRHuM:一个将人工智能表示与人类音乐判断直接匹配的平台,应用于狂欢音乐
链接:https://arxiv.org/abs/2411.14907
作者:Prashanth Thattai Ravikumar
备注:4 Pages, ICASSP workshop submission
摘要:将音乐AI模型表示与人类行为量化和对齐是MIR领域的一个重要挑战。本文提出了一个平台,用于探索人工智能音乐模型表示与人类音乐判断(DAIRHuM)之间的直接对齐。它旨在使音乐家和实验者能够标记音乐录音数据集中的相似性,并使用定量分数和视觉图来检查预训练模型与其标签的对齐。DAIRHuM被应用于分析NSynth表示之间的对齐,以及Carnatic四重奏合奏中两个演奏家之间的节奏二重奏,这是一个流派的例子,其中注释数据很少,评估对齐是不平凡的。结果表明,显着的研究结果与人类的判断,节奏和谐的模型对齐,同时突出的节奏感知和音乐相似性的判断,具体到卡纳迪音乐的关键差异。这项工作是使用户能够在卡纳蒂克音乐中探索人类-人工智能模型对齐并推进印度音乐中的MIR研究,同时处理数据稀缺性和文化特殊性的首批努力之一。该平台的开发为代表性不足的音乐类型提供了更好的音乐AI工具。
摘要:Quantifying and aligning music AI model representations with human behavioris an important challenge in the field of MIR. This paper presents a platformfor exploring the Direct alignment between AI music model Representations andHuman Musical judgments (DAIRHuM). It is designed to enable musicians andexperimentalists to label similarities in a dataset of music recordings, andexamine a pre-trained model's alignment with their labels using quantitativescores and visual plots. DAIRHuM is applied to analyze alignment between NSynthrepresentations, and a rhythmic duet between two percussionists in a Carnaticquartet ensemble, an example of a genre where annotated data is scarce andassessing alignment is non-trivial. The results demonstrate significantfindings on model alignment with human judgments of rhythmic harmony, whilehighlighting key differences in rhythm perception and music similarityjudgments specific to Carnatic music. This work is among the first efforts toenable users to explore human-AI model alignment in Carnatic music and advanceMIR research in Indian music while dealing with data scarcity and culturalspecificity. The development of this platform provides greater accessibility tomusic AI tools for under-represented genres.

【4】 Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large  Language Models
标题:谁能抵御聊天音频攻击?大型语言模型的评估基准
链接:https://arxiv.org/abs/2411.14842
作者:Wanqi Yang,  Yanda Li,  Meng Fang,  Yunchao Wei,  Tianyi Zhou,  Ling Chen
摘要:对抗性音频攻击对基于语音的人机交互中越来越多地使用大型语言模型(LLM)构成了重大威胁。虽然现有的研究主要集中在特定于模型的对抗方法上,但现实世界的应用需要一种更普遍和通用的音频对抗攻击方法。在本文中,我们介绍了聊天音频攻击(CAA)基准,包括四种不同类型的音频攻击,其目的是探索LLM在会话场景中对这些音频攻击的脆弱性。为了评估LLM的鲁棒性,我们提出了三种评估策略:标准评估,利用传统指标来量化模型在攻击下的性能;基于GPT-4 o的评估,模拟真实世界的会话复杂性;以及人类评估,提供对用户感知和信任的见解。我们评估了六个最先进的LLM与语音交互功能,包括双子座-1.5-Pro,GPT-4 o,和其他人,使用三种不同的评估方法在CAA基准。我们的综合分析揭示了四种类型的音频攻击对这些模型性能的影响,表明GPT-4 o具有最高水平的弹性。
摘要:Adversarial audio attacks pose a significant threat to the growing use oflarge language models (LLMs) in voice-based human-machine interactions. Whileexisting research has primarily focused on model-specific adversarial methods,real-world applications demand a more generalizable and universal approach toaudio adversarial attacks. In this paper, we introduce the Chat-Audio Attacks(CAA) benchmark including four distinct types of audio attacks, which aims toexplore the the vulnerabilities of LLMs to these audio attacks inconversational scenarios. To evaluate the robustness of LLMs, we propose threeevaluation strategies: Standard Evaluation, utilizing traditional metrics toquantify model performance under attacks; GPT-4o-Based Evaluation, whichsimulates real-world conversational complexities; and Human Evaluation,offering insights into user perception and trust. We evaluate sixstate-of-the-art LLMs with voice interaction capabilities, includingGemini-1.5-Pro, GPT-4o, and others, using three distinct evaluation methods onthe CAA benchmark. Our comprehensive analysis reveals the impact of four typesof audio attacks on the performance of these models, demonstrating that GPT-4oexhibits the highest level of resilience.

【5】 Mode-conditioned music learning and composition: a spiking neural  network inspired by neuroscience and psychology
标题:模式条件音乐学习和作曲:受神经科学和心理学启发的尖峰神经网络
链接:https://arxiv.org/abs/2411.14773
作者:Qian Liang,  Yi Zeng,  Menghaoran Tang
备注:18 pages, 8 figures
摘要:调式是建立音高组织框架和决定和声关系的重要因素之一。以往的研究往往采用简单、死板的对齐方法,忽视了模式的多样性。然而,与人工智能模型相比,人类拥有感知各种模式和键的认知机制。在本文中,我们提出了一种受大脑机制和心理学理论启发的尖峰神经网络来表示音乐模式和键,最终生成包含音调特征的音乐作品。具体而言,贡献如下:1)该模型设计了多个协作子系统,灵感来自相应的大脑区域的结构和功能; 2)我们引入了神经回路进化学习机制,使网络能够学习和生成音乐中的模式相关特征,反映了人类音乐感知中涉及的认知过程。3)结果表明,该模型与音乐心理学领域最重要的关键感知模型之一Krumhansl-Schmuckler模型具有相似的连接框架。4)实验表明,该模型能够生成具有给定调式和调特征的音乐作品。此外,对生成作品的定量评估表明,生成的音乐作品既具有调性特征,又具有生成多样化音乐内容所需的旋律适应性。通过将神经科学,心理学和音乐理论的见解与先进的神经网络架构相结合,我们的研究旨在创建一个系统,不仅学习和生成音乐,而且还弥合了人类认知和人工智能之间的差距。
摘要:Musical mode is one of the most critical element that establishes theframework of pitch organization and determines the harmonic relationships.Previous works often use the simplistic and rigid alignment method, andoverlook the diversity of modes. However, in contrast to AI models, humanspossess cognitive mechanisms for perceiving the various modes and keys. In thispaper, we propose a spiking neural network inspired by brain mechanisms andpsychological theories to represent musical modes and keys, ultimatelygenerating musical pieces that incorporate tonality features. Specifically, thecontributions are detailed as follows: 1) The model is designed with multiplecollaborated subsystems inspired by the structures and functions ofcorresponding brain regions; 2)We incorporate mechanisms for neural circuitevolutionary learning that enable the network to learn and generatemode-related features in music, reflecting the cognitive processes involved inhuman music perception. 3)The results demonstrate that the proposed model showsa connection framework closely similar to the Krumhansl-Schmuckler model, whichis one of the most significant key perception models in the music psychologydomain. 4) Experiments show that the model can generate music pieces withcharacteristics of the given modes and keys. Additionally, the quantitativeassessments of generated pieces reveals that the generating music pieces haveboth tonality characteristics and the melodic adaptability needed to generatediverse and musical content. By combining insights from neuroscience,psychology, and music theory with advanced neural network architectures, ourresearch aims to create a system that not only learns and generates music butalso bridges the gap between human cognition and artificial intelligence.

【6】 VQalAttent: a Transparent Speech Generation Pipeline based on  Transformer-learned VQ-VAE Latent Space
标题:VQalAttent:基于Transformer学习的VQ-VAE潜伏空间的透明语音生成管道
链接:https://arxiv.org/abs/2411.14642
作者:Armani Rodriguez,  Silvija Kokalj-Filipovic
摘要:高效地生成高质量的语音仍然是语音合成中生成模型的关键挑战。本文介绍了VQalAttent,一个轻量级的模型,旨在生成假语音与可调性能和可解释性。利用AudioMNIST数据集,包括人类的话语的十进制数字(0-9),我们的方法采用了两步架构:第一,一个可扩展的矢量量化自动编码器(VQ-VAE),压缩音频频谱到离散的潜在表示,第二,解码器只有Transformer,学习这些潜在的概率模型。经过训练的Transformer生成类似的潜在序列,通过VQ-VAE解码器可转换为音频频谱图,从中我们生成假话语。根据潜在空间的维度和外部信息来解释假货的统计和感知质量,可以在更大的商业生成模型中进行指导性改进。作为理解和改进音频合成的宝贵工具,我们的结果证明了VQalAttent能够以有限的计算资源生成可理解的语音样本,而训练管道的模块化和透明性有助于轻松地将分析与模块化修改相关联,从而为更复杂的模型提供见解。
摘要:Generating high-quality speech efficiently remains a key challenge forgenerative models in speech synthesis. This paper introduces VQalAttent, alightweight model designed to generate fake speech with tunable performance andinterpretability. Leveraging the AudioMNIST dataset, consisting of humanutterances of decimal digits (0-9), our method employs a two-step architecture:first, a scalable vector quantized autoencoder (VQ-VAE) that compresses audiospectrograms into discrete latent representations, and second, a decoder-onlytransformer that learns the probability model of these latents. Trainedtransformer generates similar latent sequences, convertible to audiospectrograms by the VQ-VAE decoder, from which we generate fake utterances.Interpreting statistical and perceptual quality of the fakes, depending on thedimension and the extrinsic information of the latent space, enables guidedimprovements in larger, commercial generative models. As a valuable tool forunderstanding and refining audio synthesis, our results demonstrateVQalAttent's capacity to generate intelligible speech samples with limitedcomputational resources, while the modularity and transparency of the trainingpipeline helps easily correlate the analytics with modular modifications, henceproviding insights for the more complex models.

【7】 Generative AI for Music and Audio
标题:音乐和音频的生成人工智能
链接:https://arxiv.org/abs/2411.14627
作者:Hao-Wen Dong
备注:PhD Dissertation
摘要:生成式人工智能一直在改变我们与技术互动和消费内容的方式。在未来十年,人工智能技术将重塑我们在各种媒体中创建音频内容的方式,包括音乐,戏剧,电影,游戏,播客和短视频。在这篇论文中,我介绍了我围绕音乐和音频生成AI的三个主要研究方向:1)多轨音乐生成,2)辅助音乐创作工具,3)音频和音乐的多模态学习。通过我的研究,我的目标是回答以下两个基本问题:1)人工智能如何帮助专业人士或业余爱好者创作音乐和音频内容?2)人工智能能像人类学习音乐一样学习创作音乐吗?我的长期目标是降低音乐创作的准入门槛,使音频内容创作民主化
摘要:Generative AI has been transforming the way we interact with technology andconsume content. In the next decade, AI technology will reshape how we createaudio content in various media, including music, theater, films, games,podcasts, and short videos. In this dissertation, I introduce the three maindirections of my research centered around generative AI for music and audio: 1)multitrack music generation, 2) assistive music creation tools, and 3)multimodal learning for audio and music. Through my research, I aim to answerthe following two fundamental questions: 1) How can AI help professionals oramateurs create music and audio content? 2) Can AI learn to create music in away similar to how humans learn music? My long-term goal is to lower thebarrier of entry for music composition and democratize audio content creation

【8】 Listening for Expert Identified Linguistic Features: Assessment of Audio  Deepfake Discernment among Undergraduate Students
标题:聆听专家识别的语言特征:本科生音频Deepfake辨别力评估
链接:https://arxiv.org/abs/2411.14586
作者:Noshaba N. Bhalli,  Nehal Naqvi,  Chloe Evered,  Christine Mallinson,  Vandana P. Janeja
摘要:本文评估了培训本科生通过聆听专家定义的语言特征来提高他们的音频deepfake识别能力的影响。这些功能已被证明可以提高AI算法的性能;在这里,我们确定AI算法的这种改进是否也可以转化为听众的感知意识和辨别能力的提高。由于人类是任何网络安全解决方案中最薄弱的环节,因此我们建议听众识别是提高音频内容可信度的关键因素。在这项研究中,我们确定是否培训,使听众熟悉英语语言的变化,可以提高他们的能力,以辨别音频deepfakes。我们专注于本科生,因为这个人口群体经常接触社交媒体以及在线欺骗和错误信息的可能性。据我们所知,我们的工作是第一个通过这种技术独特地解决英语音频deepfake识别的研究。我们的研究超越了信息训练,通过训练模块向听众引入有针对性的语言线索作为深度伪造识别机制。在实验前/实验后设计中,我们评估了264名学生的培训效果,这些学生是马里兰大学巴尔的摩县所有学生的代表性横截面,以及实验和控制部分。研究结果表明,实验组在评估音频片段时,他们的不满意度在统计学上显著降低,并且他们正确识别最初不确定的片段的能力有所提高。虽然结果很有希望,但未来的研究将探索更强大和全面的培训,以产生更大的影响。
摘要:This paper evaluates the impact of training undergraduate students to improvetheir audio deepfake discernment ability by listening for expert-definedlinguistic features. Such features have been shown to improve performance of AIalgorithms; here, we ascertain whether this improvement in AI algorithms alsotranslates to improvement of the perceptual awareness and discernment abilityof listeners. With humans as the weakest link in any cybersecurity solution, wepropose that listener discernment is a key factor for improving trustworthinessof audio content. In this study we determine whether training that familiarizeslisteners with English language variation can improve their abilities todiscern audio deepfakes. We focus on undergraduate students, as thisdemographic group is constantly exposed to social media and the potential fordeception and misinformation online. To the best of our knowledge, our work isthe first study to uniquely address English audio deepfake discernment throughsuch techniques. Our research goes beyond informational training by introducingtargeted linguistic cues to listeners as a deepfake discernment mechanism, viaa training module. In a pre-/post- experimental design, we evaluated the impactof the training across 264 students as a representative cross section of allstudents at the University of Maryland, Baltimore County, and acrossexperimental and control sections. Findings show that the experimental groupshowed a statistically significant decrease in their unsurety when evaluatingaudio clips and an improvement in their ability to correctly identify clipsthey were initially unsure about. While results are promising, future researchwill explore more robust and comprehensive trainings for greater impact.

【9】 From Statistical Methods to Pre-Trained Models; A Survey on Automatic  Speech Recognition for Resource Scarce Urdu Language
标题:从统计方法到预训练模型资源稀缺乌尔都语自动语音识别研究
链接:https://arxiv.org/abs/2411.14493
作者:Muhammad Sharif,  Zeeshan Abbas,  Jiangyan Yi,  Chenglin Liu
备注:Submitted to SN Computer Science
摘要:自动语音识别(ASR)技术近年来取得了重大进展,彻底改变了人机交互。虽然主要语言从这些发展中受益,但乌尔都语等资源较少的语言面临着独特的挑战。本文对ASR研究的动态景观进行了广泛的探索,特别关注资源受限的乌尔都语,这是南亚国家广泛使用的语言。它概述了目前的研究趋势,技术进步和未来的研究在乌尔都语ASR的潜在方向,旨在为未来的研究人员在这一领域感兴趣铺平道路。通过利用当代技术,分析现有数据集,并评估有效的算法和工具,本文旨在阐明与乌尔都语处理及其融入更广泛的语音研究领域相关的独特挑战和机遇。
摘要:Automatic Speech Recognition (ASR) technology has witnessed significantadvancements in recent years, revolutionizing human-computer interactions.While major languages have benefited from these developments, lesser-resourcedlanguages like Urdu face unique challenges. This paper provides an extensiveexploration of the dynamic landscape of ASR research, focusing particularly onthe resource-constrained Urdu language, which is widely spoken across SouthAsian nations. It outlines current research trends, technological advancements,and potential directions for future studies in Urdu ASR, aiming to pave the wayfor forthcoming researchers interested in this domain. By leveragingcontemporary technologies, analyzing existing datasets, and evaluatingeffective algorithms and tools, the paper seeks to shed light on the uniquechallenges and opportunities associated with Urdu language processing and itsintegration into the broader field of speech research.

【10】 GhostRNN: Reducing State Redundancy in RNN with Cheap Operations
标题:GhostRNN:通过廉价运营减少RNN中的状态冗余
链接:https://arxiv.org/abs/2411.14489
作者:Hang Zhou,  Xiaoxu Zheng,  Yunhe Wang,  Michael Bi Mi,  Deyi Xiong,  Kai Han
备注:None
摘要:能够对长距离依赖关系进行建模的递归神经网络(RNN)被广泛用于各种语音任务,例如,关键字识别(KWS)和语音增强(SE)。由于低资源设备的功耗和内存限制,现实世界的应用迫切需要高效的RNN模型。在本文中,我们提出了一种高效的RNN架构GhostRNN,它通过廉价的操作减少了隐藏状态冗余。特别是,我们观察到隐藏状态的部分维度与训练的RNN模型中的其他维度相似,这表明在特定的RNN中存在冗余。为了减少冗余,从而减少计算成本,我们建议首先生成一些本征态,然后应用廉价的操作来产生鬼状态的基础上的本征态。在KWS和SE任务上的实验表明,所提出的GhostRNN显着减少了内存使用(约40%)和计算成本,同时保持性能相似。
摘要:Recurrent neural network (RNNs) that are capable of modeling long-distancedependencies are widely used in various speech tasks, eg., keyword spotting(KWS) and speech enhancement (SE). Due to the limitation of power and memory inlow-resource devices, efficient RNN models are urgently required for real-worldapplications. In this paper, we propose an efficient RNN architecture,GhostRNN, which reduces hidden state redundancy with cheap operations. Inparticular, we observe that partial dimensions of hidden states are similar tothe others in trained RNN models, suggesting that redundancy exists in specificRNNs. To reduce the redundancy and hence computational cost, we propose tofirst generate a few intrinsic states, and then apply cheap operations toproduce ghost states based on the intrinsic states. Experiments on KWS and SEtasks demonstrate that the proposed GhostRNN significantly reduces the memoryusage (~40%) and computation cost while keeping performance similar.

【11】 Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre  Classification
标题:利用CNN进行注意力引导频谱图序列建模用于音乐流派分类
链接:https://arxiv.org/abs/2411.14474
作者:Aditya Sridhar
备注:6 pages, 7 figures, 17 References
摘要:音乐流派分类是音乐推荐系统、生成算法和文化分析的重要组成部分。在这项工作中,我们提出了一种使用基于注意力的时间签名建模对音乐流派进行分类的创新模型。通过卷积神经网络(CNN)和多头注意力层处理频谱图序列,我们的方法捕捉了每件作品中最具时间意义的时刻,为流派识别制作了一个独特的“签名”。这种时间焦点不仅提高了分类精度,而且还揭示了可以直观地映射到听众感知的特定类型特征的见解。我们的研究结果通过突出跨流派的相似性和独特性,与人类的音乐直觉紧密结合,在个性化音乐推荐系统中提供了潜在的应用。这项工作弥合了技术分类任务和细致入微的人类体裁体验之间的差距。
摘要:Music genre classification is a critical component of music recommendationsystems, generation algorithms, and cultural analytics. In this work, wepresent an innovative model for classifying music genres using attention-basedtemporal signature modeling. By processing spectrogram sequences throughConvolutional Neural Networks (CNNs) and multi-head attention layers, ourapproach captures the most temporally significant moments within each piece,crafting a unique "signature" for genre identification. This temporal focus notonly enhances classification accuracy but also reveals insights intogenre-specific characteristics that can be intuitively mapped to listenerperceptions. Our findings offer potential applications in personalized musicrecommendation systems by highlighting cross-genre similarities anddistinctiveness, aligning closely with human musical intuition. This workbridges the gap between technical classification tasks and the nuanced, humanexperience of genre.

【12】 Direct Speech-to-Speech Neural Machine Translation: A Survey
标题:直接语音到语音神经机器翻译:调查
链接:https://arxiv.org/abs/2411.14453
作者:Mahendra Gupta,  Maitreyee Dutta,  Chandresh Kumar Maurya
摘要:语音到语音翻译(S2ST)模型将语音从一种语言转换为具有相同语言信息的另一种目标语言。S2ST对于弥合社区之间的通信差距非常重要,并且具有多种应用。近年来,研究人员引入了直接S2ST模型,该模型具有在不依赖中间文本生成的情况下翻译语音的潜力,具有更好的解码延迟,并且能够保留非语言和非语言特征。然而,直接S2ST尚未实现无缝通信的高质量性能,并且在性能方面仍然落后于级联模型,特别是在现实世界的翻译中。据我们所知,在直接S2ST系统上没有全面的调查,初学者和高级研究人员可以快速调查。目前的工作提供了一个全面的审查直接S2ST模型,数据和应用程序的问题,性能指标。我们批判性地分析了模型在基准数据集上的性能,并提供了研究挑战和未来的方向。
摘要:Speech-to-Speech Translation (S2ST) models transform speech from one languageto another target language with the same linguistic information. S2ST isimportant for bridging the communication gap among communities and has diverseapplications. In recent years, researchers have introduced direct S2ST models,which have the potential to translate speech without relying on intermediatetext generation, have better decoding latency, and the ability to preserveparalinguistic and non-linguistic features. However, direct S2ST has yet toachieve quality performance for seamless communication and still lags behindthe cascade models in terms of performance, especially in real-worldtranslation. To the best of our knowledge, no comprehensive survey is availableon the direct S2ST system, which beginners and advanced researchers can lookupon for a quick survey. The present work provides a comprehensive review ofdirect S2ST models, data and application issues, and performance metrics. Wecritically analyze the models' performance over the benchmark datasets andprovide research challenges and future directions.

机器翻译由腾讯交互翻译提供,仅供参考