cs.SD语音【1】Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer标题:通过音频风格转移评估自动语音识别系统的稳健性链接:https://arxiv.org/abs/2405.09470作者:Weifei Jin,Yuxin Cao,Junjie Su,Qi Shen,Kai Ye,Derui Wang,Jie Hao,Ziyao Liu备注:Accepted to SecTL (AsiaCCS Workshop) 2024摘要:鉴于自动语音识别(ASR)系统的广泛应用,其安全问题比以往任何时候都受到了更多的关注,这主要是由于深度神经网络的敏感性。以前的研究表明,秘密制作对抗性扰动可以操纵语音识别系统,从而产生恶意命令。这些攻击方法大多需要在$\ell_p$范数约束下添加噪声扰动,不可避免地留下手动修改的伪影。最近的研究通过操纵风格向量来基于文本到语音(TTS)合成音频合成对抗性示例,从而缓解了这一限制。然而,基于优化目标的风格修改显著降低了音频风格的可控性和可编辑性。在本文中,我们提出了一种攻击ASR系统的基础上,用户定制的风格转移。我们首先测试了风格转移攻击(STA)的效果,它将风格转移和对抗性攻击按顺序结合起来。然后,作为改进,我们提出了一个迭代的风格代码攻击(SCA),以保持音频质量。实验结果表明,该方法可以满足用户定制风格的需要,并取得了82%的攻击成功率,同时保持声音的自然性,由于我们的用户学习。摘要:In light of the widespread application of Automatic Speech Recognition (ASR) systems, their security concerns have received much more attention than ever before, primarily due to the susceptibility of Deep Neural Networks. Previous studies have illustrated that surreptitiously crafting adversarial perturbations enables the manipulation of speech recognition systems, resulting in the production of malicious commands. These attack methods mostly require adding noise perturbations under $\ell_p$ norm constraints, inevitably leaving behind artifacts of manual modifications. Recent research has alleviated this limitation by manipulating style vectors to synthesize adversarial examples based on Text-to-Speech (TTS) synthesis audio. However, style modifications based on optimization objectives significantly reduce the controllability and editability of audio styles. In this paper, we propose an attack on ASR systems based on user-customized style transfer. We first test the effect of Style Transfer Attack (STA) which combines style transfer and adversarial attack in sequential order. And then, as an improvement, we propose an iterative Style Code Attack (SCA) to maintain audio quality. Experimental results show that our method can meet the need for user-customized styles and achieve a success rate of 82% in attacks, while keeping sound naturalness due to our user study.
【2】 Dance Any Beat: Blending Beats with Visuals in Dance Video Generation标题:舞蹈任意节拍:在舞蹈视频生成中将节拍与视觉效果融为一体链接:https://arxiv.org/abs/2405.09266作者:Xuanchen Wang,Heng Wang,Dongnan Liu,Weidong Cai备注:11 pages, 6 figures, demo page: this https URL摘要:从音乐生成舞蹈的任务是至关重要的,但目前的方法,主要产生联合序列,导致输出缺乏直观性和复杂的数据收集,由于精确的联合注释的必要性。我们介绍了一个舞蹈任何节拍扩散模型,即DabFusion,采用音乐作为有条件的输入,直接创建舞蹈视频从静止图像,利用有条件的图像到视频生成原理。这种方法开创了使用音乐作为图像到视频合成的调节因素。我们的方法分为两个阶段:训练自动编码器来预测参考帧和驱动帧之间的潜在光流,消除对联合注释的需要,并训练基于U-Net的扩散模型来产生这些由CLAP编码的音乐节奏引导的潜在光流。虽然能够制作高质量的舞蹈视频,但基线模型在节奏调整方面很困难。我们通过增加节拍信息,提高同步性来增强模型。我们引入了一个2D运动音乐对齐分数(2D-MM对齐)的定量评估。在AIST++数据集上进行评估,我们的增强模型显示出2D-MM Align评分和既定指标的显着改善。视频结果可以在我们的项目页面上找到:https://DabFusion.github.io。摘要:The task of generating dance from music is crucial, yet current methods, which mainly produce joint sequences, lead to outputs that lack intuitiveness and complicate data collection due to the necessity for precise joint annotations. We introduce a Dance Any Beat Diffusion model, namely DabFusion, that employs music as a conditional input to directly create dance videos from still images, utilizing conditional image-to-video generation principles. This approach pioneers the use of music as a conditioning factor in image-to-video synthesis. Our method unfolds in two stages: training an auto-encoder to predict latent optical flow between reference and driving frames, eliminating the need for joint annotation, and training a U-Net-based diffusion model to produce these latent optical flows guided by music rhythm encoded by CLAP. Although capable of producing high-quality dance videos, the baseline model struggles with rhythm alignment. We enhance the model by adding beat information, improving synchronization. We introduce a 2D motion-music alignment score (2D-MM Align) for quantitative assessment. Evaluated on the AIST++ dataset, our enhanced model shows marked improvements in 2D-MM Align score and established metrics. Video results can be found on our project page: https://DabFusion.github.io.
【3】 SMUG-Explain: A Framework for Symbolic Music Graph Explanations标题:SMUG-Explain:符号音乐图形解释的框架链接:https://arxiv.org/abs/2405.09241作者:Emmanouil Karystinaios,Francesco Foscarin,Gerhard Widmer备注:In Proceedings of the Sound and Music Computing Conference 2024 (SMC2024), Porto, Portugal摘要:在这项工作中,我们提出了乐谱音乐图(SMUG)-解释,一个框架,用于生成和可视化的解释图神经网络应用于任意预测任务的乐谱。我们的系统允许用户直接在乐谱的上下文中可视化输入音符(和音符特征)对网络输出的贡献。我们提供了一个基于音乐符号雕刻库Verovio的交互式界面。我们展示了SMUG-Explain在古典音乐节奏检测任务中的使用。所有代码都可以在https://github.com/manoskary/SMUG-Explain上找到。摘要:In this work, we present Score MUsic Graph (SMUG)-Explain, a framework for generating and visualizing explanations of graph neural networks applied to arbitrary prediction tasks on musical scores. Our system allows the user to visualize the contribution of input notes (and note features) to the network output, directly in the context of the musical score. We provide an interactive interface based on the music notation engraving library Verovio. We showcase the usage of SMUG-Explain on the task of cadence detection in classical music. All code is available on https://github.com/manoskary/SMUG-Explain.
【4】 Perception-Inspired Graph Convolution for Music Understanding Tasks标题:音乐理解任务的感知启发图卷积链接:https://arxiv.org/abs/2405.09224作者:Emmanouil Karystinaios,Francesco Foscarin,Gerhard Widmer备注:Accepted at the 33rd International Joint Conference on Artificial Intelligence (IJCAI-24)摘要:我们提出了一个新的图卷积块,称为MusGConv,专门为有效处理乐谱数据而设计,并受到一般感知原则的激励。它侧重于音乐的两个基本维度,音高和节奏,并考虑这些组件的相对和绝对表示。我们评估我们的方法对四个不同的音乐理解问题:单声道的声音分离,谐波分析,节奏检测,作曲家识别,在抽象方面,转化为不同的图学习问题,即,节点分类,链接预测,和图形分类。我们的实验表明,MusGConv提高了上述三个任务的性能,同时在概念上非常简单和高效。我们将此解释为证据,即在乐谱数据上开发图形网络应用程序时,包括基本音乐概念的感知信息处理是有益的。摘要:We propose a new graph convolutional block, called MusGConv, specifically designed for the efficient processing of musical score data and motivated by general perceptual principles. It focuses on two fundamental dimensions of music, pitch and rhythm, and considers both relative and absolute representations of these components. We evaluate our approach on four different musical understanding problems: monophonic voice separation, harmonic analysis, cadence detection, and composer identification which, in abstract terms, translate to different graph learning problems, namely, node classification, link prediction, and graph classification. Our experiments demonstrate that MusGConv improves the performance on three of the aforementioned tasks while being conceptually very simple and efficient. We interpret this as evidence that it is beneficial to include perception-informed processing of fundamental musical concepts when developing graph network applications on musical score data.
【5】 Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis标题:文本到语音合成中的分层情绪预测与控制链接:https://arxiv.org/abs/2405.09171作者:Sho Inoue,Kun Zhou,Shuai Wang,Haizhou Li备注:This is accepted to IEEE ICASSP 2024摘要:如何有效地控制文本到语音(TTS)合成中的情感渲染仍然是一个挑战。先前的研究主要集中在学习一个全球的韵律表示在话语水平,这与语言韵律密切相关。我们的目标是构建一个分层的情感分布(ED),有效地封装在不同级别的粒度,包括音素,单词和话语的情感强度变化。在TTS训练期间,从地面实况音频中提取分层ED,并指导预测器在情感和语言韵律之间建立联系。在运行时的推理,TTS模型产生情感的语音,并在同一时间,提供定量控制的情绪的语音成分。客观和主观评价都验证了所提出的框架在情感预测和控制方面的有效性。摘要:It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with linguistic prosody. Our goal is to construct a hierarchical emotion distribution (ED) that effectively encapsulates intensity variations of emotions at various levels of granularity, encompassing phonemes, words, and utterances. During TTS training, the hierarchical ED is extracted from the ground-truth audio and guides the predictor to establish a connection between emotional and linguistic prosody. At run-time inference, the TTS model generates emotional speech and, at the same time, provides quantitative control of emotion over the speech constituents. Both objective and subjective evaluations validate the effectiveness of the proposed framework in terms of emotion prediction and control.
【6】 Naturalistic Music Decoding from EEG Data via Latent Diffusion Models标题:通过潜在扩散模型从脑电数据中解码自然主义音乐链接:https://arxiv.org/abs/2405.09062作者:Emilian Postolache,Natalia Polouliakh,Hiroaki Kitano,Akima Connelly,Emanuele Rodolà,Taketo Akama摘要:在这篇文章中,我们探讨了潜在的使用潜在的扩散模型,一个家庭的强大的生成模型,从脑电图(EEG)记录重建自然主义音乐的任务。与音色有限的简单音乐(如MIDI生成的曲调或单声道作品)不同,这里的重点是复杂的音乐,包括各种乐器,声音和效果,丰富的和声和音色。这项研究代表了使用非侵入性EEG数据实现高质量一般音乐重建的初步尝试,直接在原始数据上采用端到端训练方法,而无需手动预处理和通道选择。我们在公共NMED-T数据集上训练我们的模型,并进行定量评估,提出基于神经嵌入的指标。我们还根据生成的曲目进行歌曲分类。我们的工作有助于神经解码和脑-机接口的持续研究,为使用EEG数据进行复杂听觉信息重建的可行性提供了见解。摘要:In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike simpler music with limited timbres, such as MIDI-generated tunes or monophonic pieces, the focus here is on intricate music featuring a diverse array of instruments, voices, and effects, rich in harmonics and timbre. This study represents an initial foray into achieving general music reconstruction of high-quality using non-invasive EEG data, employing an end-to-end training approach directly on raw data without the need for manual pre-processing and channel selection. We train our models on the public NMED-T dataset and perform quantitative evaluation proposing neural embedding-based metrics. We additionally perform song classification based on the generated tracks. Our work contributes to the ongoing research in neural decoding and brain-computer interfaces, offering insights into the feasibility of using EEG data for complex auditory information reconstruction.
【7】 PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset标题:PolyGlotFake:一种新型的多语言和多模式DeepFake数据集链接:https://arxiv.org/abs/2405.08838作者:Yang Hou,Haitao Fu,Chuankai Chen,Zida Li,Haoyu Zhang,Jianjun Zhao备注:13 page, 4 figures摘要:随着生成式人工智能的快速发展,操纵音频和视觉模式的多模态deepfakes引起了越来越多的公众关注。目前,deepfake检测已成为应对这些日益增长的威胁的关键策略。然而,作为训练和验证deepfake检测器的关键因素,大多数现有的deepfake数据集主要集中在视觉模式上,少数多模态数据集采用过时的技术,其音频内容仅限于单一语言,因此无法代表当前deepfake技术的前沿进步和全球化趋势。为了解决这一差距,我们提出了一个新颖的、多语言的、多模式的deepfake数据集:PolyGlotFake。它包括七种语言的内容,使用各种尖端和流行的文本到语音,语音克隆和嘴唇同步技术创建。我们使用PolyGlotFake数据集上最先进的检测方法进行全面的实验。这些实验证明了数据集的重大挑战及其在推进多模态deepfake检测研究方面的实用价值。摘要:With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these growing threats. However, as a key factor in training and validating deepfake detectors, most existing deepfake datasets primarily focus on the visual modal, and the few that are multimodal employ outdated techniques, and their audio content is limited to a single language, thereby failing to represent the cutting-edge advancements and globalization trends in current deepfake technologies. To address this gap, we propose a novel, multilingual, and multimodal deepfake dataset: PolyGlotFake. It includes content in seven languages, created using a variety of cutting-edge and popular Text-to-Speech, voice cloning, and lip-sync technologies. We conduct comprehensive experiments using state-of-the-art detection methods on PolyGlotFake dataset. These experiments demonstrate the dataset's significant challenges and its practical value in advancing research into multimodal deepfake detection.
【8】 Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization标题:使用弱监督语音活动检测嵌入说话人以实现有效的说话人Diary化链接:https://arxiv.org/abs/2405.09142作者:Jenthe Thienpondt,Kris Demuynck备注:Proceedings of Odyssey 2024: The Speaker and Language Recognition Workshop摘要:当前的说话人日志化系统依赖于在检测到的语音段上的说话人嵌入提取之前的外部语音活动检测模型。在本文中,我们建立了一个扬声器嵌入提取器的注意力系统作为一个弱监督的内部VAD模型,并表现得等同于或优于可比的监督VAD系统。随后,说话人日记可以有效地执行提取VAD logits和相应的说话人嵌入的同时,减轻需要和计算开销的外部VAD模型。我们提供了一个广泛的分析的帧级注意力系统在当前的说话人验证模型的行为,并提出了一种新的说话人日记流水线使用ECAPA2说话人嵌入VAD和嵌入提取。所提出的策略在AMI,VoxConverse和DIHARD III日志基准测试中获得了最先进的性能。摘要:Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding extractor acts as a weakly supervised internal VAD model and performs equally or better than comparable supervised VAD systems. Subsequently, speaker diarization can be performed efficiently by extracting the VAD logits and corresponding speaker embedding simultaneously, alleviating the need and computational overhead of an external VAD model. We provide an extensive analysis of the behavior of the frame-level attention system in current speaker verification models and propose a novel speaker diarization pipeline using ECAPA2 speaker embeddings for both VAD and embedding extraction. The proposed strategy gains state-of-the-art performance on the AMI, VoxConverse and DIHARD III diarization benchmarks. eess.AS音频处理【1】 Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization标题:使用弱监督语音活动检测嵌入说话人以实现有效的说话人Diary化链接:https://arxiv.org/abs/2405.09142作者:Jenthe Thienpondt,Kris Demuynck备注:Proceedings of Odyssey 2024: The Speaker and Language Recognition Workshop摘要:当前的说话人日志化系统依赖于在检测到的语音段上的说话人嵌入提取之前的外部语音活动检测模型。在本文中,我们建立了一个扬声器嵌入提取器的注意力系统作为一个弱监督的内部VAD模型,并表现得等同于或优于可比的监督VAD系统。随后,说话人日记可以有效地执行提取VAD logits和相应的说话人嵌入的同时,减轻需要和计算开销的外部VAD模型。我们提供了一个广泛的分析的帧级注意力系统在当前的说话人验证模型的行为,并提出了一种新的说话人日记流水线使用ECAPA2说话人嵌入VAD和嵌入提取。所提出的策略在AMI,VoxConverse和DIHARD III日志基准测试中获得了最先进的性能。摘要:Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding extractor acts as a weakly supervised internal VAD model and performs equally or better than comparable supervised VAD systems. Subsequently, speaker diarization can be performed efficiently by extracting the VAD logits and corresponding speaker embedding simultaneously, alleviating the need and computational overhead of an external VAD model. We provide an extensive analysis of the behavior of the frame-level attention system in current speaker verification models and propose a novel speaker diarization pipeline using ECAPA2 speaker embeddings for both VAD and embedding extraction. The proposed strategy gains state-of-the-art performance on the AMI, VoxConverse and DIHARD III diarization benchmarks. 【2】 Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer标题:通过音频风格转移评估自动语音识别系统的稳健性链接:https://arxiv.org/abs/2405.09470作者:Weifei Jin,Yuxin Cao,Junjie Su,Qi Shen,Kai Ye,Derui Wang,Jie Hao,Ziyao Liu备注:Accepted to SecTL (AsiaCCS Workshop) 2024摘要:鉴于自动语音识别(ASR)系统的广泛应用,其安全问题比以往任何时候都受到了更多的关注,这主要是由于深度神经网络的敏感性。以前的研究表明,秘密制作对抗性扰动可以操纵语音识别系统,从而产生恶意命令。这些攻击方法大多需要在$\ell_p $范数约束下添加噪声扰动,不可避免地留下手动修改的伪影。最近的研究通过操纵风格向量来基于文本到语音(TTS)合成音频合成对抗性示例,从而缓解了这一限制。然而,基于优化目标的风格修改显著降低了音频风格的可控性和可编辑性。在本文中,我们提出了一种攻击ASR系统的基础上,用户定制的风格转移。我们首先测试了风格转移攻击(STA)的效果,它将风格转移和对抗性攻击按顺序结合起来。然后,作为改进,我们提出了一个迭代的风格代码攻击(SCA),以保持音频质量。实验结果表明,该方法可以满足用户定制风格的需要,并取得了82%的攻击成功率,同时保持声音的自然性,由于我们的用户学习。摘要:In light of the widespread application of Automatic Speech Recognition (ASR) systems, their security concerns have received much more attention than ever before, primarily due to the susceptibility of Deep Neural Networks. Previous studies have illustrated that surreptitiously crafting adversarial perturbations enables the manipulation of speech recognition systems, resulting in the production of malicious commands. These attack methods mostly require adding noise perturbations under $\ell_p$ norm constraints, inevitably leaving behind artifacts of manual modifications. Recent research has alleviated this limitation by manipulating style vectors to synthesize adversarial examples based on Text-to-Speech (TTS) synthesis audio. However, style modifications based on optimization objectives significantly reduce the controllability and editability of audio styles. In this paper, we propose an attack on ASR systems based on user-customized style transfer. We first test the effect of Style Transfer Attack (STA) which combines style transfer and adversarial attack in sequential order. And then, as an improvement, we propose an iterative Style Code Attack (SCA) to maintain audio quality. Experimental results show that our method can meet the need for user-customized styles and achieve a success rate of 82% in attacks, while keeping sound naturalness due to our user study. 【3】 Dance Any Beat: Blending Beats with Visuals in Dance Video Generation标题:舞蹈任意节拍:在舞蹈视频生成中将节拍与视觉效果融为一体链接:https://arxiv.org/abs/2405.09266作者:Xuanchen Wang,Heng Wang,Dongnan Liu,Weidong Cai备注:11 pages, 6 figures, demo page: this https URL摘要:从音乐生成舞蹈的任务是至关重要的,但目前的方法,主要产生联合序列,导致输出缺乏直观性和复杂的数据收集,由于精确的联合注释的必要性。我们介绍了一个舞蹈任何节拍扩散模型,即DabFusion,采用音乐作为有条件的输入,直接创建舞蹈视频从静止图像,利用有条件的图像到视频生成原理。这种方法开创了使用音乐作为图像到视频合成的调节因素。我们的方法分为两个阶段:训练自动编码器来预测参考帧和驱动帧之间的潜在光流,消除对联合注释的需要,并训练基于U-Net的扩散模型来产生这些由CLAP编码的音乐节奏引导的潜在光流。虽然能够制作高质量的舞蹈视频,但基线模型在节奏调整方面很困难。我们通过增加节拍信息,提高同步性来增强模型。我们引入了一个2D运动音乐对齐分数(2D-MM对齐)的定量评估。在AIST++数据集上进行评估,我们的增强模型显示出2D-MM Align评分和既定指标的显着改善。视频结果可以在我们的项目页面上找到:https://DabFusion.github.io。摘要:The task of generating dance from music is crucial, yet current methods, which mainly produce joint sequences, lead to outputs that lack intuitiveness and complicate data collection due to the necessity for precise joint annotations. We introduce a Dance Any Beat Diffusion model, namely DabFusion, that employs music as a conditional input to directly create dance videos from still images, utilizing conditional image-to-video generation principles. This approach pioneers the use of music as a conditioning factor in image-to-video synthesis. Our method unfolds in two stages: training an auto-encoder to predict latent optical flow between reference and driving frames, eliminating the need for joint annotation, and training a U-Net-based diffusion model to produce these latent optical flows guided by music rhythm encoded by CLAP. Although capable of producing high-quality dance videos, the baseline model struggles with rhythm alignment. We enhance the model by adding beat information, improving synchronization. We introduce a 2D motion-music alignment score (2D-MM Align) for quantitative assessment. Evaluated on the AIST++ dataset, our enhanced model shows marked improvements in 2D-MM Align score and established metrics. Video results can be found on our project page: https://DabFusion.github.io.
【4】 SMUG-Explain: A Framework for Symbolic Music Graph Explanations标题:SMUG-Explain:符号音乐图形解释的框架链接:https://arxiv.org/abs/2405.09241作者:Emmanouil Karystinaios,Francesco Foscarin,Gerhard Widmer备注:In Proceedings of the Sound and Music Computing Conference 2024 (SMC2024), Porto, Portugal摘要:在这项工作中,我们提出了乐谱音乐图(SMUG)—解释,一个框架,用于生成和可视化的解释图神经网络应用于任意预测任务的乐谱。我们的系统允许用户直接在乐谱的上下文中可视化输入音符(和音符特征)对网络输出的贡献。我们提供了一个基于音乐符号雕刻库Verovio的交互式界面。我们展示了SMUG—Explain在古典音乐节奏检测任务中的使用。所有代码都可以在www.example.com上找到。摘要:In this work, we present Score MUsic Graph (SMUG)-Explain, a framework for generating and visualizing explanations of graph neural networks applied to arbitrary prediction tasks on musical scores. Our system allows the user to visualize the contribution of input notes (and note features) to the network output, directly in the context of the musical score. We provide an interactive interface based on the music notation engraving library Verovio. We showcase the usage of SMUG-Explain on the task of cadence detection in classical music. All code is available on https://github.com/manoskary/SMUG-Explain.
【5】 Perception-Inspired Graph Convolution for Music Understanding Tasks标题:音乐理解任务的感知启发图卷积链接:https://arxiv.org/abs/2405.09224作者:Emmanouil Karystinaios,Francesco Foscarin,Gerhard Widmer备注:Accepted at the 33rd International Joint Conference on Artificial Intelligence (IJCAI-24)摘要:我们提出了一个新的图卷积块,称为MusGConv,专门为有效处理乐谱数据而设计,并受到一般感知原则的激励。它侧重于音乐的两个基本维度,音高和节奏,并考虑这些组件的相对和绝对表示。我们评估我们的方法对四个不同的音乐理解问题:单声道的声音分离,谐波分析,节奏检测,作曲家识别,在抽象方面,转化为不同的图学习问题,即,节点分类,链接预测,和图形分类。我们的实验表明,MusGConv提高了上述三个任务的性能,同时在概念上非常简单和高效。我们将此解释为证据,即在乐谱数据上开发图形网络应用程序时,包括基本音乐概念的感知信息处理是有益的。摘要:We propose a new graph convolutional block, called MusGConv, specifically designed for the efficient processing of musical score data and motivated by general perceptual principles. It focuses on two fundamental dimensions of music, pitch and rhythm, and considers both relative and absolute representations of these components. We evaluate our approach on four different musical understanding problems: monophonic voice separation, harmonic analysis, cadence detection, and composer identification which, in abstract terms, translate to different graph learning problems, namely, node classification, link prediction, and graph classification. Our experiments demonstrate that MusGConv improves the performance on three of the aforementioned tasks while being conceptually very simple and efficient. We interpret this as evidence that it is beneficial to include perception-informed processing of fundamental musical concepts when developing graph network applications on musical score data.
【6】 Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis标题:文本到语音合成中的分层情绪预测与控制链接:https://arxiv.org/abs/2405.09171作者:Sho Inoue,Kun Zhou,Shuai Wang,Haizhou Li备注:This is accepted to IEEE ICASSP 2024摘要:如何有效地控制文本到语音(TTS)合成中的情感渲染仍然是一个挑战。先前的研究主要集中在学习一个全球的韵律表示在话语水平,这与语言韵律密切相关。我们的目标是构建一个分层的情感分布(ED),有效地封装在不同级别的粒度,包括音素,单词和话语的情感强度变化。在TTS训练期间,从地面实况音频中提取分层ED,并指导预测器在情感和语言韵律之间建立联系。在运行时的推理,TTS模型产生情感的语音,并在同一时间,提供定量控制的情绪的语音成分。客观和主观评价都验证了所提出的框架在情感预测和控制方面的有效性。摘要:It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with linguistic prosody. Our goal is to construct a hierarchical emotion distribution (ED) that effectively encapsulates intensity variations of emotions at various levels of granularity, encompassing phonemes, words, and utterances. During TTS training, the hierarchical ED is extracted from the ground-truth audio and guides the predictor to establish a connection between emotional and linguistic prosody. At run-time inference, the TTS model generates emotional speech and, at the same time, provides quantitative control of emotion over the speech constituents. Both objective and subjective evaluations validate the effectiveness of the proposed framework in terms of emotion prediction and control.
【7】 Naturalistic Music Decoding from EEG Data via Latent Diffusion Models标题:通过潜在扩散模型从脑电数据中解码自然主义音乐链接:https://arxiv.org/abs/2405.09062作者:Emilian Postolache,Natalia Polouliakh,Hiroaki Kitano,Akima Connelly,Emanuele Rodolà,Taketo Akama摘要:在这篇文章中,我们探讨了潜在的使用潜在的扩散模型,一个家庭的强大的生成模型,从脑电图(EEG)记录重建自然主义音乐的任务。与音色有限的简单音乐(如MIDI生成的曲调或单声道作品)不同,这里的重点是复杂的音乐,具有多种乐器,声音和效果,谐波和音色丰富。这项研究代表了使用非侵入性EEG数据实现高质量一般音乐重建的初步尝试,直接在原始数据上采用端到端训练方法,而无需手动预处理和通道选择。我们在公共NMED-T数据集上训练我们的模型,并进行定量评估,提出基于神经嵌入的指标。我们还根据生成的曲目进行歌曲分类。我们的工作有助于神经解码和脑-机接口的持续研究,为使用EEG数据进行复杂听觉信息重建的可行性提供了见解。摘要:In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike simpler music with limited timbres, such as MIDI-generated tunes or monophonic pieces, the focus here is on intricate music featuring a diverse array of instruments, voices, and effects, rich in harmonics and timbre. This study represents an initial foray into achieving general music reconstruction of high-quality using non-invasive EEG data, employing an end-to-end training approach directly on raw data without the need for manual pre-processing and channel selection. We train our models on the public NMED-T dataset and perform quantitative evaluation proposing neural embedding-based metrics. We additionally perform song classification based on the generated tracks. Our work contributes to the ongoing research in neural decoding and brain-computer interfaces, offering insights into the feasibility of using EEG data for complex auditory information reconstruction. 【8】 PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset标题:PolyGlotFake:一种新型的多语言和多模式DeepFake数据集链接:https://arxiv.org/abs/2405.08838作者:Yang Hou,Haitao Fu,Chuankai Chen,Zida Li,Haoyu Zhang,Jianjun Zhao备注:13 page, 4 figures摘要:随着生成式人工智能的快速发展,操纵音频和视觉模式的多模态deepfakes引起了越来越多的公众关注。目前,deepfake检测已成为应对这些日益增长的威胁的关键策略。然而,作为训练和验证deepfake检测器的关键因素,大多数现有的deepfake数据集主要集中在视觉模式上,少数多模态数据集采用过时的技术,其音频内容仅限于单一语言,因此无法代表当前deepfake技术的前沿进步和全球化趋势。为了解决这一差距,我们提出了一个新颖的、多语言的、多模式的deepfake数据集:PolyGlotFake。它包括七种语言的内容,使用各种尖端和流行的文本到语音,语音克隆和嘴唇同步技术创建。我们使用PolyGlotFake数据集上最先进的检测方法进行全面的实验。这些实验证明了数据集的重大挑战及其在推进多模态deepfake检测研究方面的实用价值。摘要:With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these growing threats. However, as a key factor in training and validating deepfake detectors, most existing deepfake datasets primarily focus on the visual modal, and the few that are multimodal employ outdated techniques, and their audio content is limited to a single language, thereby failing to represent the cutting-edge advancements and globalization trends in current deepfake technologies. To address this gap, we propose a novel, multilingual, and multimodal deepfake dataset: PolyGlotFake. It includes content in seven languages, created using a variety of cutting-edge and popular Text-to-Speech, voice cloning, and lip-sync technologies. We conduct comprehensive experiments using state-of-the-art detection methods on PolyGlotFake dataset. These experiments demonstrate the dataset's significant challenges and its practical value in advancing research into multimodal deepfake detection.