本文经arXiv每日学术速递授权转载
【1】Musical composition and 2D cellular automata based on music intervals
链接:https://arxiv.org/abs/2411.19844
备注:17 pages, 3 figures
摘要:本研究是一种理论方法,探索的适用性的2D元胞自动机的基础上旋律和谐波间隔的随机阵列的音符。本研究的目的是探讨元胞自动机在音乐情境中的替代用途,以更好地理解音乐创造力。我们以复杂系统和人文方法为框架,以音乐理论的规律为基础,捕捉音乐创作的本质。研究结果表明,这些规则的问题产生大规模的模式有组织的笔记。因此,我们的配方提供了一种新的方法来理解和复制方面的音乐创造力。
摘要:This study is a theoretical approach for exploring the applicability of a 2Dcellular automaton based on melodic and harmonic intervals in random arrays ofmusical notes. The aim of this study was to explore alternatives uses for acellular automaton in the musical context for better understanding the musicalcreativity. We used the complex systems and humanities approaches as aframework for capturing the essence of creating music based on rules of musictheory. Findings suggested that such rules matter for generating large-scalepatterns of organized notes. Therefore, our formulation provides a novelapproach for understanding and replicating aspects of the musical creativity.
标题:并行堆叠聚合网络用于支持物联网的智能设备中的语音认证
链接:https://arxiv.org/abs/2411.19841
备注:arXiv admin note: text overlap with arXiv:2309.10560
摘要:近年来,由于对用户隐私和安全的担忧日益增加,物联网智能设备上的语音认证已经变得突出。当前的认证系统容易受到不同的语音欺骗攻击(例如,重放、语音克隆和音频深度伪造),其模仿合法语音以欺骗认证系统并实现欺诈活动(例如,假冒、未经授权的访问、金融欺诈等)。现有的解决方案通常被设计为应对单一类型的攻击,从而导致针对不可见攻击的性能受损。另一方面,现有的统一语音反欺骗解决方案不是专门为物联网设计的,具有复杂的架构,因此无法部署在支持物联网的智能设备上。此外,这些统一解决方案中的大多数都表现出严重的性能问题,包括更高的等错误率或特定攻击的更低准确性。为了克服这些问题,我们提出了并行堆栈聚合网络(PSA-Net),这是一个轻量级框架,旨在为语音控制的智能物联网设备提供反欺骗防御系统。PSA-Net直接处理原始音频,无需依赖于音频的手工制作功能或预先计算的频谱图。此外,PSA-Net采用分裂-变换-聚合方法,该方法涉及话语的分割,通过卷积提取内在可微嵌入,以及将它们聚合以区分合法音频和欺骗音频。与现有的面向Resnet的深度解决方案相比,我们将基数作为网络中的一个额外维度,这增强了PSA-Net在不同攻击中的泛化能力。结果表明,PSA-Net对当前反欺骗解决方案中存在的不同攻击实现了更一致的性能。
摘要:Voice authentication on IoT-enabled smart devices has gained prominence inrecent years due to increasing concerns over user privacy and security. Thecurrent authentication systems are vulnerable to different voice-spoofingattacks (e.g., replay, voice cloning, and audio deepfakes) that mimiclegitimate voices to deceive authentication systems and enable fraudulentactivities (e.g., impersonation, unauthorized access, financial fraud, etc.).Existing solutions are often designed to tackle a single type of attack,leading to compromised performance against unseen attacks. On the other hand,existing unified voice anti-spoofing solutions, not designed specifically forIoT, possess complex architectures and thus cannot be deployed on IoT-enabledsmart devices. Additionally, most of these unified solutions exhibitsignificant performance issues, including higher equal error rates or loweraccuracy for specific attacks. To overcome these issues, we present theparallel stacked aggregation network (PSA-Net), a lightweight frameworkdesigned as an anti-spoofing defense system for voice-controlled smart IoTdevices. The PSA-Net processes raw audios directly and eliminates the need fordataset-dependent handcrafted features or pre-computed spectrograms.Furthermore, PSA-Net employs a split-transform-aggregate approach, whichinvolves the segmentation of utterances, the extraction of intrinsicdifferentiable embeddings through convolutions, and the aggregation of them todistinguish legitimate from spoofed audios. In contrast to existing deepResnet-oriented solutions, we incorporate cardinality as an additionaldimension in our network, which enhances the PSA-Net ability to generalizeacross diverse attacks. The results show that the PSA-Net achieves moreconsistent performance for different attacks that exist in currentanti-spoofing solutions.
标题:采用关节嵌入预测架构的Zero-Shot音乐主干检索
链接:https://arxiv.org/abs/2411.19806
备注:Submitted to ICASSP 2025
摘要:在本文中,我们解决的任务,音乐干检索。给定一个音乐组合,它包括检索一个与之匹配的词干,即,一起演奏的话会很好听为此,我们引入了一种基于联合嵌入预测架构的新方法,其中编码器和预测器被联合训练以产生上下文的潜在表示并预测目标的潜在表示。特别是,我们设计我们的预测器是以任意仪器为条件的,使我们的模型能够执行zero-shot干检索。此外,我们发现使用对比学习对编码器进行预训练大大提高了模型的性能。 我们使用MUSDB18和MoisesDB数据集验证了我们的模型的检索性能。我们表明,它在两个数据集上的表现都明显优于以前的基线,展示了它支持或多或少精确(可能是看不见的)条件的能力。我们还评估了学习嵌入的节拍跟踪任务,证明他们保留时间结构和本地信息。
摘要:In this paper, we tackle the task of musical stem retrieval. Given a musicalmix, it consists in retrieving a stem that would fit with it, i.e., that wouldsound pleasant if played together. To do so, we introduce a new method based onJoint-Embedding Predictive Architectures, where an encoder and a predictor arejointly trained to produce latent representations of a context and predictlatent representations of a target. In particular, we design our predictor tobe conditioned on arbitrary instruments, enabling our model to performzero-shot stem retrieval. In addition, we discover that pretraining the encoderusing contrastive learning drastically improves the model's performance. We validate the retrieval performances of our model using the MUSDB18 andMoisesDB datasets. We show that it significantly outperforms previous baselineson both datasets, showcasing its ability to support more or less precise (andpossibly unseen) conditioning. We also evaluate the learned embeddings on abeat tracking task, demonstrating that they retain temporal structure and localinformation.
标题:一种基于监督对比学习的跨数据库语音情感识别方法
链接:https://arxiv.org/abs/2411.19803
摘要:语音情感识别(SER)的研究往往面临着缺乏大规模公共数据集和处理不同分布数据时泛化能力有限等挑战。针对这一问题,提出一种基于监督对比学习的跨语料语音情感识别方法。该方法采用两阶段的微调过程:首先,使用多个语音情感数据集上的监督对比学习微调自监督语音表示模型;然后,在目标数据集上微调分类器。实验结果表明,基于WavLM的模型在IEMOCAP数据集和CASIA数据集上的未加权准确率(UA)分别为77.41%和96.49%,优于这两个数据集上的最新结果。
摘要:Research on Speech Emotion Recognition (SER) often faces challenges such asthe lack of large-scale public datasets and limited generalization capabilitywhen dealing with data from different distributions. To solve this problem,this paper proposes a cross-corpus speech emotion recognition method based onsupervised contrast learning. The method employs a two-stage fine-tuningprocess: first, the self-supervised speech representation model is fine-tunedusing supervised contrastive learning on multiple speech emotion datasets;then, the classifier is fine-tuned on the target dataset. The experimentalresults show that the WavLM-based model achieved unweighted accuracy (UA) of77.41% on the IEMOCAP dataset and 96.49% on the CASIA dataset, outperformingthe state-of-the-art results on the two datasets.
标题:电子竞技中的语音沟通分析
链接:https://arxiv.org/abs/2411.19793
备注:17 pages, 11 figures. Independent research
摘要:在大多数基于团队的电子竞技中,语音通信在团队效率和协同作用方面非常突出。事实上,已经观察到,当试图在正式比赛中有良好表现时,不仅团队的技能方面,而且团队有效的语音通信也会发挥作用。随着最近出现的LLM(大型语言模型)工具关于NLP(自然语言处理)(Vaswani et.等),我们决定尝试应用它们,以便更好地了解如何提高语音通信的有效性。本文以《英雄联盟》电子竞技为视角进行研究。然而,主要的概念和想法可以很容易地应用于任何其他团队相关的电子竞技。
摘要:In most team-based esports, voice communications are prominent in the teamefficiency and synergy. In fact it has been observed that not only the skillaspect of the team but also the team effective voice communication comes intoplay when trying to have good performance in official matches. With the recentemergence of LLM (Large Language Models) tools regarding NLP (Natural LanguageProcessing) (Vaswani et. al.), we decided to try applying them in order to havea better understanding on how to improve the effectiveness of the voicecommunications. In this paper the study has been made through the prism ofLeague of Legends esport. However the main concepts and ideas can be easilyapplicable in any other team related esports.
标题:Noro:具有隐藏说话者表示能力的噪音稳健单次语音转换系统
链接:https://arxiv.org/abs/2411.19770
备注:Submitted to IEEE OJSP
摘要:单次语音转换(One-shot voice conversion,VC)的目的是在保持原语音语义的前提下,仅使用目标语音中的一个参考语音,将源语音的音色转换为目标语音的音色。尽管一次性VC取得了进步,但在现实世界的场景中,其有效性会下降,其中参考语音通常来自互联网,包含各种干扰,如背景噪声。为了解决这个问题,我们引入了Noro,一种噪声鲁棒的单次VC系统。Noro具有专为VC使用噪声参考语音量身定制的创新组件,包括双分支参考编码模块和与噪声无关的对比扬声器损失。实验结果表明,Noro优于我们的基线系统在干净和嘈杂的情况下,突出了其在现实世界中的应用的功效。此外,我们调查了隐藏的扬声器表示能力,我们的基线系统重新利用其参考编码器作为扬声器编码器。实验结果表明,在SUPERB环境下,该方法与几种先进的自监督学习模型相比具有较强的竞争力,突出了通过一次性VC任务推进说话人表征学习的潜力。
摘要:One-shot voice conversion (VC) aims to alter the timbre of speech from asource speaker to match that of a target speaker using just a single referencespeech from the target, while preserving the semantic content of the originalsource speech. Despite advancements in one-shot VC, its effectiveness decreasesin real-world scenarios where reference speeches, often sourced from theinternet, contain various disturbances like background noise. To address thisissue, we introduce Noro, a Noise Robust One-shot VC system. Noro featuresinnovative components tailored for VC using noisy reference speeches, includinga dual-branch reference encoding module and a noise-agnostic contrastivespeaker loss. Experimental results demonstrate that Noro outperforms ourbaseline system in both clean and noisy scenarios, highlighting its efficacyfor real-world applications. Additionally, we investigate the hidden speakerrepresentation capabilities of our baseline system by repurposing its referenceencoder as a speaker encoder. The results shows that it is competitive withseveral advanced self-supervised learning models for speaker representationunder the SUPERB settings, highlighting the potential for advancing speakerrepresentation learning through one-shot VC task.
标题:用于节能音频分类的记忆纳米线网络:无需预处理、缩短延迟的水库计算
链接:https://arxiv.org/abs/2411.19611
备注:17 pages, 6 Figures
摘要:语音识别是自然语言处理中的一个关键挑战,要求实时应用具有低延迟、高效计算和强泛化能力。虽然基于软件的人工神经网络(ANN)擅长这项任务,但它们是计算密集型的,并且严重依赖于数据预处理。神经形态计算以其低延迟和节能的优势,有望用于音频分类。忆阻纳米线网络与梅尔频率倒谱系数提取等预处理技术相结合,已被广泛用于联想学习,但这种预处理可能是功耗密集型的,破坏了延迟优势。这项研究开创了使用纳米线网络的忆阻和时空特性进行音频信号分类,而无需预处理。纳米线网络仿真与三个线性分类器配对,用于10类MNIST音频分类和二进制扬声器泛化测试。该混合系统实现了显著的优势:出色的数据压缩,仅利用3%的纳米线输出,计算延迟减少10倍,分类准确性提高高达28.5%(使用逻辑回归分类器)。与原始数据分类器相比,多说话人数据集的准确率和召回率分别提高了10%和17%,单个说话人数据集的准确率和召回率分别提高了24%和17%。这项工作为在边缘计算设备中利用忆阻纳米线网络(NWN)提供了基本的概念证明,展示了它们在有效实时音频信号处理方面的潜力,同时降低了计算开销和功耗,并实现先进神经形态计算解决方案的开发。
摘要:Speech recognition is a key challenge in natural language processing,requiring low latency, efficient computation, and strong generalization forreal-time applications. While software-based artificial neural networks (ANNs)excel at this task, they are computationally intensive and depend heavily ondata pre-processing. Neuromorphic computing, with its low-latency andenergy-efficient advantages, holds promise for audio classification. Memristivenanowire networks, combined with pre-processing techniques like Mel-FrequencyCepstrum Coefficient extraction, have been widely used for associativelearning, but such pre-processing can be power-intensive, undermining latencybenefits. This study pioneers the use of memristive and spatio-temporalproperties of nanowire networks for audio signal classification withoutpre-processing. A nanowire network simulation is paired with three linearclassifiers for 10-class MNIST audio classification and binary speakergeneralization tests. The hybrid system achieves significant benefits:excellent data compression with only 3% of nanowire output utilized, a 10-foldreduction in computational latency, and up to 28.5% improved classificationaccuracy (using a logistic regression classifier). Precision and recall improveby 10% and 17% for multispeaker datasets, and by 24% and 17% for individualspeaker datasets, compared to raw data classifiers.This work provides afoundational proof of concept for utilizing memristive nanowire networks (NWN)in edge-computing devices, showcasing their potential for efficient, real-timeaudio signal processing with reduced computational overhead and powerconsumption, and enabling the development of advanced neuromorphic computingsolutions.
标题:生成人工智能时代的Deepfake媒体生成和检测:调查和展望
链接:https://arxiv.org/abs/2411.19537
摘要:随着生成建模的最新进展,deepfake内容的真实性一直在稳步增长,甚至达到了人们经常无法在线检测到被操纵的媒体内容的程度,从而被欺骗到各种骗局中。在本文中,我们综述了deepfake生成和检测技术,包括该领域的最新发展,如扩散模型和神经辐射场。我们的文献综述涵盖了所有deepfake媒体类型,包括图像、视频、音频和多模式(视听)内容。我们根据用于更改或生成虚假内容的程序来识别各种类型的deepfake。我们进一步构建了deepfake生成和检测方法的分类,说明了重要的方法组以及这些方法的应用领域。接下来,我们收集用于deepfake检测的数据集,并在最流行的数据集上提供最佳deepfake检测器的更新排名。此外,我们还开发了一种新的多模态基准测试,以评估deepfake检测器对分发内容的影响。结果表明,最先进的检测器无法推广到由看不见的深度伪造生成器生成的深度伪造内容。最后,我们提出了未来的方向,以获得强大而强大的深度伪造检测器。我们的项目页面和新的基准可以在https://github.com/CroitoruAlin/biodeep上找到。
摘要:With the recent advancements in generative modeling, the realism of deepfakecontent has been increasing at a steady pace, even reaching the point wherepeople often fail to detect manipulated media content online, thus beingdeceived into various kinds of scams. In this paper, we survey deepfakegeneration and detection techniques, including the most recent developments inthe field, such as diffusion models and Neural Radiance Fields. Our literaturereview covers all deepfake media types, comprising image, video, audio andmultimodal (audio-visual) content. We identify various kinds of deepfakes,according to the procedure used to alter or generate the fake content. Wefurther construct a taxonomy of deepfake generation and detection methods,illustrating the important groups of methods and the domains where thesemethods are applied. Next, we gather datasets used for deepfake detection andprovide updated rankings of the best performing deepfake detectors on the mostpopular datasets. In addition, we develop a novel multimodal benchmark toevaluate deepfake detectors on out-of-distribution content. The resultsindicate that state-of-the-art detectors fail to generalize to deepfake contentgenerated by unseen deepfake generators. Finally, we propose future directionsto obtain robust and powerful deepfake detectors. Our project page and newbenchmark are available at https://github.com/CroitoruAlin/biodeep.
标题:同图:可控实时说话头部合成的运动空间扩散
链接:https://arxiv.org/abs/2411.19509
摘要:扩散模型的最新进展彻底改变了音频驱动的说话头合成。除了精确的嘴唇同步,基于扩散的方法在生成与音频信号良好对齐的微妙表情和自然头部运动方面表现出色。然而,这些方法都面临着缓慢的推理速度,不足的细粒度控制面部运动,偶尔的视觉伪影,主要是由于隐式的潜在空间来自变分自动编码器(VAE),这阻止了他们在实时交互应用中的采用。为了解决这些问题,我们引入Ditto,一个基于扩散的框架,使可控的实时说话头合成。我们的关键创新在于通过一个明确的身份不可知的运动空间,取代传统的VAE表示,桥接运动生成和真实感神经渲染。这种设计大大降低了扩散学习的复杂性,同时能够精确控制合成的说话头。我们进一步提出了一个推理策略,共同优化三个关键组成部分:音频特征提取,运动生成和视频合成。这种优化实现了流处理、实时推理和低第一帧延迟,这些功能对于AI助手等交互式应用至关重要。大量的实验结果表明,Ditto生成引人注目的说话头部视频,并大大优于现有的方法在运动控制和实时性能。
摘要:Recent advances in diffusion models have revolutionized audio-driven talkinghead synthesis. Beyond precise lip synchronization, diffusion-based methodsexcel in generating subtle expressions and natural head movements that arewell-aligned with the audio signal. However, these methods are confronted byslow inference speed, insufficient fine-grained control over facial motions,and occasional visual artifacts largely due to an implicit latent space derivedfrom Variational Auto-Encoders (VAE), which prevent their adoption in realtimeinteraction applications. To address these issues, we introduce Ditto, adiffusion-based framework that enables controllable realtime talking headsynthesis. Our key innovation lies in bridging motion generation andphotorealistic neural rendering through an explicit identity-agnostic motionspace, replacing conventional VAE representations. This design substantiallyreduces the complexity of diffusion learning while enabling precise controlover the synthesized talking heads. We further propose an inference strategythat jointly optimizes three key components: audio feature extraction, motiongeneration, and video synthesis. This optimization enables streamingprocessing, realtime inference, and low first-frame delay, which are thefunctionalities crucial for interactive applications such as AI assistants.Extensive experimental results demonstrate that Ditto generates compellingtalking head videos and substantially outperforms existing methods in bothmotion control and realtime performance.
标题:V2 SFlow:具有语音分解和纠正流的视频转语音生成
链接:https://arxiv.org/abs/2411.19486
摘要:在本文中,我们介绍了V2SFlow,一种新颖的视频到语音(V2S)框架,旨在直接从无声的说话人脸视频生成自然和可理解的语音。虽然最近的V2S系统在具有有限扬声器和词汇的约束数据集上显示出有希望的结果,但由于语音信号的固有可变性和复杂性,它们的性能在现实世界的无约束数据集上通常会下降。为了解决这些挑战,我们将语音信号分解为可管理的子空间(内容,音高和说话人信息),每个子空间代表不同的语音属性,并直接从视觉输入中预测它们。为了从这些预测的属性生成连贯和逼真的语音,我们采用了一个整流匹配解码器建立在一个Transformer架构,模型有效的概率路径从随机噪声的目标语音分布。大量的实验表明,V2SFlow的性能明显优于最先进的方法,甚至超过了地面真实话语的自然性。
摘要:In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) frameworkdesigned to generate natural and intelligible speech directly from silenttalking face videos. While recent V2S systems have shown promising results onconstrained datasets with limited speakers and vocabularies, their performanceoften degrades on real-world, unconstrained datasets due to the inherentvariability and complexity of speech signals. To address these challenges, wedecompose the speech signal into manageable subspaces (content, pitch, andspeaker information), each representing distinct speech attributes, and predictthem directly from the visual input. To generate coherent and realistic speechfrom these predicted attributes, we employ a rectified flow matching decoderbuilt on a Transformer architecture, which models efficient probabilisticpathways from random noise to the target speech distribution. Extensiveexperiments demonstrate that V2SFlow significantly outperforms state-of-the-artmethods, even surpassing the naturalness of ground truth utterances.
标题:音乐基金会模型的参数高效迁移学习
链接:https://arxiv.org/abs/2411.19371
备注:6+2 pages
摘要:最近发布了更多的音乐基础模型,承诺对音乐信息进行通用的、主要与任务无关的编码。使音乐基础模型适应下游任务的常见方法是探测和微调。然而,这些常见的迁移学习方法面临挑战。探测可能会导致次优性能,因为预训练的权重被冻结,而微调的计算成本很高,并且容易出现过拟合。我们的工作研究了使用参数有效的迁移学习(PETL)的音乐基础模型,它集成了探测和微调的优势。我们介绍了三种类型的PETL方法:基于适配器的方法,基于适配器的方法,和基于重新参数化的方法。这些方法仅训练少量参数,因此不需要大量的计算资源。结果表明,PETL方法优于探测和微调的音乐自动标记。在关键检测和节奏估计方面,它们实现了与微调类似的结果,但训练成本显著降低。然而,通过从头开始训练小模型所取得的类似结果,当前一代基础模型在关键和节奏任务上的有用性受到质疑。代码可在https://github.com/suncerock/peft-music/上获得
摘要:More music foundation models are recently being released, promising ageneral, mostly task independent encoding of musical information. Common waysof adapting music foundation models to downstream tasks are probing andfine-tuning. These common transfer learning approaches, however, facechallenges. Probing might lead to suboptimal performance because thepre-trained weights are frozen, while fine-tuning is computationally expensiveand is prone to overfitting. Our work investigates the use ofparameter-efficient transfer learning (PETL) for music foundation models whichintegrates the advantage of probing and fine-tuning. We introduce three typesof PETL methods: adapter-based methods, prompt-based methods, andreparameterization-based methods. These methods train only a small number ofparameters, and therefore do not require significant computational resources.Results show that PETL methods outperform both probing and fine-tuning on musicauto-tagging. On key detection and tempo estimation, they achieve similarresults as fine-tuning with significantly less training cost. However, theusefulness of the current generation of foundation model on key and tempo tasksis questioned by the similar results achieved by training a small model fromscratch. Code available at https://github.com/suncerock/peft-music/
标题:在家庭环境中使用对话虚拟助理进行基于语音的2型糖尿病分类
链接:https://arxiv.org/abs/2411.19204
备注:8 pages
摘要:随着机器学习和深度学习技术的出现,将云技术与医疗物联网结合起来,实现无处不在的医疗保健,在过去十年中已经出现了许多成功的应用。这些应用之一,即基于语音的病理学,尚未得到学术界和工业界的显着关注。将语音分析应用于致命疾病的早期检测,有望改善患者的健康状况和生活质量。在本文中,我们提出了一种基于声学机器学习的分类到商品化会话虚拟助理系统中的新应用,以预先筛选糖尿病的发作。具体来说,我们开发了一个分类系统,当他们与虚拟助手交谈时,从n=24名老年人的声音中提取声学特征,并预测糖尿病(2型)的发病率。我们的分流系统实现了命中率为70%和60%的男性和女性老年人受试者,分别。我们提出的分流使用7个不可识别的基于语音的功能,可以在资源受限的嵌入式系统运行基于语音的虚拟助理。该应用程序证明了应用基于语音的病理学分析的可行性,通过早期检测改变生活的慢性疾病(如糖尿病)来改善家庭环境中老年人的健康状况。
摘要:Incorporating cloud technology with Internet of Medical Things for ubiquitoushealthcare has seen many successful applications in the last decade with theadvent of machine learning and deep learning techniques. One of theseapplications, namely voice-based pathology, has yet to receive notableattention from academia and industry. Applying voice analysis to earlydetection of fatal diseases holds much promise to improve health outcomes andquality of life of patients. In this paper, we propose a novel application ofacoustic machine learning based triaging into commoditised conversationalvirtual assistant systems to pre-screen for onset of diabetes. Specifically, wedeveloped a triaging system which extracts acoustic features from the voices ofn=24 older adults when they converse with a virtual assistant and predict theincidence of Diabetes Mellitus (Type 2) or not. Our triaging system achievedhit-rates of 70% and 60% for male and female older adult subjects,respectively. Our proposed triaging uses 7 non-identifiable voice-basedfeatures and can operate within resource-constrained embedded systems runningvoice-based virtual assistants. This application demonstrates the feasibilityof applying voice-based pathology analysis to improve health outcomes of olderadults within the home environment by early detection of life-changing chronicconditions like diabetes.
标题:CoDiff-VC:用于零激发语音转换的编解码器辅助扩散模型
链接:https://arxiv.org/abs/2411.18918
备注:Submitted to ICASSP2025
摘要:Zero-shot语音转换(VC)的目的是将原说话人的音色转换为任何目标说话人,同时保持语言内容。当前主流的zero-shot语音转换方法依赖于预先训练的识别模型来解开语言内容和说话人表示。这导致解耦语言内容中的音色残留以及说话者表示建模的不足。在这项研究中,我们提出了CoDiff-VC,一个端到端的框架zero-shot语音转换,集成了语音编解码器和扩散模型,以产生高保真波形。我们的方法涉及采用单码本编解码器从源语音中分离语言内容。为了增强内容分解,我们引入混合风格层归一化(MSLN)来扰动原始音色。此外,我们采用了多尺度扬声器音色建模方法,以确保音色的一致性,提高语音细节的相似性。为了提高语音质量和说话人相似性,我们引入了双分类器免费指导,在生成过程中提供内容和音色指导。客观和主观实验证实,CoDiff-VC显着提高说话人相似度,生成自然和更高质量的语音。
摘要:Zero-shot voice conversion (VC) aims to convert the original speaker's timbreto any target speaker while keeping the linguistic content. Current mainstreamzero-shot voice conversion approaches depend on pre-trained recognition modelsto disentangle linguistic content and speaker representation. This results in atimbre residue within the decoupled linguistic content and inadequacies inspeaker representation modeling. In this study, we propose CoDiff-VC, anend-to-end framework for zero-shot voice conversion that integrates a speechcodec and a diffusion model to produce high-fidelity waveforms. Our approachinvolves employing a single-codebook codec to separate linguistic content fromthe source speech. To enhance content disentanglement, we introduce Mix-Stylelayer normalization (MSLN) to perturb the original timbre. Additionally, weincorporate a multi-scale speaker timbre modeling approach to ensure timbreconsistency and improve voice detail similarity. To improve speech quality andspeaker similarity, we introduce dual classifier-free guidance, providing bothcontent and timbre guidance during the generation process. Objective andsubjective experiments affirm that CoDiff-VC significantly improves speakersimilarity, generating natural and higher-quality speech.
标题:高斯语音:音频驱动的高斯化身
链接:https://arxiv.org/abs/2411.18675
备注:Paper Video: this https URL Project Page: this https URL
摘要:我们介绍GaussianSpeech,一种新的方法,合成高保真动画序列的照片般逼真的,个性化的3D人类头部化身从语音音频。为了捕捉人类头部的表现力,细节的性质,包括皮肤皱纹和精细尺度的面部运动,我们建议将语音信号与3D高斯飞溅耦合,以创建逼真的,时间相干的运动序列。我们提出了一个紧凑和高效的基于3DGS的化身表示,生成表情相关的颜色,并利用皱纹和感知为基础的损失,合成面部细节,包括皱纹,发生不同的表情。为了使序列建模的3D高斯splats与音频,我们设计了一个音频条件下的Transformer模型,能够直接从音频输入中提取嘴唇和表情特征。由于缺乏与音频相对应的高质量说话人数据集,我们捕获了一个新的大规模多视图说话人视听序列数据集,这些人具有母语英语口音和不同的面部几何形状。GaussianSpeech始终以实时渲染速率实现最先进的视觉自然运动性能,同时包含各种面部表情和风格。
摘要:We introduce GaussianSpeech, a novel approach that synthesizes high-fidelityanimation sequences of photo-realistic, personalized 3D human head avatars fromspoken audio. To capture the expressive, detailed nature of human heads,including skin furrowing and finer-scale facial movements, we propose to couplespeech signal with 3D Gaussian splatting to create realistic, temporallycoherent motion sequences. We propose a compact and efficient 3DGS-based avatarrepresentation that generates expression-dependent color and leverages wrinkle-and perceptually-based losses to synthesize facial details, including wrinklesthat occur with different expressions. To enable sequence modeling of 3DGaussian splats with audio, we devise an audio-conditioned transformer modelcapable of extracting lip and expression features directly from audio input.Due to the absence of high-quality datasets of talking humans in correspondencewith audio, we captured a new large-scale multi-view dataset of audio-visualsequences of talking humans with native English accents and diverse facialgeometry. GaussianSpeech consistently achieves state-of-the-art performancewith visually natural motion at real time rendering rates, while encompassingdiverse facial expressions and styles.
标题:迈向高级语音信号处理:基于卷积的架构及其应用的统计视角
链接:https://arxiv.org/abs/2411.18636
摘要:本文综述了基于卷积的模型,包括卷积神经网络(CNN),Conformer,ResNets和CRNN-作为语音信号处理模型,并提供了它们的统计背景和语音识别,说话人识别,情感识别和语音增强应用。通过比较训练成本评估、模型大小、准确性和速度评估,我们比较了每个模型的优缺点,识别了潜在的错误,并提出了进一步研究的途径,强调了它在推进语音技术应用中的核心作用。
摘要:This article surveys convolution-based models including convolutional neuralnetworks (CNNs), Conformers, ResNets, and CRNNs-as speech signal processingmodels and provide their statistical backgrounds and speech recognition,speaker identification, emotion recognition, and speech enhancementapplications. Through comparative training cost assessment, model size,accuracy and speed assessment, we compare the strengths and weaknesses of eachmodel, identify potential errors and propose avenues for further research,emphasizing the central role it plays in advancing applications of speechtechnologies.
标题:用于低比特率高质量语音编码的缩放变换器
链接:https://arxiv.org/abs/2411.19842
摘要:使用神经音频编解码器模型对语音进行标记化是现代人工智能管道的重要组成部分,用于单独或在多模态环境中生成或理解语音。传统上,这种标记化模型集中在低参数数架构上,仅使用具有强归纳偏差的组件。在这项工作中,我们表明,通过缩放一个Transformer架构与大参数计数这个问题,并应用灵活的有限标量量化(FSQ)为基础的瓶颈,它是可能达到国家的最先进的语音质量在极低的比特率为400美元或700美元比特每秒。经过训练的模型在客观和主观测试中都远远优于现有的基线。
摘要:The tokenization of speech with neural audio codec models is a vital part ofmodern AI pipelines for the generation or understanding of speech, alone or ina multimodal context. Traditionally such tokenization models have concentratedon low parameter-count architectures using only components with stronginductive biases. In this work we show that by scaling a transformerarchitecture with large parameter count to this problem, and applying aflexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible toreach state-of-the-art speech quality at extremely low bit-rates of $400$ or$700$ bits-per-second. The trained models strongly out-perform existingbaselines in both objective and subjective tests.
标题:用于低比特率高质量语音编码的缩放变换器
链接:https://arxiv.org/abs/2411.19842
摘要:使用神经音频编解码器模型对语音进行标记化是现代人工智能管道的重要组成部分,用于单独或在多模态环境中生成或理解语音。传统上,这种标记化模型集中在低参数数架构上,仅使用具有强归纳偏差的组件。在这项工作中,我们表明,通过缩放一个Transformer架构与大参数计数这个问题,并应用灵活的有限标量量化(FSQ)为基础的瓶颈,它是可能达到国家的最先进的语音质量在极低的比特率为400美元或700美元比特每秒。经过训练的模型在客观和主观测试中都远远优于现有的基线。
摘要:The tokenization of speech with neural audio codec models is a vital part ofmodern AI pipelines for the generation or understanding of speech, alone or ina multimodal context. Traditionally such tokenization models have concentratedon low parameter-count architectures using only components with stronginductive biases. In this work we show that by scaling a transformerarchitecture with large parameter count to this problem, and applying aflexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible toreach state-of-the-art speech quality at extremely low bit-rates of $400$ or$700$ bits-per-second. The trained models strongly out-perform existingbaselines in both objective and subjective tests.
标题:AudioSetCaps:使用自动生成管道和大型音频和语言模型的丰富音频字幕数据集
链接:https://arxiv.org/abs/2411.18953
摘要:随着音频语言模型的出现,构建大规模成对的音频语言数据集对于模型开发来说已经变得至关重要但具有挑战性,主要是由于所涉及的时间密集型和劳动密集型需求。虽然大型语言模型(LLM)已经提高了合成音频字幕生成的效率,但是当前的方法难以有效地提取和合并详细的音频信息。在本文中,我们提出了一个自动化的管道,集成了音频语言模型的细粒度内容提取,LLM合成字幕生成,和对比语言音频预训练(CLAP)模型为基础的细化过程,以提高字幕的质量。具体来说,我们在内容提取阶段采用提示链接技术来获得准确和细粒度的音频信息,同时我们使用细化过程来减轻生成的字幕中的潜在幻觉。利用AudioSet数据集和所提出的方法,我们创建了AudioSetCaps,这是一个包含190万个音频字幕对的数据集,是撰写本文时最大的音频字幕数据集。使用AudioSetCaps训练的模型在音频文本检索方面实现了最先进的性能,文本到音频的R@1得分为46.3%,音频到文本检索和自动音频字幕的R@1得分为59.7%,CIDER得分为84.8。由于我们的方法在AudioSetCaps上显示出有希望的结果,我们创建了另一个包含410万个基于Youtube-8 M和VGGSound数据集的合成音频语言对的数据集。为了促进音频语言学习的研究,我们在https://github.com/JishengBai/AudioSetCaps上公开了我们的管道,600万个音频语言对的数据集和预训练模型。
摘要:With the emergence of audio-language models, constructing large-scale pairedaudio-language datasets has become essential yet challenging for modeldevelopment, primarily due to the time-intensive and labour-heavy demandsinvolved. While large language models (LLMs) have improved the efficiency ofsynthetic audio caption generation, current approaches struggle to effectivelyextract and incorporate detailed audio information. In this paper, we proposean automated pipeline that integrates audio-language models for fine-grainedcontent extraction, LLMs for synthetic caption generation, and a contrastivelanguage-audio pretraining (CLAP) model-based refinement process to improve thequality of captions. Specifically, we employ prompt chaining techniques in thecontent extraction stage to obtain accurate and fine-grained audio information,while we use the refinement process to mitigate potential hallucinations in thegenerated captions. Leveraging the AudioSet dataset and the proposed approach,we create AudioSetCaps, a dataset comprising 1.9 million audio-caption pairs,the largest audio-caption dataset at the time of writing. The models trainedwith AudioSetCaps achieve state-of-the-art performance on audio-text retrievalwith R@1 scores of 46.3% for text-to-audio and 59.7% for audio-to-textretrieval and automated audio captioning with the CIDEr score of 84.8. As ourapproach has shown promising results with AudioSetCaps, we create anotherdataset containing 4.1 million synthetic audio-language pairs based on theYoutube-8M and VGGSound datasets. To facilitate research in audio-languagelearning, we have made our pipeline, datasets with 6 million audio-languagepairs, and pre-trained models publicly available athttps://github.com/JishengBai/AudioSetCaps.
标题:TS 3-编解码器:基于转换器的简单流媒体单一编解码器
链接:https://arxiv.org/abs/2411.18803
摘要:神经音频编解码器(NAC)作为音频压缩以及语音语言模型的音频表示的关键技术已经获得了极大的关注。虽然主流NAC模型主要是基于卷积的,但具有纯粹基于变压器和无卷积架构的NAC的性能仍然未被探索。介绍了一种基于转换器的简单流单路编解码器TS 3-Codec。TS 3-Codec仅由一堆Transformer层和几个线性层组成,通过完全消除需要仔细调整超参数和大量计算的卷积层,提供更大的简单性和表现力。在流式传输设置下,与具有最先进的基于卷积的架构的编解码器相比,所提出的TS 3编解码器实现了相当或更高的性能,同时仅需要12%的计算和77%的比特率。此外,当使用类似的计算资源时,它显著优于基于卷积的编解码器。
摘要:Neural audio codecs (NACs) have garnered significant attention as keytechnologies for audio compression as well as audio representation for speechlanguage models. While mainstream NAC models are predominantlyconvolution-based, the performance of NACs with a purely transformer-based, andconvolution-free architecture remains unexplored. This paper introducesTS3-Codec, a Transformer-Based Simple Streaming Single Codec. TS3-Codecconsists of only a stack of transformer layers with a few linear layers,offering greater simplicity and expressiveness by fully eliminating convolutionlayers that require careful hyperparameter tuning and large computations. Underthe streaming setup, the proposed TS3-Codec achieves comparable or superiorperformance compared to the codec with state-of-the-art convolution-basedarchitecture while requiring only 12% of the computation and 77% of bitrate.Furthermore, it significantly outperforms the convolution-based codec whenusing similar computational resources.
标题:基于音乐间隔的音乐创作和2D元胞自动机
链接:https://arxiv.org/abs/2411.19844
备注:17 pages, 3 figures
摘要:本研究是一种理论方法,探索的适用性的2D元胞自动机的基础上旋律和谐波间隔的随机阵列的音符。本研究的目的是探讨元胞自动机在音乐情境中的替代用途,以更好地理解音乐创造力。我们以复杂系统和人文方法为框架,以音乐理论的规律为基础,捕捉音乐创作的本质。研究结果表明,这些规则的问题产生大规模的模式有组织的笔记。因此,我们的配方提供了一种新的方法来理解和复制方面的音乐创造力。
摘要:This study is a theoretical approach for exploring the applicability of a 2Dcellular automaton based on melodic and harmonic intervals in random arrays ofmusical notes. The aim of this study was to explore alternatives uses for acellular automaton in the musical context for better understanding the musicalcreativity. We used the complex systems and humanities approaches as aframework for capturing the essence of creating music based on rules of musictheory. Findings suggested that such rules matter for generating large-scalepatterns of organized notes. Therefore, our formulation provides a novelapproach for understanding and replicating aspects of the musical creativity.
标题:并行堆叠聚合网络用于支持物联网的智能设备中的语音认证
链接:https://arxiv.org/abs/2411.19841
备注:arXiv admin note: text overlap with arXiv:2309.10560
摘要:近年来,由于对用户隐私和安全的担忧日益增加,物联网智能设备上的语音认证已经变得突出。当前的认证系统容易受到不同的语音欺骗攻击(例如,重放、语音克隆和音频深度伪造),其模仿合法语音以欺骗认证系统并实现欺诈活动(例如,假冒、未经授权的访问、金融欺诈等)。现有的解决方案通常被设计为应对单一类型的攻击,从而导致针对不可见攻击的性能受损。另一方面,现有的统一语音反欺骗解决方案并非专门针对物联网设计,具有复杂的架构,因此无法部署在支持物联网的智能设备上。此外,这些统一解决方案中的大多数都表现出严重的性能问题,包括较高的相等错误率或针对特定攻击的较低准确性。为了克服这些问题,我们提出了并行堆栈聚合网络(PSA-Net),这是一个轻量级框架,旨在为语音控制的智能物联网设备提供反欺骗防御系统。PSA-Net直接处理原始音频,无需依赖于音频的手工制作功能或预先计算的频谱图。此外,PSA-Net采用分裂-变换-聚合方法,该方法涉及话语的分割,通过卷积提取内在可微嵌入,以及将它们聚合以区分合法音频和欺骗音频。与现有的面向Resnet的深度解决方案相比,我们将基数作为网络中的一个额外维度,这增强了PSA-Net在不同攻击中的泛化能力。结果表明,PSA-Net对当前反欺骗解决方案中存在的不同攻击实现了更一致的性能。
摘要:Voice authentication on IoT-enabled smart devices has gained prominence inrecent years due to increasing concerns over user privacy and security. Thecurrent authentication systems are vulnerable to different voice-spoofingattacks (e.g., replay, voice cloning, and audio deepfakes) that mimiclegitimate voices to deceive authentication systems and enable fraudulentactivities (e.g., impersonation, unauthorized access, financial fraud, etc.).Existing solutions are often designed to tackle a single type of attack,leading to compromised performance against unseen attacks. On the other hand,existing unified voice anti-spoofing solutions, not designed specifically forIoT, possess complex architectures and thus cannot be deployed on IoT-enabledsmart devices. Additionally, most of these unified solutions exhibitsignificant performance issues, including higher equal error rates or loweraccuracy for specific attacks. To overcome these issues, we present theparallel stacked aggregation network (PSA-Net), a lightweight frameworkdesigned as an anti-spoofing defense system for voice-controlled smart IoTdevices. The PSA-Net processes raw audios directly and eliminates the need fordataset-dependent handcrafted features or pre-computed spectrograms.Furthermore, PSA-Net employs a split-transform-aggregate approach, whichinvolves the segmentation of utterances, the extraction of intrinsicdifferentiable embeddings through convolutions, and the aggregation of them todistinguish legitimate from spoofed audios. In contrast to existing deepResnet-oriented solutions, we incorporate cardinality as an additionaldimension in our network, which enhances the PSA-Net ability to generalizeacross diverse attacks. The results show that the PSA-Net achieves moreconsistent performance for different attacks that exist in currentanti-spoofing solutions.
标题:采用关节嵌入预测架构的Zero-Shot音乐主干检索
链接:https://arxiv.org/abs/2411.19806
备注:Submitted to ICASSP 2025
摘要:在本文中,我们解决的任务,音乐干检索。给定一个音乐组合,它包括检索一个与之匹配的词干,即,一起演奏的话会很好听为此,我们引入了一种基于联合嵌入预测架构的新方法,其中编码器和预测器被联合训练以产生上下文的潜在表示并预测目标的潜在表示。特别是,我们将预测器设计为以任意仪器为条件,使我们的模型能够执行零次(zero-shot)词干检索。此外,我们发现使用对比学习对编码器进行预训练大大提高了模型的性能。 我们使用MUSDB18和MoisesDB数据集验证了我们的模型的检索性能。我们表明,它在两个数据集上的表现都明显优于以前的基线,展示了它支持或多或少精确(可能是看不见的)条件的能力。我们还评估了学习嵌入的节拍跟踪任务,证明他们保留时间结构和本地信息。
摘要:In this paper, we tackle the task of musical stem retrieval. Given a musicalmix, it consists in retrieving a stem that would fit with it, i.e., that wouldsound pleasant if played together. To do so, we introduce a new method based onJoint-Embedding Predictive Architectures, where an encoder and a predictor arejointly trained to produce latent representations of a context and predictlatent representations of a target. In particular, we design our predictor tobe conditioned on arbitrary instruments, enabling our model to performzero-shot stem retrieval. In addition, we discover that pretraining the encoderusing contrastive learning drastically improves the model's performance. We validate the retrieval performances of our model using the MUSDB18 andMoisesDB datasets. We show that it significantly outperforms previous baselineson both datasets, showcasing its ability to support more or less precise (andpossibly unseen) conditioning. We also evaluate the learned embeddings on abeat tracking task, demonstrating that they retain temporal structure and localinformation.
标题:一种基于监督对比学习的跨数据库语音情感识别方法
链接:https://arxiv.org/abs/2411.19803
摘要:语音情感识别(SER)研究在处理不同分布的数据时经常面临缺乏大规模公共数据集和泛化能力有限等挑战。针对这一问题,提出一种基于监督对比学习的跨语料语音情感识别方法。该方法采用两阶段的微调过程:首先,使用多个语音情感数据集上的监督对比学习微调自监督语音表示模型;然后,在目标数据集上微调分类器。实验结果表明,基于WavLM的模型在IEMOCAP数据集和CASIA数据集上的未加权准确率(UA)分别为77.41%和96.49%,优于这两个数据集上的最新结果。
摘要:Research on Speech Emotion Recognition (SER) often faces challenges such asthe lack of large-scale public datasets and limited generalization capabilitywhen dealing with data from different distributions. To solve this problem,this paper proposes a cross-corpus speech emotion recognition method based onsupervised contrast learning. The method employs a two-stage fine-tuningprocess: first, the self-supervised speech representation model is fine-tunedusing supervised contrastive learning on multiple speech emotion datasets;then, the classifier is fine-tuned on the target dataset. The experimentalresults show that the WavLM-based model achieved unweighted accuracy (UA) of77.41% on the IEMOCAP dataset and 96.49% on the CASIA dataset, outperformingthe state-of-the-art results on the two datasets.
标题:电子竞技中的语音沟通分析
链接:https://arxiv.org/abs/2411.19793
备注:17 pages, 11 figures. Independent research
摘要:在大多数基于团队的电子竞技中,语音通信在团队效率和协同作用方面非常突出。事实上,已经观察到,当试图在正式比赛中有良好表现时,不仅团队的技能方面,而且团队有效的语音通信也会发挥作用。随着最近出现的LLM(大型语言模型)工具关于NLP(自然语言处理)(Vaswani et.等),我们决定尝试应用它们,以便更好地了解如何提高语音通信的有效性。本文以《英雄联盟》电子竞技为视角进行研究。然而,主要的概念和想法可以很容易地应用于任何其他团队相关的电子竞技。
摘要:In most team-based esports, voice communications are prominent in the teamefficiency and synergy. In fact it has been observed that not only the skillaspect of the team but also the team effective voice communication comes intoplay when trying to have good performance in official matches. With the recentemergence of LLM (Large Language Models) tools regarding NLP (Natural LanguageProcessing) (Vaswani et. al.), we decided to try applying them in order to havea better understanding on how to improve the effectiveness of the voicecommunications. In this paper the study has been made through the prism ofLeague of Legends esport. However the main concepts and ideas can be easilyapplicable in any other team related esports.
标题:Noro:具有隐藏说话者表示能力的噪音稳健单次语音转换系统
链接:https://arxiv.org/abs/2411.19770
备注:Submitted to IEEE OJSP
摘要:单次语音转换(One-shot voice conversion,VC)的目的是在保持原语音语义的前提下,仅使用目标语音中的一个参考语音,将源语音的音色转换为目标语音的音色。尽管一次性VC取得了进步,但在现实世界的场景中,其有效性会下降,其中参考语音通常来自互联网,包含各种干扰,如背景噪声。为了解决这个问题,我们引入了Noro,一种噪声鲁棒的单次VC系统。Noro具有专为VC使用噪声参考语音量身定制的创新组件,包括双分支参考编码模块和与噪声无关的对比扬声器损失。实验结果表明,Noro优于我们的基线系统在干净和嘈杂的情况下,突出了其在现实世界中的应用的功效。此外,我们通过将其参考编码器重新利用为扬声器编码器来研究基线系统的隐藏扬声器表示能力。实验结果表明,在SUPERB环境下,该方法与几种先进的自监督学习模型相比具有较强的竞争力,突出了通过一次性VC任务推进说话人表征学习的潜力。
摘要:One-shot voice conversion (VC) aims to alter the timbre of speech from asource speaker to match that of a target speaker using just a single referencespeech from the target, while preserving the semantic content of the originalsource speech. Despite advancements in one-shot VC, its effectiveness decreasesin real-world scenarios where reference speeches, often sourced from theinternet, contain various disturbances like background noise. To address thisissue, we introduce Noro, a Noise Robust One-shot VC system. Noro featuresinnovative components tailored for VC using noisy reference speeches, includinga dual-branch reference encoding module and a noise-agnostic contrastivespeaker loss. Experimental results demonstrate that Noro outperforms ourbaseline system in both clean and noisy scenarios, highlighting its efficacyfor real-world applications. Additionally, we investigate the hidden speakerrepresentation capabilities of our baseline system by repurposing its referenceencoder as a speaker encoder. The results shows that it is competitive withseveral advanced self-supervised learning models for speaker representationunder the SUPERB settings, highlighting the potential for advancing speakerrepresentation learning through one-shot VC task.
标题:用于节能音频分类的记忆纳米线网络:无需预处理、缩短延迟的水库计算
链接:https://arxiv.org/abs/2411.19611
备注:17 pages, 6 Figures
摘要:语音识别是自然语言处理中的一个关键挑战,要求实时应用具有低延迟、高效计算和强泛化能力。虽然基于软件的人工神经网络(ANN)擅长这项任务,但它们是计算密集型的,并且严重依赖于数据预处理。神经形态计算以其低延迟和节能的优势,有望用于音频分类。忆阻纳米线网络与梅尔频率倒谱系数提取等预处理技术相结合,已被广泛用于联想学习,但这种预处理可能是功耗密集型的,破坏了延迟优势。这项研究开创了使用纳米线网络的忆阻和时空特性进行音频信号分类,而无需预处理。纳米线网络仿真与三个线性分类器配对,用于10类MNIST音频分类和二进制扬声器泛化测试。该混合系统实现了显著的优势:出色的数据压缩,仅利用3%的纳米线输出,计算延迟减少10倍,分类准确性提高高达28.5%(使用逻辑回归分类器)。与原始数据分类器相比,多说话人数据集的准确率和召回率分别提高了10%和17%,单个说话人数据集的准确率和召回率分别提高了24%和17%。这项工作为在边缘计算设备中利用忆阻纳米线网络(NWN)提供了基本的概念证明,展示了它们在有效实时音频信号处理方面的潜力,同时降低了计算开销和功耗,并使先进的神经形态计算解决方案的发展成为可能。
摘要:Speech recognition is a key challenge in natural language processing,requiring low latency, efficient computation, and strong generalization forreal-time applications. While software-based artificial neural networks (ANNs)excel at this task, they are computationally intensive and depend heavily ondata pre-processing. Neuromorphic computing, with its low-latency andenergy-efficient advantages, holds promise for audio classification. Memristivenanowire networks, combined with pre-processing techniques like Mel-FrequencyCepstrum Coefficient extraction, have been widely used for associativelearning, but such pre-processing can be power-intensive, undermining latencybenefits. This study pioneers the use of memristive and spatio-temporalproperties of nanowire networks for audio signal classification withoutpre-processing. A nanowire network simulation is paired with three linearclassifiers for 10-class MNIST audio classification and binary speakergeneralization tests. The hybrid system achieves significant benefits:excellent data compression with only 3% of nanowire output utilized, a 10-foldreduction in computational latency, and up to 28.5% improved classificationaccuracy (using a logistic regression classifier). Precision and recall improveby 10% and 17% for multispeaker datasets, and by 24% and 17% for individualspeaker datasets, compared to raw data classifiers.This work provides afoundational proof of concept for utilizing memristive nanowire networks (NWN)in edge-computing devices, showcasing their potential for efficient, real-timeaudio signal processing with reduced computational overhead and powerconsumption, and enabling the development of advanced neuromorphic computingsolutions.
标题:生成人工智能时代的Deepfake媒体生成和检测:调查和展望
链接:https://arxiv.org/abs/2411.19537
摘要:随着生成建模的最新进展,deepfake内容的真实性一直在稳步增长,甚至达到了人们经常无法在线检测到被操纵的媒体内容的程度,从而被欺骗到各种骗局中。在本文中,我们综述了deepfake生成和检测技术,包括该领域的最新发展,如扩散模型和神经辐射场。我们的文献综述涵盖了所有deepfake媒体类型,包括图像、视频、音频和多模式(视听)内容。我们根据用于更改或生成虚假内容的程序来识别各种类型的deepfake。我们进一步构建了deepfake生成和检测方法的分类,说明了重要的方法组以及这些方法的应用领域。接下来,我们收集用于deepfake检测的数据集,并在最流行的数据集上提供最佳deepfake检测器的更新排名。此外,我们还开发了一种新的多模态基准测试,以评估deepfake检测器对分发内容的影响。结果表明,最先进的检测器无法推广到由看不见的深度伪造生成器生成的深度伪造内容。最后,我们提出了未来的方向,以获得强大而强大的深度伪造检测器。我们的项目页面和新的基准可以在https://github.com/CroitoruAlin/biodeep上找到。
摘要:With the recent advancements in generative modeling, the realism of deepfakecontent has been increasing at a steady pace, even reaching the point wherepeople often fail to detect manipulated media content online, thus beingdeceived into various kinds of scams. In this paper, we survey deepfakegeneration and detection techniques, including the most recent developments inthe field, such as diffusion models and Neural Radiance Fields. Our literaturereview covers all deepfake media types, comprising image, video, audio andmultimodal (audio-visual) content. We identify various kinds of deepfakes,according to the procedure used to alter or generate the fake content. Wefurther construct a taxonomy of deepfake generation and detection methods,illustrating the important groups of methods and the domains where thesemethods are applied. Next, we gather datasets used for deepfake detection andprovide updated rankings of the best performing deepfake detectors on the mostpopular datasets. In addition, we develop a novel multimodal benchmark toevaluate deepfake detectors on out-of-distribution content. The resultsindicate that state-of-the-art detectors fail to generalize to deepfake contentgenerated by unseen deepfake generators. Finally, we propose future directionsto obtain robust and powerful deepfake detectors. Our project page and newbenchmark are available at https://github.com/CroitoruAlin/biodeep.
标题:同图:可控实时说话头部合成的运动空间扩散
链接:https://arxiv.org/abs/2411.19509
摘要:扩散模型的最新进展彻底改变了音频驱动的说话头合成。除了精确的嘴唇同步,基于扩散的方法在生成与音频信号良好对齐的微妙表情和自然头部运动方面表现出色。然而,这些方法都面临着缓慢的推理速度,不足的细粒度控制面部运动,偶尔的视觉伪影,主要是由于隐式的潜在空间来自变分自动编码器(VAE),这阻止了他们在实时交互应用中的采用。为了解决这些问题,我们引入Ditto,一个基于扩散的框架,使可控的实时说话头合成。我们的关键创新在于通过一个明确的身份不可知的运动空间,取代传统的VAE表示,桥接运动生成和真实感神经渲染。这种设计大大降低了扩散学习的复杂性,同时能够精确控制合成的说话头。我们进一步提出了一个推理策略,共同优化三个关键组成部分:音频特征提取,运动生成和视频合成。这种优化实现了流处理、实时推理和低第一帧延迟,这些功能对于AI助手等交互式应用至关重要。大量的实验结果表明,Ditto生成引人注目的说话头部视频,并大大优于现有的方法在运动控制和实时性能。
摘要:Recent advances in diffusion models have revolutionized audio-driven talkinghead synthesis. Beyond precise lip synchronization, diffusion-based methodsexcel in generating subtle expressions and natural head movements that arewell-aligned with the audio signal. However, these methods are confronted byslow inference speed, insufficient fine-grained control over facial motions,and occasional visual artifacts largely due to an implicit latent space derivedfrom Variational Auto-Encoders (VAE), which prevent their adoption in realtimeinteraction applications. To address these issues, we introduce Ditto, adiffusion-based framework that enables controllable realtime talking headsynthesis. Our key innovation lies in bridging motion generation andphotorealistic neural rendering through an explicit identity-agnostic motionspace, replacing conventional VAE representations. This design substantiallyreduces the complexity of diffusion learning while enabling precise controlover the synthesized talking heads. We further propose an inference strategythat jointly optimizes three key components: audio feature extraction, motiongeneration, and video synthesis. This optimization enables streamingprocessing, realtime inference, and low first-frame delay, which are thefunctionalities crucial for interactive applications such as AI assistants.Extensive experimental results demonstrate that Ditto generates compellingtalking head videos and substantially outperforms existing methods in bothmotion control and realtime performance.
标题:V2 SFlow:具有语音分解和纠正流的视频转语音生成
链接:https://arxiv.org/abs/2411.19486
摘要:在本文中,我们介绍了V2SFlow,一种新颖的视频到语音(V2S)框架,旨在直接从无声的说话人脸视频生成自然和可理解的语音。虽然最近的V2S系统在具有有限扬声器和词汇的约束数据集上显示出有希望的结果,但由于语音信号的固有可变性和复杂性,它们的性能在现实世界的无约束数据集上通常会下降。为了解决这些挑战,我们将语音信号分解为可管理的子空间(内容,音高和说话人信息),每个子空间代表不同的语音属性,并直接从视觉输入中预测它们。为了从这些预测的属性生成连贯和逼真的语音,我们采用了一个整流匹配解码器建立在一个Transformer架构,模型有效的概率路径从随机噪声的目标语音分布。大量的实验表明,V2SFlow的性能明显优于最先进的方法,甚至超过了地面真实话语的自然性。
摘要:In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) frameworkdesigned to generate natural and intelligible speech directly from silenttalking face videos. While recent V2S systems have shown promising results onconstrained datasets with limited speakers and vocabularies, their performanceoften degrades on real-world, unconstrained datasets due to the inherentvariability and complexity of speech signals. To address these challenges, wedecompose the speech signal into manageable subspaces (content, pitch, andspeaker information), each representing distinct speech attributes, and predictthem directly from the visual input. To generate coherent and realistic speechfrom these predicted attributes, we employ a rectified flow matching decoderbuilt on a Transformer architecture, which models efficient probabilisticpathways from random noise to the target speech distribution. Extensiveexperiments demonstrate that V2SFlow significantly outperforms state-of-the-artmethods, even surpassing the naturalness of ground truth utterances.
标题:音乐基金会模型的参数高效迁移学习
链接:https://arxiv.org/abs/2411.19371
备注:6+2 pages
摘要:最近发布了更多的音乐基础模型,承诺对音乐信息进行通用的、主要与任务无关的编码。使音乐基础模型适应下游任务的常见方法是探测和微调。然而,这些常见的迁移学习方法面临挑战。探测可能导致次优性能,因为预先训练的权重被冻结,而微调在计算上是昂贵的,并且容易过拟合。我们的工作研究了使用参数有效的迁移学习(PETL)的音乐基础模型,它集成了探测和微调的优势。我们介绍了三种类型的PETL方法:基于适配器的方法,基于适配器的方法,和基于重新参数化的方法。这些方法仅训练少量参数,因此不需要大量的计算资源。结果表明,PETL方法优于探测和微调的音乐自动标记。在关键检测和节奏估计方面,它们实现了与微调类似的结果,但训练成本显著降低。然而,通过从头开始训练小模型所取得的类似结果,当前一代基础模型在关键和节奏任务上的有用性受到质疑。代码可在https://github.com/suncerock/peft-music/上获得
摘要:More music foundation models are recently being released, promising ageneral, mostly task independent encoding of musical information. Common waysof adapting music foundation models to downstream tasks are probing andfine-tuning. These common transfer learning approaches, however, facechallenges. Probing might lead to suboptimal performance because thepre-trained weights are frozen, while fine-tuning is computationally expensiveand is prone to overfitting. Our work investigates the use ofparameter-efficient transfer learning (PETL) for music foundation models whichintegrates the advantage of probing and fine-tuning. We introduce three typesof PETL methods: adapter-based methods, prompt-based methods, andreparameterization-based methods. These methods train only a small number ofparameters, and therefore do not require significant computational resources.Results show that PETL methods outperform both probing and fine-tuning on musicauto-tagging. On key detection and tempo estimation, they achieve similarresults as fine-tuning with significantly less training cost. However, theusefulness of the current generation of foundation model on key and tempo tasksis questioned by the similar results achieved by training a small model fromscratch. Code available at https://github.com/suncerock/peft-music/
标题:在家庭环境中使用对话虚拟助理进行基于语音的2型糖尿病分类
链接:https://arxiv.org/abs/2411.19204
备注:8 pages
摘要:在过去十年中,随着机器学习和深度学习技术的出现,将云技术与医疗物联网结合起来用于无处不在的医疗保健已经取得了许多成功的应用。这些应用之一,即基于语音的病理学,尚未得到学术界和工业界的显着关注。将语音分析应用于致命疾病的早期检测,有望改善患者的健康状况和生活质量。在本文中,我们提出了一种基于声学机器学习的分类到商品化会话虚拟助理系统中的新应用,以预先筛选糖尿病的发作。具体来说,我们开发了一个分类系统,当他们与虚拟助手交谈时,从n=24名老年人的声音中提取声学特征,并预测糖尿病(2型)的发病率。我们的分流系统实现了命中率为70%和60%的男性和女性老年人受试者,分别。我们提出的分流使用7个不可识别的基于语音的功能,可以在资源受限的嵌入式系统运行基于语音的虚拟助理。该应用程序证明了应用基于语音的病理学分析的可行性,通过早期检测改变生活的慢性疾病(如糖尿病)来改善家庭环境中老年人的健康状况。
摘要:Incorporating cloud technology with Internet of Medical Things for ubiquitoushealthcare has seen many successful applications in the last decade with theadvent of machine learning and deep learning techniques. One of theseapplications, namely voice-based pathology, has yet to receive notableattention from academia and industry. Applying voice analysis to earlydetection of fatal diseases holds much promise to improve health outcomes andquality of life of patients. In this paper, we propose a novel application ofacoustic machine learning based triaging into commoditised conversationalvirtual assistant systems to pre-screen for onset of diabetes. Specifically, wedeveloped a triaging system which extracts acoustic features from the voices ofn=24 older adults when they converse with a virtual assistant and predict theincidence of Diabetes Mellitus (Type 2) or not. Our triaging system achievedhit-rates of 70% and 60% for male and female older adult subjects,respectively. Our proposed triaging uses 7 non-identifiable voice-basedfeatures and can operate within resource-constrained embedded systems runningvoice-based virtual assistants. This application demonstrates the feasibilityof applying voice-based pathology analysis to improve health outcomes of olderadults within the home environment by early detection of life-changing chronicconditions like diabetes.
标题:CoDiff-VC:用于零激发语音转换的编解码器辅助扩散模型
链接:https://arxiv.org/abs/2411.18918
备注:Submitted to ICASSP2025
摘要:Zero-shot语音转换(VC)的目的是将原说话人的音色转换为任何目标说话人,同时保持语言内容。当前主流的zero-shot语音转换方法依赖于预先训练的识别模型来解开语言内容和说话人表示。这导致解耦语言内容中的音色残留以及说话者表示建模的不足。在这项研究中,我们提出了CoDiff-VC,一个端到端的框架zero-shot语音转换,集成了语音编解码器和扩散模型,以产生高保真波形。我们的方法涉及采用单码本编解码器从源语音中分离语言内容。为了增强内容分解,我们引入混合风格层归一化(MSLN)来扰动原始音色。此外,我们采用了多尺度扬声器音色建模方法,以确保音色的一致性,提高语音细节的相似性。为了提高语音质量和说话人相似性,我们引入了双分类器免费指导,在生成过程中提供内容和音色指导。客观和主观实验证实,CoDiff-VC显着提高说话人相似度,生成自然和更高质量的语音。
摘要:Zero-shot voice conversion (VC) aims to convert the original speaker's timbreto any target speaker while keeping the linguistic content. Current mainstreamzero-shot voice conversion approaches depend on pre-trained recognition modelsto disentangle linguistic content and speaker representation. This results in atimbre residue within the decoupled linguistic content and inadequacies inspeaker representation modeling. In this study, we propose CoDiff-VC, anend-to-end framework for zero-shot voice conversion that integrates a speechcodec and a diffusion model to produce high-fidelity waveforms. Our approachinvolves employing a single-codebook codec to separate linguistic content fromthe source speech. To enhance content disentanglement, we introduce Mix-Stylelayer normalization (MSLN) to perturb the original timbre. Additionally, weincorporate a multi-scale speaker timbre modeling approach to ensure timbreconsistency and improve voice detail similarity. To improve speech quality andspeaker similarity, we introduce dual classifier-free guidance, providing bothcontent and timbre guidance during the generation process. Objective andsubjective experiments affirm that CoDiff-VC significantly improves speakersimilarity, generating natural and higher-quality speech.
标题:高斯语音:音频驱动的高斯化身
链接:https://arxiv.org/abs/2411.18675
备注:Paper Video: this https URL Project Page: this https URL
摘要:我们介绍GaussianSpeech,一种新的方法,合成高保真动画序列的照片般逼真的,个性化的3D人类头部化身从语音音频。为了捕捉人类头部的表现力,细节的性质,包括皮肤皱纹和精细尺度的面部运动,我们建议将语音信号与3D高斯飞溅耦合,以创建逼真的,时间相干的运动序列。我们提出了一个紧凑和高效的基于3DGS的化身表示,生成表情相关的颜色,并利用皱纹和感知为基础的损失,合成面部细节,包括皱纹,发生不同的表情。为了使序列建模的3D高斯splats与音频,我们设计了一个音频条件下的Transformer模型,能够直接从音频输入中提取嘴唇和表情特征。由于缺乏与音频相对应的高质量说话人数据集,我们捕获了一个新的大规模多视图说话人视听序列数据集,这些人具有母语英语口音和不同的面部几何形状。GaussianSpeech始终以实时渲染速率实现最先进的视觉自然运动性能,同时包含各种面部表情和风格。
摘要:We introduce GaussianSpeech, a novel approach that synthesizes high-fidelityanimation sequences of photo-realistic, personalized 3D human head avatars fromspoken audio. To capture the expressive, detailed nature of human heads,including skin furrowing and finer-scale facial movements, we propose to couplespeech signal with 3D Gaussian splatting to create realistic, temporallycoherent motion sequences. We propose a compact and efficient 3DGS-based avatarrepresentation that generates expression-dependent color and leverages wrinkle-and perceptually-based losses to synthesize facial details, including wrinklesthat occur with different expressions. To enable sequence modeling of 3DGaussian splats with audio, we devise an audio-conditioned transformer modelcapable of extracting lip and expression features directly from audio input.Due to the absence of high-quality datasets of talking humans in correspondencewith audio, we captured a new large-scale multi-view dataset of audio-visualsequences of talking humans with native English accents and diverse facialgeometry. GaussianSpeech consistently achieves state-of-the-art performancewith visually natural motion at real time rendering rates, while encompassingdiverse facial expressions and styles.
标题:迈向高级语音信号处理:基于卷积的架构及其应用的统计视角
链接:https://arxiv.org/abs/2411.18636
摘要:本文综述了基于卷积的模型,包括卷积神经网络(CNN),Conformer,ResNets和CRNN-作为语音信号处理模型,并提供了它们的统计背景和语音识别,说话人识别,情感识别和语音增强应用。通过比较训练成本评估、模型大小、准确性和速度评估,我们比较了每个模型的优缺点,识别了潜在的错误,并提出了进一步研究的途径,强调了它在推进语音技术应用中的核心作用。
摘要:This article surveys convolution-based models including convolutional neuralnetworks (CNNs), Conformers, ResNets, and CRNNs-as speech signal processingmodels and provide their statistical backgrounds and speech recognition,speaker identification, emotion recognition, and speech enhancementapplications. Through comparative training cost assessment, model size,accuracy and speed assessment, we compare the strengths and weaknesses of eachmodel, identify potential errors and propose avenues for further research,emphasizing the central role it plays in advancing applications of speechtechnologies.
