今日论文合集:cs.SD语音4篇,eess.AS音频处理4篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】CLAD: Robust Audio Deepfake Detection Against Manipulation Attacks with Contrastive Learning

标题:CLAD:通过对比学习对抗操纵攻击的稳健音频Deepfake检测
链接:http://arxiv.org/pdf/2404.15854v1
作者:Haolin Wu,Jing Chen,Ruiying Du,Cong Wu,Kun He,Xingcan Shang,Hao Ren,Guowen Xu
备注:Submitted to IEEE TDSC
摘要:音频deepfake的日益流行构成了重大的安全威胁,需要强大的检测方法。虽然现有的检测系统表现出的承诺,其对恶意音频操纵的鲁棒性仍然未充分发掘。为了弥补这一差距,我们首次全面研究了最广泛采用的音频深度伪造检测器对操纵攻击的敏感性。令人惊讶的是,即使是像音量控制这样的操作也可以在不影响人类感知的情况下显著绕过检测。为了解决这个问题,我们提出了CLAD(基于对比学习的音频deepfake检测器)来增强对操纵攻击的鲁棒性。其关键思想是结合对比学习,以最大限度地减少由操作引入的变化,从而提高检测的鲁棒性。此外,我们还引入了长度损失,旨在通过在特征空间中更紧密地聚类真实音频来提高检测精度。我们全面评估了最广泛采用的音频deepfake检测模型和我们提出的CLAD对各种操纵攻击的影响。检测模型存在漏洞,在音量控制、衰落和噪声注入下,FAR分别上升到36.69%、31.23%和51.28%。CLAD增强了鲁棒性,在噪声注入下将FAR降低到0.81%,并在所有测试中始终保持FAR低于1.63%。我们的源代码和文档可以在工件存储库(https:github.com CLAD23 CLAD)中找到。
摘要:The increasing prevalence of audio deepfakes poses significant security threats, necessitating robust detection methods. While existing detection systems exhibit promise, their robustness against malicious audio manipulations remains underexplored. To bridge the gap, we undertake the first comprehensive study of the susceptibility of the most widely adopted audio deepfake detectors to manipulation attacks. Surprisingly, even manipulations like volume control can significantly bypass detection without affecting human perception. To address this, we propose CLAD (Contrastive Learning-based Audio deepfake Detector) to enhance the robustness against manipulation attacks. The key idea is to incorporate contrastive learning to minimize the variations introduced by manipulations, therefore enhancing detection robustness. Additionally, we incorporate a length loss, aiming to improve the detection accuracy by clustering real audios more closely in the feature space. We comprehensively evaluated the most widely adopted audio deepfake detection models and our proposed CLAD against various manipulation attacks. The detection models exhibited vulnerabilities, with FAR rising to 36.69%, 31.23%, and 51.28% under volume control, fading, and noise injection, respectively. CLAD enhanced robustness, reducing the FAR to 0.81% under noise injection and consistently maintaining an FAR below 1.63% across all tests. Our source code and documentation are available in the artifact repository (https: github.com CLAD23 CLAD).


【2】Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
标题:采用对抗互补表示学习的高效多模型融合
链接:http://arxiv.org/pdf/2404.15704v1
作者:Zuheng Kang,Yayun He,Jianzong Wang,Junqing Peng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:单模型系统在诸如说话人确认(SV)和图像分类等任务中经常存在缺陷,在决策过程中严重依赖部分先验知识,从而导致次优性能。虽然多模型融合(MMF)可以缓解这些问题中的一些,但学习表示中的冗余可能会限制改进。为此,我们提出了一个对抗性互补表示学习(ACoRL)框架,使新训练的模型能够避免以前获得的知识,允许每个单独的组件模型学习最大限度地不同的互补表示。我们做了三个详细的解释,为什么这个作品和实验结果表明,我们的方法更有效地提高性能相比,传统的MMF。此外,归因分析验证了在ACoRL下训练的模型获得了更多的互补知识,突出了我们的方法在提高跨任务的效率和鲁棒性方面的功效。
摘要:Single-model systems often suffer from deficiencies in tasks such as speaker verification (SV) and image classification, relying heavily on partial prior knowledge during decision-making, resulting in suboptimal performance. Although multi-model fusion (MMF) can mitigate some of these issues, redundancy in learned representations may limits improvements. To this end, we propose an adversarial complementary representation learning (ACoRL) framework that enables newly trained models to avoid previously acquired knowledge, allowing each individual component model to learn maximally distinct, complementary representations. We make three detailed explanations of why this works and experimental results demonstrate that our method more efficiently improves performance compared to traditional MMF. Furthermore, attribution analysis validates the model trained under ACoRL acquires more complementary knowledge, highlighting the efficacy of our approach in enhancing efficiency and robustness across tasks.

【3】HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
标题:HybridVC:通过文本和音频脚本进行高效的语音风格转换
链接:http://arxiv.org/pdf/2404.15637v1
作者:Xinlei Niu,Jing Zhang,Charles Patrick Martin
摘要:我们介绍HybridVC,这是一个基于预训练条件变分自编码器(CVAE)的语音转换(VC)框架,它结合了潜在模型与对比学习的优势。HybridVC支持文本和音频提示,实现更灵活的语音风格转换。HybridVC对以预先训练的说话人编码器获取的说话人嵌入为条件的潜在分布进行建模,并通过并行对比学习优化风格文本嵌入以与说话人风格信息保持一致。因此,HybridVC可以在有限的计算资源下有效地训练。我们的实验证明了HybridVC的优越的训练效率和先进的多模态语音风格转换的能力。这凸显了其广泛应用的潜力,例如各种社交媒体平台中的用户定义个性化语音。一个全面的消融研究进一步验证了我们的方法的有效性。
摘要:We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts, enabling more flexible voice style conversion. HybridVC models a latent distribution conditioned on speaker embeddings acquired by a pretrained speaker encoder and optimises style text embeddings to align with the speaker style information through contrastive learning in parallel. Therefore, HybridVC can be efficiently trained under limited computational resources. Our experiments demonstrate HybridVC's superior training efficiency and its capability for advanced multi-modal voice style conversion. This underscores its potential for widespread applications such as user-defined personalised voice in various social media platforms. A comprehensive ablation study further validates the effectiveness of our method.

【4】Characteristics-Based Design of Multi-Exponent Bandpass Filters
标题:基于特性的多指数带通过滤器设计
链接:http://arxiv.org/pdf/2404.15321v1
作者:Samiya A Alkhairy
备注:14 pages, 5 figures, 2 tables, 62 equations. Submitted to IEEE Transactions on Circuits and Systems I: Regular Papers
摘要:我们开发的方法来设计所需的特性,如峰值频率,带宽和群延迟的带通滤波器。我们开发了这种滤波器的设计方法,我们称为广义听觉滤波器(GAF),它被表示为二阶滤波器提出的非酉指数,因此有三个自由度的过滤器。我们的滤波器设计方法可满足三个频域特性的规范,包括峰值频率、凸度、3dB、ndB品质因数、等效矩形带宽、最大群延迟和相位累积。为了发展我们的基于特性的设计方法,我们推导出的滤波器常数直接在滤波器特性的表达式。GAF在特性集合方面的参数化允许同时指定基于幅度的特性(例如带宽)和基于相位的特性(例如群延迟)。这使得能够设计没有显著群延迟的急剧调谐滤波器,并且在频率选择性和同步都是设计的重要方面的滤波器组中特别重要。使用我们的方法,我们直接规定所需的滤波器特性值-不同于迭代滤波器设计方法。这允许更直接地设计GAF,用于从地震信号、耳蜗植入物和彩虹传感器中拾取相位。该方法也直接适用于相关的带通和多频带滤波器。
摘要:We develop methods to design bandpass filters given desired characteristics such as peak frequency, bandwidth, and group delay. We develop this filter design method for filters we refer to as Generalized Auditory Filters (GAFs) which are represented as second order filters raised to non-unitary exponents and hence have three degrees of freedom. Our method for filter design accommodates specification of a trio of frequency-domain characteristics from amongst the peak frequency, convexity, 3dB, ndB quality factor, equivalent rectangular bandwidth, maximum group delay, and phase accumulation. To develop our characteristics-based design methods, we derive expressions for the filter constants directly in terms of filter characteristics. The parameterization of GAFs in terms of sets of characteristics allows for specifying magnitude-based characteristics (e.g. bandwidths) and phase-based characteristics (e.g. group delays) simultaneously. This enables designing sharply tuned filters without significant group delay, and is particularly important in filterbanks where frequency selectivity and synchronization are both important aspects of design. Using our methods, we directly dictate values for desired filter characteristics - unlike iterative filter design methods. This allows for more direct design of GAFs for phase-picking from seismic signals, cochlear implants, and rainbow sensors. The methods also directly apply to related bandpass and multi-band filters.

eess.AS音频处理
【1】Characteristics-Based Design of Multi-Exponent Bandpass Filters
标题:基于特性的多指数带通过滤器设计
链接:http://arxiv.org/pdf/2404.15321v1
作者:Samiya A Alkhairy
备注:14 pages, 5 figures, 2 tables, 62 equations. Submitted to IEEE Transactions on Circuits and Systems I: Regular Papers
摘要:我们开发的方法来设计所需的特性,如峰值频率,带宽和群延迟的带通滤波器。我们开发了这种滤波器的设计方法,我们称为广义听觉滤波器(GAF),它被表示为二阶滤波器提出的非酉指数,因此有三个自由度的过滤器。我们的滤波器设计方法可满足三个频域特性的规范,包括峰值频率、凸度、3dB、ndB品质因数、等效矩形带宽、最大群延迟和相位累积。为了发展我们的基于特性的设计方法,我们推导出的滤波器常数直接在滤波器特性的表达式。GAF在特性集合方面的参数化允许同时指定基于幅度的特性(例如带宽)和基于相位的特性(例如群延迟)。这使得能够设计没有显著群延迟的急剧调谐滤波器,并且在频率选择性和同步都是设计的重要方面的滤波器组中特别重要。使用我们的方法,我们直接规定所需的滤波器特性值-不同于迭代滤波器设计方法。这允许更直接地设计GAF,用于从地震信号、耳蜗植入物和彩虹传感器中拾取相位。该方法也直接适用于相关的带通和多频带滤波器。
摘要:We develop methods to design bandpass filters given desired characteristics such as peak frequency, bandwidth, and group delay. We develop this filter design method for filters we refer to as Generalized Auditory Filters (GAFs) which are represented as second order filters raised to non-unitary exponents and hence have three degrees of freedom. Our method for filter design accommodates specification of a trio of frequency-domain characteristics from amongst the peak frequency, convexity, 3dB, ndB quality factor, equivalent rectangular bandwidth, maximum group delay, and phase accumulation. To develop our characteristics-based design methods, we derive expressions for the filter constants directly in terms of filter characteristics. The parameterization of GAFs in terms of sets of characteristics allows for specifying magnitude-based characteristics (e.g. bandwidths) and phase-based characteristics (e.g. group delays) simultaneously. This enables designing sharply tuned filters without significant group delay, and is particularly important in filterbanks where frequency selectivity and synchronization are both important aspects of design. Using our methods, we directly dictate values for desired filter characteristics - unlike iterative filter design methods. This allows for more direct design of GAFs for phase-picking from seismic signals, cochlear implants, and rainbow sensors. The methods also directly apply to related bandpass and multi-band filters.

【2】CLAD: Robust Audio Deepfake Detection Against Manipulation Attacks with Contrastive Learning

标题:CLAD:通过对比学习对抗操纵攻击的稳健音频Deepfake检测
链接:http://arxiv.org/pdf/2404.15854v1
作者:Haolin Wu,Jing Chen,Ruiying Du,Cong Wu,Kun He,Xingcan Shang,Hao Ren,Guowen Xu
备注:Submitted to IEEE TDSC
摘要:音频deepfake的日益流行构成了重大的安全威胁,需要强大的检测方法。虽然现有的检测系统表现出的承诺,其对恶意音频操纵的鲁棒性仍然未充分发掘。为了弥补这一差距,我们首次全面研究了最广泛采用的音频深度伪造检测器对操纵攻击的敏感性。令人惊讶的是,即使是像音量控制这样的操作也可以在不影响人类感知的情况下显著绕过检测。为了解决这个问题,我们提出了CLAD(基于对比学习的音频deepfake检测器)来增强对操纵攻击的鲁棒性。其关键思想是结合对比学习,以最大限度地减少由操作引入的变化,从而提高检测的鲁棒性。此外,我们还引入了长度损失,旨在通过在特征空间中更紧密地聚类真实音频来提高检测精度。我们全面评估了最广泛采用的音频deepfake检测模型和我们提出的CLAD对各种操纵攻击的影响。检测模型存在漏洞,在音量控制、衰落和噪声注入下,FAR分别上升到36.69%、31.23%和51.28%。CLAD增强了鲁棒性,在噪声注入下将FAR降低到0.81%,并在所有测试中始终保持FAR低于1.63%。我们的源代码和文档可以在工件存储库(https:github.com CLAD23 CLAD)中找到。
摘要:The increasing prevalence of audio deepfakes poses significant security threats, necessitating robust detection methods. While existing detection systems exhibit promise, their robustness against malicious audio manipulations remains underexplored. To bridge the gap, we undertake the first comprehensive study of the susceptibility of the most widely adopted audio deepfake detectors to manipulation attacks. Surprisingly, even manipulations like volume control can significantly bypass detection without affecting human perception. To address this, we propose CLAD (Contrastive Learning-based Audio deepfake Detector) to enhance the robustness against manipulation attacks. The key idea is to incorporate contrastive learning to minimize the variations introduced by manipulations, therefore enhancing detection robustness. Additionally, we incorporate a length loss, aiming to improve the detection accuracy by clustering real audios more closely in the feature space. We comprehensively evaluated the most widely adopted audio deepfake detection models and our proposed CLAD against various manipulation attacks. The detection models exhibited vulnerabilities, with FAR rising to 36.69%, 31.23%, and 51.28% under volume control, fading, and noise injection, respectively. CLAD enhanced robustness, reducing the FAR to 0.81% under noise injection and consistently maintaining an FAR below 1.63% across all tests. Our source code and documentation are available in the artifact repository (https: github.com CLAD23 CLAD).


【3】Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
标题:采用对抗互补表示学习的高效多模型融合
链接:http://arxiv.org/pdf/2404.15704v1
作者:Zuheng Kang,Yayun He,Jianzong Wang,Junqing Peng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:单模型系统在诸如说话人确认(SV)和图像分类等任务中经常存在缺陷,在决策过程中严重依赖部分先验知识,从而导致次优性能。虽然多模型融合(MMF)可以缓解这些问题中的一些,但学习表示中的冗余可能会限制改进。为此,我们提出了一个对抗性互补表示学习(ACoRL)框架,使新训练的模型能够避免以前获得的知识,允许每个单独的组件模型学习最大限度地不同的互补表示。我们做了三个详细的解释,为什么这个作品和实验结果表明,我们的方法更有效地提高性能相比,传统的MMF。此外,归因分析验证了在ACoRL下训练的模型获得了更多的互补知识,突出了我们的方法在提高跨任务的效率和鲁棒性方面的功效。
摘要:Single-model systems often suffer from deficiencies in tasks such as speaker verification (SV) and image classification, relying heavily on partial prior knowledge during decision-making, resulting in suboptimal performance. Although multi-model fusion (MMF) can mitigate some of these issues, redundancy in learned representations may limits improvements. To this end, we propose an adversarial complementary representation learning (ACoRL) framework that enables newly trained models to avoid previously acquired knowledge, allowing each individual component model to learn maximally distinct, complementary representations. We make three detailed explanations of why this works and experimental results demonstrate that our method more efficiently improves performance compared to traditional MMF. Furthermore, attribution analysis validates the model trained under ACoRL acquires more complementary knowledge, highlighting the efficacy of our approach in enhancing efficiency and robustness across tasks.

【4】HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
标题:HybridVC:通过文本和音频脚本进行高效的语音风格转换
链接:http://arxiv.org/pdf/2404.15637v1
作者:Xinlei Niu,Jing Zhang,Charles Patrick Martin
摘要:我们介绍HybridVC,这是一个基于预训练条件变分自编码器(CVAE)的语音转换(VC)框架,它结合了潜在模型与对比学习的优势。HybridVC支持文本和音频提示,实现更灵活的语音风格转换。HybridVC对以预先训练的说话人编码器获取的说话人嵌入为条件的潜在分布进行建模,并通过并行对比学习优化风格文本嵌入以与说话人风格信息保持一致。因此,HybridVC可以在有限的计算资源下有效地训练。我们的实验证明了HybridVC的优越的训练效率和先进的多模态语音风格转换的能力。这凸显了其广泛应用的潜力,例如各种社交媒体平台中的用户定义个性化语音。一个全面的消融研究进一步验证了我们的方法的有效性。
摘要:We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts, enabling more flexible voice style conversion. HybridVC models a latent distribution conditioned on speaker embeddings acquired by a pretrained speaker encoder and optimises style text embeddings to align with the speaker style information through contrastive learning in parallel. Therefore, HybridVC can be efficiently trained under limited computational resources. Our experiments demonstrate HybridVC's superior training efficiency and its capability for advanced multi-modal voice style conversion. This underscores its potential for widespread applications such as user-defined personalised voice in various social media platforms. A comprehensive ablation study further validates the effectiveness of our method.

机器翻译由腾讯交互翻译提供,仅供参考