【1】Enhancing Generalization in Audio Deepfake Detection: A Neural Collapse based Sampling and Training Approach标题:增强音频深度伪造检测的概括:基于神经崩溃的采样和训练方法链接:https://arxiv.org/abs/2404.13008作者:Mohammed Yousif,Jonat John Mathew,Huzaifa Pallan,Agamjeet Singh Padda,Syed Daniyal Shah,Sara Adamski,Madhu Reddiboina,Arjun Pankajakshan摘要:音频deepfake检测的泛化提出了一个重大挑战,在特定数据集上训练的模型通常难以检测在不同条件和未知算法下生成的deepfake。虽然使用不同的数据集集体训练模型可以增强其泛化能力,但它会带来很高的计算成本。为了解决这个问题,我们提出了一种基于神经网络的采样方法,应用于在不同数据集上训练的预训练模型,以创建一个新的训练数据库。使用ASVspoof 2019数据集作为概念验证,我们使用Resnet和ConvNext架构实现预训练模型。我们的方法在看不见的数据上表现出相当的泛化能力,同时计算效率高,需要更少的训练数据。使用野外数据集进行评价。摘要:Generalization in audio deepfake detection presents a significant challenge, with models trained on specific datasets often struggling to detect deepfakes generated under varying conditions and unknown algorithms. While collectively training a model using diverse datasets can enhance its generalization ability, it comes with high computational costs. To address this, we propose a neural collapse-based sampling approach applied to pre-trained models trained on distinct datasets to create a new training database. Using ASVspoof 2019 dataset as a proof-of-concept, we implement pre-trained models with Resnet and ConvNext architectures. Our approach demonstrates comparable generalization on unseen data while being computationally efficient, requiring less training data. Evaluation is conducted using the In-the-wild dataset. 【2】 TRNet: Two-level Refinement Network leveraging Speech Enhancement for Noise Robust Speech Emotion Recognition标题:TRNet:利用语音增强进行噪音稳健语音情感识别的两级细化网络链接:https://arxiv.org/abs/2404.12979作者:Chengxin Chen,Pengyuan Zhang备注:13 pages, 3 figures摘要:语音情感识别(SER)的一个持续挑战是无处不在的环境噪声,这经常导致在实际应用中降低SER性能。在本文中,我们介绍了一个两级细化网络,称为TRNet,以解决这一挑战。具体地,采用预训练的语音增强模块用于前端降噪和噪声水平估计。之后,我们利用干净的语音频谱图及其相应的深度表示作为参考信号,以在模型训练期间改进增强语音的频谱图失真和表示偏移。实验结果验证了所提出的TRNet在匹配和不匹配的噪声环境中大大提高了系统的鲁棒性,而不会影响其在清洁环境中的性能。摘要:One persistent challenge in Speech Emotion Recognition (SER) is the ubiquitous environmental noise, which frequently results in diminished SER performance in practical use. In this paper, we introduce a Two-level Refinement Network, dubbed TRNet, to address this challenge. Specifically, a pre-trained speech enhancement module is employed for front-end noise reduction and noise level estimation. Later, we utilize clean speech spectrograms and their corresponding deep representations as reference signals to refine the spectrogram distortion and representation shift of enhanced speech during model training. Experimental results validate that the proposed TRNet substantially increases the system's robustness in both matched and unmatched noisy environments, without compromising its performance in clean environments. 【3】 Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction标题:语音链中的分离:跨模式条件视听目标语音提取链接:https://arxiv.org/abs/2404.12725作者:Zhaoxi Mu,Xinyu Yang备注:Accepted by IJCAI 2024摘要:视觉线索的整合使目标语音提取任务的性能恢复了活力,将其提升到了该领域的最前沿。然而,这种多模态学习范式经常遇到模态失衡的挑战。在视听目标语音提取任务中,音频模态往往占主导地位,潜在地掩盖了视觉引导的重要性。为了解决这个问题,我们提出了AVSepChain,从语音链概念中汲取灵感。我们的方法划分的视听目标语音提取任务分为两个阶段:语音感知和语音生产。在言语感知阶段,听觉信息是主导模态,视觉信息是条件模态。相反,在言语产生阶段,角色是颠倒的。这种情态地位的转换旨在缓解情态失衡的问题。此外,我们引入了一个对比语义匹配损失,以确保所产生的语音传达的语义信息与在语音生产阶段的嘴唇运动传达的语义信息对齐。通过在多个视听目标语音提取基准数据集上进行的大量实验,我们展示了我们所提出的方法所取得的优异性能。摘要:The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the challenge of modality imbalance. In audio-visual target speech extraction tasks, the audio modality tends to dominate, potentially overshadowing the importance of visual guidance. To tackle this issue, we propose AVSepChain, drawing inspiration from the speech chain concept. Our approach partitions the audio-visual target speech extraction task into two stages: speech perception and speech production. In the speech perception stage, audio serves as the dominant modality, while visual information acts as the conditional modality. Conversely, in the speech production stage, the roles are reversed. This transformation of modality status aims to alleviate the problem of modality imbalance. Additionally, we introduce a contrastive semantic matching loss to ensure that the semantic information conveyed by the generated speech aligns with the semantic information conveyed by lip movements during the speech production stage. Through extensive experiments conducted on multiple benchmark datasets for audio-visual target speech extraction, we showcase the superior performance achieved by our proposed method.
eess.AS音频处理【1】 Enhancing Generalization in Audio Deepfake Detection: A Neural Collapse based Sampling and Training Approach标题:增强音频深度伪造检测的概括:基于神经崩溃的采样和训练方法链接:https://arxiv.org/abs/2404.13008作者:Mohammed Yousif,Jonat John Mathew,Huzaifa Pallan,Agamjeet Singh Padda,Syed Daniyal Shah,Sara Adamski,Madhu Reddiboina,Arjun Pankajakshan摘要:音频deepfake检测的泛化提出了一个重大挑战,在特定数据集上训练的模型通常难以检测在不同条件和未知算法下生成的deepfake。虽然使用不同的数据集集体训练模型可以增强其泛化能力,但它会带来很高的计算成本。为了解决这个问题,我们提出了一种基于神经网络的采样方法,应用于在不同数据集上训练的预训练模型,以创建一个新的训练数据库。使用ASVspoof 2019数据集作为概念验证,我们使用Resnet和ConvNext架构实现预训练模型。我们的方法在看不见的数据上表现出相当的泛化能力,同时计算效率高,需要更少的训练数据。使用野外数据集进行评价。摘要:Generalization in audio deepfake detection presents a significant challenge, with models trained on specific datasets often struggling to detect deepfakes generated under varying conditions and unknown algorithms. While collectively training a model using diverse datasets can enhance its generalization ability, it comes with high computational costs. To address this, we propose a neural collapse-based sampling approach applied to pre-trained models trained on distinct datasets to create a new training database. Using ASVspoof 2019 dataset as a proof-of-concept, we implement pre-trained models with Resnet and ConvNext architectures. Our approach demonstrates comparable generalization on unseen data while being computationally efficient, requiring less training data. Evaluation is conducted using the In-the-wild dataset.
【2】 TRNet: Two-level Refinement Network leveraging Speech Enhancement for Noise Robust Speech Emotion Recognition标题:TRNet:利用语音增强进行噪音稳健语音情感识别的两级细化网络链接:https://arxiv.org/abs/2404.12979作者:Chengxin Chen,Pengyuan Zhang备注:13 pages, 3 figures摘要:语音情感识别(SER)的一个持续挑战是无处不在的环境噪声,这经常导致在实际应用中降低SER性能。在本文中,我们介绍了一个两级细化网络,称为TRNet,以解决这一挑战。具体地,采用预训练的语音增强模块用于前端降噪和噪声水平估计。之后,我们利用干净的语音频谱图及其相应的深度表示作为参考信号,以在模型训练期间改进增强语音的频谱图失真和表示偏移。实验结果验证了所提出的TRNet在匹配和不匹配的噪声环境中大大提高了系统的鲁棒性,而不会影响其在清洁环境中的性能。摘要:One persistent challenge in Speech Emotion Recognition (SER) is the ubiquitous environmental noise, which frequently results in diminished SER performance in practical use. In this paper, we introduce a Two-level Refinement Network, dubbed TRNet, to address this challenge. Specifically, a pre-trained speech enhancement module is employed for front-end noise reduction and noise level estimation. Later, we utilize clean speech spectrograms and their corresponding deep representations as reference signals to refine the spectrogram distortion and representation shift of enhanced speech during model training. Experimental results validate that the proposed TRNet substantially increases the system's robustness in both matched and unmatched noisy environments, without compromising its performance in clean environments. 【3】 Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction标题:语音链中的分离:跨模式条件视听目标语音提取链接:https://arxiv.org/abs/2404.12725作者:Zhaoxi Mu,Xinyu Yang备注:Accepted by IJCAI 2024摘要:视觉线索的整合使目标语音提取任务的性能恢复了活力,将其提升到了该领域的最前沿。然而,这种多模态学习范式经常遇到模态失衡的挑战。在视听目标语音提取任务中,音频模态往往占主导地位,潜在地掩盖了视觉引导的重要性。为了解决这个问题,我们提出了AVSepChain,从语音链概念中汲取灵感。我们的方法划分的视听目标语音提取任务分为两个阶段:语音感知和语音生产。在言语感知阶段,听觉信息是主导模态,视觉信息是条件模态。相反,在言语产生阶段,角色是颠倒的。这种情态地位的转换旨在缓解情态失衡的问题。此外,我们引入了一个对比语义匹配损失,以确保所产生的语音传达的语义信息与在语音生产阶段的嘴唇运动传达的语义信息对齐。通过在多个视听目标语音提取基准数据集上进行的大量实验,我们展示了我们所提出的方法所取得的优异性能。摘要:The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the challenge of modality imbalance. In audio-visual target speech extraction tasks, the audio modality tends to dominate, potentially overshadowing the importance of visual guidance. To tackle this issue, we propose AVSepChain, drawing inspiration from the speech chain concept. Our approach partitions the audio-visual target speech extraction task into two stages: speech perception and speech production. In the speech perception stage, audio serves as the dominant modality, while visual information acts as the conditional modality. Conversely, in the speech production stage, the roles are reversed. This transformation of modality status aims to alleviate the problem of modality imbalance. Additionally, we introduce a contrastive semantic matching loss to ensure that the semantic information conveyed by the generated speech aligns with the semantic information conveyed by lip movements during the speech production stage. Through extensive experiments conducted on multiple benchmark datasets for audio-visual target speech extraction, we showcase the superior performance achieved by our proposed method. 机器翻译由腾讯交互翻译提供,仅供参考