本文经arXiv每日学术速递授权转载
【1】 ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
标题: ElasticAST:适用于所有长度和分辨率的音频频谱图Transformer
作者:Jiu Feng,Mehmet Hamza Erol,Joon Son Chung,Arda Senocak
备注:Interspeech 2024. Code is available at this https URL
链接:点击下载PDF文件
【2】 Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
标题: 评估无人机控制的语音命令管道:从STT和LLM到直接分类和连体网络
作者:Lucca Emmanuel Pineli Simões,Lucas Brandão Rodrigues,Rafaela Mota Silva,Gustavo Rodrigues da Silva
链接:点击下载PDF文件
【3】 Speech dereverberation constrained on room impulse response characteristics
标题: 语音去回响受限于房间脉冲响应特性
作者:Louis Bahrman,Mathieu Fontaine,Jonathan Le Roux,Gaël Richard
Journal-ref:INTERSPEECH, Sep 2024, Kos Island, Greece
链接:点击下载PDF文件
【4】 From Real to Cloned Singer Identification
标题: 从真实歌手到克隆歌手识别
作者:Dorian Desblancs,Gabriel Meseguer-Brocal,Romain Hennequin,Manuel Moussallam
备注:To be published at ISMIR 2024
链接:点击下载PDF文件
【5】 Autoregressive Speech Synthesis without Vector Quantization
标题: 无需量化的自回归语音合成
作者:Lingwei Meng,Long Zhou,Shujie Liu,Sanyuan Chen,Bing Han,Shujie Hu,Yanqing Liu,Jinyu Li,Sheng Zhao,Xixin Wu,Helen Meng,Furu Wei
链接:点击下载PDF文件
【6】 Adversarial-MidiBERT: Symbolic Music Understanding Model Based on Unbias Pre-training and Mask Fine-tuning
标题: Adversarial-MidiBERT:基于Unbias预训练和Mass微调的象征性音乐理解模型
作者:Zijian Zhao
链接:点击下载PDF文件
【7】 An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
标题: 一种定位部分虚假音频中操纵区域的无监督域自适应方法
作者:Siding Zeng,Jiangyan Yi,Jianhua Tao,Yujie Chen,Shan Liang,Yong Ren,Xiaohui Zhang
链接:点击下载PDF文件
【8】 Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
标题: 用于视听Zero-Shot学习的Spiking Tucker Fusion Transformer
作者:Wenrui Li,Penghong Wang,Ruiqin Xiong,Xiaopeng Fan
备注:Accepted by TIP
链接:点击下载PDF文件
【9】 Phonetic Richness for Improved Automatic Speaker Verification
标题: 语音丰富度用于改进自动说话人验证
作者:Nicholas Klein,Ganesh Sivaraman,Elie Khoury
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
【10】 Source Tracing of Audio Deepfake Systems
标题: 音频Deepfake系统的来源追踪
作者:Nicholas Klein,Tianxiang Chen,Hemlata Tak,Ricardo Casal,Elie Khoury
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
标题: 语音丰富度用于改进自动说话人验证
作者:Nicholas Klein,Ganesh Sivaraman,Elie Khoury
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
【2】 Source Tracing of Audio Deepfake Systems
标题: 音频Deepfake系统的来源追踪
作者:Nicholas Klein,Tianxiang Chen,Hemlata Tak,Ricardo Casal,Elie Khoury
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【3】 ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
标题: ElasticAST:适用于所有长度和分辨率的音频频谱图Transformer
作者:Jiu Feng,Mehmet Hamza Erol,Joon Son Chung,Arda Senocak
备注:Interspeech 2024. Code is available at this https URL
链接:点击下载PDF文件
【4】 Speech dereverberation constrained on room impulse response characteristics
标题: 语音去回响受限于房间脉冲响应特性
作者:Louis Bahrman,Mathieu Fontaine,Jonathan Le Roux,Gaël Richard
Journal-ref:INTERSPEECH, Sep 2024, Kos Island, Greece
链接:点击下载PDF文件
【5】 From Real to Cloned Singer Identification
标题: 从真实歌手到克隆歌手识别
作者:Dorian Desblancs,Gabriel Meseguer-Brocal,Romain Hennequin,Manuel Moussallam
备注:To be published at ISMIR 2024
链接:点击下载PDF文件
【6】 Autoregressive Speech Synthesis without Vector Quantization
标题: 无需量化的自回归语音合成
作者:Lingwei Meng,Long Zhou,Shujie Liu,Sanyuan Chen,Bing Han,Shujie Hu,Yanqing Liu,Jinyu Li,Sheng Zhao,Xixin Wu,Helen Meng,Furu Wei
链接:点击下载PDF文件
【7】 Adversarial-MidiBERT: Symbolic Music Understanding Model Based on Unbias Pre-training and Mask Fine-tuning
标题: Adversarial-MidiBERT:基于Unbias预训练和Mass微调的象征性音乐理解模型
作者:Zijian Zhao
链接:点击下载PDF文件
【8】 An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
标题: 一种定位部分虚假音频中操纵区域的无监督域自适应方法
作者:Siding Zeng,Jiangyan Yi,Jianhua Tao,Yujie Chen,Shan Liang,Yong Ren,Xiaohui Zhang
链接:点击下载PDF文件
【9】 Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
标题: 用于视听Zero-Shot学习的Spiking Tucker Fusion Transformer
作者:Wenrui Li,Penghong Wang,Ruiqin Xiong,Xiaopeng Fan
备注:Accepted by TIP
链接:点击下载PDF文件
标题: ElasticAST:适用于所有长度和分辨率的音频频谱图Transformer
作者:Jiu Feng,Mehmet Hamza Erol,Joon Son Chung,Arda Senocak
备注:Interspeech 2024. Code is available at this https URL
链接:点击下载PDF文件
摘要:Transformers已迅速取代基于CNN的架构,成为音频分类的新标准。基于变换器的模型,如音频频谱图Transformers(AST),也继承了CNN的固定大小输入范例。然而,当输入长度与训练不同时,这会导致AST在推理中的性能下降。本文介绍了一种在训练和推理过程中使用可变长度音频输入和AST模型的方法。通过采用序列打包,我们的方法ElasticAST在训练期间适应任何音频长度,从而在推理时提供所有长度和分辨率的灵活性。这种灵活性使ElasticAST能够在各种长度或分辨率下保持评估能力,并实现与在特定长度或分辨率下训练的标准AST相似的性能。此外,实验证明了ElasticAST在本机长度音频数据集上训练和评估时的更好性能。摘要:Transformers have rapidly overtaken CNN-based architectures as the new standard in audio classification. Transformer-based models, such as the Audio Spectrogram Transformers (AST), also inherit the fixed-size input paradigm from CNNs. However, this leads to performance degradation for ASTs in the inference when input lengths vary from the training. This paper introduces an approach that enables the use of variable-length audio inputs with AST models during both training and inference. By employing sequence packing, our method ElasticAST, accommodates any audio length during training, thereby offering flexibility across all lengths and resolutions at the inference. This flexibility allows ElasticAST to maintain evaluation capabilities at various lengths or resolutions and achieve similar performance to standard ASTs trained at specific lengths or resolutions. Moreover, experiments demonstrate ElasticAST's better performance when trained and evaluated on native-length audio datasets.
【2】 Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
标题: 评估无人机控制的语音命令管道:从STT和LLM到直接分类和连体网络
作者:Lucca Emmanuel Pineli Simões,Lucas Brandão Rodrigues,Rafaela Mota Silva,Gustavo Rodrigues da Silva
链接:点击下载PDF文件
摘要:本文介绍了使用语音识别和深度学习技术控制Tello无人机的三种语音命令管道的开发和比较评估。其目的是通过实现对无人机动作的直观语音控制来增强人机交互。开发的管道包括:(1)传统的语音到文本(STT),随后是大语言模型(LLM)方法,(2)直接语音到功能映射模型,以及(3)基于暹罗神经网络的系统。每个流水线都根据推理时间、准确性、效率和灵活性进行评估。提供了详细的方法、数据集准备和评估指标,全面分析了每个管道在不同场景中的优势和适用性。摘要:This paper presents the development and comparative evaluation of three voice command pipelines for controlling a Tello drone, using speech recognition and deep learning techniques. The aim is to enhance human-machine interaction by enabling intuitive voice control of drone actions. The pipelines developed include: (1) a traditional Speech-to-Text (STT) followed by a Large Language Model (LLM) approach, (2) a direct voice-to-function mapping model, and (3) a Siamese neural network-based system. Each pipeline was evaluated based on inference time, accuracy, efficiency, and flexibility. Detailed methodologies, dataset preparation, and evaluation metrics are provided, offering a comprehensive analysis of each pipeline's strengths and applicability across different scenarios.
【3】 Speech dereverberation constrained on room impulse response characteristics
标题: 语音去回响受限于房间脉冲响应特性
作者:Louis Bahrman,Mathieu Fontaine,Jonathan Le Roux,Gaël Richard
Journal-ref:INTERSPEECH, Sep 2024, Kos Island, Greece
链接:点击下载PDF文件
摘要:单通道语音去混响的目的是从受室内声反射影响的录音中提取干语音信号。然而,目前大多数基于深度学习的语音去混响方法对于室内声学是不可解释的,并且在这方面可以被认为是黑箱系统。在这项工作中,我们解决了这个问题,通过正规化的训练损失使用一种新的物理相干损失,鼓励房间脉冲响应(RIR)引起的去混响输出的模型,以匹配的声学特性的房间中的信号被记录。我们的调查表明,保存的原始dereverberated信号旁边提供一个更物理相干的RIR。摘要:Single-channel speech dereverberation aims at extracting a dry speech signal from a recording affected by the acoustic reflections in a room. However, most current deep learning-based approaches for speech dereverberation are not interpretable for room acoustics, and can be considered as black-box systems in that regard. In this work, we address this problem by regularizing the training loss using a novel physical coherence loss which encourages the room impulse response (RIR) induced by the dereverberated output of the model to match the acoustic properties of the room in which the signal was recorded. Our investigation demonstrates the preservation of the original dereverberated signal alongside the provision of a more physically coherent RIR.
【4】 From Real to Cloned Singer Identification
标题: 从真实歌手到克隆歌手识别
作者:Dorian Desblancs,Gabriel Meseguer-Brocal,Romain Hennequin,Manuel Moussallam
备注:To be published at ISMIR 2024
链接:点击下载PDF文件
摘要:流行歌手的克隆声音听起来越来越逼真,在过去几年里越来越受欢迎。然而,由于人格权问题,他们对该行业构成了威胁。因此,需要识别合成语音中的原始歌手的方法。在本文中,我们探讨如何歌手识别方法可以用于这样的任务。我们提出了三个嵌入模型,使用歌手级对比学习方案进行训练,其中正对由来自相同歌手的声乐片段组成。这些片段可以是第一个模型的混合,第二个模型的人声,以及第三个模型的两者。我们证明,这三个模型是非常有能力识别真正的歌手。然而,他们的表现恶化时,分类克隆版本的歌手在我们的评估集。这对于使用混合物作为输入的模型尤其如此。这些发现强调了理解歌手识别系统中存在的偏见的必要性,以及它们如何影响音乐中声音深度伪造的识别。摘要:Cloned voices of popular singers sound increasingly realistic and have gained popularity over the past few years. They however pose a threat to the industry due to personality rights concerns. As such, methods to identify the original singer in synthetic voices are needed. In this paper, we investigate how singer identification methods could be used for such a task. We present three embedding models that are trained using a singer-level contrastive learning scheme, where positive pairs consist of segments with vocals from the same singers. These segments can be mixtures for the first model, vocals for the second, and both for the third. We demonstrate that all three models are highly capable of identifying real singers. However, their performance deteriorates when classifying cloned versions of singers in our evaluation set. This is especially true for models that use mixtures as an input. These findings highlight the need to understand the biases that exist within singer identification systems, and how they can influence the identification of voice deepfakes in music.
【5】 Autoregressive Speech Synthesis without Vector Quantization
标题: 无需量化的自回归语音合成
作者:Lingwei Meng,Long Zhou,Shujie Liu,Sanyuan Chen,Bing Han,Shujie Hu,Yanqing Liu,Jinyu Li,Sheng Zhao,Xixin Wu,Helen Meng,Furu Wei
链接:点击下载PDF文件
摘要:我们提出了MELLE,一种新的基于连续值标记的文本到语音合成(TTS)的语言建模方法。MELLE自回归直接从文本条件生成连续的梅尔频谱图帧,绕过了矢量量化的需要,矢量量化最初是为音频压缩而设计的,与梅尔频谱图相比牺牲了保真度。具体来说,(i)代替交叉熵损失,我们应用回归损失和提出的频谱图通量损失函数来模拟连续值令牌的概率分布。(ii)我们在MELLE中加入了变分推理,以简化采样机制,从而增强输出多样性和模型鲁棒性。实验表明,与两阶段编解码器语言模型VALL-E及其变体相比,单阶段MELLE通过避免采样离散代码的固有缺陷来减轻鲁棒性问题,在多个指标上实现卓越的性能,最重要的是,提供了一个更精简的范例。请访问https: aka.ms melle查看我们工作的演示。摘要:We present MELLE, a novel continuous-valued tokens based language modeling approach for text to speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need for vector quantization, which are originally designed for audio compression and sacrifice fidelity compared to mel-spectrograms. Specifically, (i) instead of cross-entropy loss, we apply regression loss with a proposed spectrogram flux loss function to model the probability distribution of the continuous-valued tokens. (ii) we have incorporated variational inference into MELLE to facilitate sampling mechanisms, thereby enhancing the output diversity and model robustness. Experiments demonstrate that, compared to the two-stage codec language models VALL-E and its variants, the single-stage MELLE mitigates robustness issues by avoiding the inherent flaws of sampling discrete codes, achieves superior performance across multiple metrics, and, most importantly, offers a more streamlined paradigm. See https: aka.ms melle for demos of our work.
【6】 Adversarial-MidiBERT: Symbolic Music Understanding Model Based on Unbias Pre-training and Mask Fine-tuning
标题: Adversarial-MidiBERT:基于Unbias预训练和Mass微调的象征性音乐理解模型
作者:Zijian Zhao
链接:点击下载PDF文件
摘要:作为音乐信息检索的重要组成部分,符号音乐理解(SMU)因其能够帮助音乐爱好者学习和创作音乐而受到广泛关注。近年来,由于符号音乐与自然语言具有巨大的相似性,预训练的语言模型在SMU中得到了广泛的应用,并且预训练的方式也有助于充分利用有限的音乐数据。然而,在预先训练的语言模型中观察到偏见问题,如性别歧视,年龄歧视和种族主义,这归因于训练数据的不平衡分布。它对下游任务的性能也有显著影响,这也发生在SMU中。为了解决这一挑战,我们提出了一个符号音乐理解模型的基础上双向编码器表示从Transformers(BERT)的Adversarial-MidiBERT。我们引入了一种基于对抗学习的无偏预训练方法,以最大限度地减少在训练过程中导致偏见的令牌的参与。此外,我们还提出了一种模板微调方法来缩小预训练和微调之间的数据差距,这可以帮助模型更快地收敛并表现得更好。我们评估我们的方法对四个音乐理解任务,我们的方法在所有这些都表现出出色的性能。我们模型的代码可在https: github.com RS2002 Adversarial-MidiBERT上公开获得。摘要:As an important part of Music Information Retrieval (MIR), Symbolic Music Understanding (SMU) has gained substantial attention, as it can assist musicians and amateurs in learning and creating music. Recently, pre-trained language models have been widely adopted in SMU because the symbolic music shares a huge similarity with natural language, and the pre-trained manner also helps make full use of limited music data. However, the issue of bias, such as sexism, ageism, and racism, has been observed in pre-trained language models, which is attributed to the imbalanced distribution of training data. It also has a significant influence on the performance of downstream tasks, which also happens in SMU. To address this challenge, we propose Adversarial-MidiBERT, a symbolic music understanding model based on Bidirectional Encoder Representations from Transformers (BERT). We introduce an unbiased pre-training method based on adversarial learning to minimize the participation of tokens that lead to biases during training. Furthermore, we propose a mask fine-tuning method to narrow the data gap between pre-training and fine-tuning, which can help the model converge faster and perform better. We evaluate our method on four music understanding tasks, and our approach demonstrates excellent performance in all of them. The code for our model is publicly available at https: github.com RS2002 Adversarial-MidiBERT.
【7】 An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
标题: 一种定位部分虚假音频中操纵区域的无监督域自适应方法
作者:Siding Zeng,Jiangyan Yi,Jianhua Tao,Yujie Chen,Shan Liang,Yong Ren,Xiaohui Zhang
链接:点击下载PDF文件
摘要:当在部分伪音频(PFA)中定位操作区域的任务涉及跨域数据集时,由于源域和目标域之间的偏移,深度学习模型的性能会显著下降。为了解决这个问题,现有的方法通常在训练之前使用数据增强。然而,它们忽略了源域中不存在的目标域的特征。受混合专家模型的启发,本文提出了一种无监督的基于多样性和熵的样本挖掘方法。我们的方法首先从一组不同的专家那里学习,这些专家从源域的不同角度获得了很好的性能,但在目标样本上存在模糊性。我们利用这些不同的专家,通过计算熵来选择信息量最大的样本。此外,我们还介绍了一种为这些选定的样本量身定制的标签生成方法,这些样本被纳入源域的训练过程中,并整合目标域信息。我们将我们的方法应用于跨域部分虚假音频检测数据集ADD2023Track2。通过从目标域引入10%的未知样本,我们获得了43.84%的F1分数,与第二好的方法相比,相对增加了77.2%。摘要:When the task of locating manipulation regions in partially-fake audio (PFA) involves cross-domain datasets, the performance of deep learning models drops significantly due to the shift between the source and target domains. To address this issue, existing approaches often employ data augmentation before training. However, they overlook the characteristics in target domain that are absent in source domain. Inspired by the mixture-of-experts model, we propose an unsupervised method named Samples mining with Diversity and Entropy (SDE). Our method first learns from a collection of diverse experts that achieve great performance from different perspectives in the source domain, but with ambiguity on target samples. We leverage these diverse experts to select the most informative samples by calculating their entropy. Furthermore, we introduced a label generation method tailored for these selected samples that are incorporated in the training process in source domain integrating the target domain information. We applied our method to a cross-domain partially fake audio detection dataset, ADD2023Track2. By introducing 10% of unknown samples from the target domain, we achieved an F1 score of 43.84%, which represents a relative increase of 77.2% compared to the second-best method.
【8】 Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
标题: 用于视听Zero-Shot学习的Spiking Tucker Fusion Transformer
作者:Wenrui Li,Penghong Wang,Ruiqin Xiong,Xiaopeng Fan
备注:Accepted by TIP
链接:点击下载PDF文件
摘要:脉冲神经网络(SNN),有效地编码时间序列已显示出巨大的潜力,在提取视听联合特征表示。然而,耦合SNN(二进制尖峰序列)和Transformers(浮点序列),以共同探索的时间语义信息仍然面临挑战。在本文中,我们介绍了一种新的尖峰塔克融合Transformer(STFT)的视听zero-shot学习(STNL)。STFT利用来自不同时间步长的时间和语义信息来生成鲁棒的表示。引入时间步长因子(TSF)动态合成后续推理信息。为了引导输入膜电位的形成并降低尖峰噪声,我们提出了一种全局-局部池化(GLP),它结合了最大和平均池化操作。此外,基于语义和时间线索动态调整尖峰神经元的阈值。由于在简单的双线性模型中增加了参数的数量,因此很难集成SNN和Transformers提取的时间和语义信息。为了解决这个问题,我们引入了一个时间语义Tucker融合模块,它实现了SNN和Transformer输出的多尺度融合,同时保持完整的二阶交互。我们的实验结果表明,所提出的方法在实现国家的最先进的性能在三个基准数据集的有效性。VGGSound、UCF 101和ActivityNet的谐波平均(HM)改善分别约为15.4%、3.9%和14.9%。摘要:The spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling SNNs (binary spike sequences) with transformers (float-point sequences) to jointly explore the temporal-semantic information still facing challenges. In this paper, we introduce a novel Spiking Tucker Fusion Transformer (STFT) for audio-visual zero-shot learning (ZSL). The STFT leverage the temporal and semantic information from different time steps to generate robust representations. The time-step factor (TSF) is introduced to dynamically synthesis the subsequent inference information. To guide the formation of input membrane potentials and reduce the spike noise, we propose a global-local pooling (GLP) which combines the max and average pooling operations. Furthermore, the thresholds of the spiking neurons are dynamically adjusted based on semantic and temporal cues. Integrating the temporal and semantic information extracted by SNNs and Transformers are difficult due to the increased number of parameters in a straightforward bilinear model. To address this, we introduce a temporal-semantic Tucker fusion module, which achieves multi-scale fusion of SNN and Transformer outputs while maintaining full second-order interactions. Our experimental results demonstrate the effectiveness of the proposed approach in achieving state-of-the-art performance in three benchmark datasets. The harmonic mean (HM) improvement of VGGSound, UCF101 and ActivityNet are around 15.4 %, 3.9 %, and 14.9 %, respectively.
【9】 Phonetic Richness for Improved Automatic Speaker Verification
标题: 语音丰富度用于改进自动说话人验证
作者:Nicholas Klein,Ganesh Sivaraman,Elie Khoury
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
摘要:当涉及到说话人验证系统中的身份验证时,并非所有的话语都是平等的。为了考虑不同的声学条件,必须估计测试话语的质量。除了话语的网络语音持续时间之外,本文观察到语音丰富度也是话语质量的关键指标,在准确的说话人确认中起着重要作用。几个语音直方图为基础的公式的语音丰富度进行了探讨,使用自动说话人识别系统获得的成绩单。建议的语音丰富度的措施被发现是正相关的语音认证分数在整个评估基准。此外,所提出的措施与净语音相结合,有助于校准说话人验证分数,在Voxceleb1评估协议上获得5.8%的相对EER改善。所提出的基于语音丰富度的校准为具有重复单词的短话语提供了更高的益处。摘要:When it comes to authentication in speaker verification systems, not all utterances are created equal. It is essential to estimate the quality of test utterances in order to account for varying acoustic conditions. In addition to the net-speech duration of an utterance, it is observed in this paper that phonetic richness is also a key indicator of utterance quality, playing a significant role in accurate speaker verification. Several phonetic histogram based formulations of phonetic richness are explored using transcripts obtained from an automatic speaker recognition system. The proposed phonetic richness measure is found to be positively correlated with voice authentication scores across evaluation benchmarks. Additionally, the proposed measure in combination with net speech helps in calibrating the speaker verification scores, obtaining a relative EER improvement of 5.8% on the Voxceleb1 evaluation protocol. The proposed phonetic richness based calibration provides higher benefit for short utterances with repeated words.
【10】 Source Tracing of Audio Deepfake Systems
标题: 音频Deepfake系统的来源追踪
作者:Nicholas Klein,Tianxiang Chen,Hemlata Tak,Ricardo Casal,Elie Khoury
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:生成AI技术的最新进展使音频deepfake变得更加逼真。虽然目前对反欺骗系统的研究主要集中在评估给定的音频样本是假的还是真的,但对识别创建音频deepfakes的特定技术的关注有限。音频deepfake生成中常用的算法,如文本到语音(TTS)和语音转换(VC),经历不同的阶段,包括输入处理,声学建模和波形生成。在这项工作中,我们介绍了一个系统,旨在对各种欺骗属性进行分类,在整个生成管道中捕获各个模块的独特功能。我们在两个数据集上评估了我们的系统:ASVspoof 2019逻辑访问和多语言音频反欺骗数据集(MLAAD)。两个实验的结果都证明了系统识别deepfake生成系统的不同欺骗属性的鲁棒性。摘要:Recent progress in generative AI technology has made audio deepfakes remarkably more realistic. While current research on anti-spoofing systems primarily focuses on assessing whether a given audio sample is fake or genuine, there has been limited attention on discerning the specific techniques to create the audio deepfakes. Algorithms commonly used in audio deepfake generation, like text-to-speech (TTS) and voice conversion (VC), undergo distinct stages including input processing, acoustic modeling, and waveform generation. In this work, we introduce a system designed to classify various spoofing attributes, capturing the distinctive features of individual modules throughout the entire generation pipeline. We evaluate our system on two datasets: the ASVspoof 2019 Logical Access and the Multi-Language Audio Anti-Spoofing Dataset (MLAAD). Results from both experiments demonstrate the robustness of the system to identify the different spoofing attributes of deepfake generation systems.
eess.AS音频处理
【1】 Phonetic Richness for Improved Automatic Speaker Verification标题: 语音丰富度用于改进自动说话人验证
作者:Nicholas Klein,Ganesh Sivaraman,Elie Khoury
备注:Accepted by EUSIPCO 2024
链接:点击下载PDF文件
摘要:当涉及到说话人验证系统中的身份验证时,并非所有的话语都是平等的。为了考虑不同的声学条件,必须估计测试话语的质量。除了话语的网络语音持续时间之外,本文观察到语音丰富度也是话语质量的关键指标,在准确的说话人确认中起着重要作用。几个语音直方图为基础的公式的语音丰富度进行了探讨,使用自动说话人识别系统获得的成绩单。建议的语音丰富度的措施被发现是正相关的语音认证分数在整个评估基准。此外,所提出的措施与净语音相结合,有助于校准说话人验证分数,在Voxceleb1评估协议上获得5.8%的相对EER改善。所提出的基于语音丰富度的校准为具有重复单词的短话语提供了更高的益处。摘要:When it comes to authentication in speaker verification systems, not all utterances are created equal. It is essential to estimate the quality of test utterances in order to account for varying acoustic conditions. In addition to the net-speech duration of an utterance, it is observed in this paper that phonetic richness is also a key indicator of utterance quality, playing a significant role in accurate speaker verification. Several phonetic histogram based formulations of phonetic richness are explored using transcripts obtained from an automatic speaker recognition system. The proposed phonetic richness measure is found to be positively correlated with voice authentication scores across evaluation benchmarks. Additionally, the proposed measure in combination with net speech helps in calibrating the speaker verification scores, obtaining a relative EER improvement of 5.8% on the Voxceleb1 evaluation protocol. The proposed phonetic richness based calibration provides higher benefit for short utterances with repeated words.
【2】 Source Tracing of Audio Deepfake Systems
标题: 音频Deepfake系统的来源追踪
作者:Nicholas Klein,Tianxiang Chen,Hemlata Tak,Ricardo Casal,Elie Khoury
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:生成AI技术的最新进展使音频deepfake变得更加逼真。虽然目前对反欺骗系统的研究主要集中在评估给定的音频样本是假的还是真的,但对识别创建音频deepfakes的特定技术的关注有限。音频deepfake生成中常用的算法,如文本到语音(TTS)和语音转换(VC),经历不同的阶段,包括输入处理,声学建模和波形生成。在这项工作中,我们介绍了一个系统,旨在对各种欺骗属性进行分类,在整个生成管道中捕获各个模块的独特功能。我们在两个数据集上评估了我们的系统:ASVspoof 2019逻辑访问和多语言音频反欺骗数据集(MLAAD)。两个实验的结果都证明了系统识别deepfake生成系统的不同欺骗属性的鲁棒性。摘要:Recent progress in generative AI technology has made audio deepfakes remarkably more realistic. While current research on anti-spoofing systems primarily focuses on assessing whether a given audio sample is fake or genuine, there has been limited attention on discerning the specific techniques to create the audio deepfakes. Algorithms commonly used in audio deepfake generation, like text-to-speech (TTS) and voice conversion (VC), undergo distinct stages including input processing, acoustic modeling, and waveform generation. In this work, we introduce a system designed to classify various spoofing attributes, capturing the distinctive features of individual modules throughout the entire generation pipeline. We evaluate our system on two datasets: the ASVspoof 2019 Logical Access and the Multi-Language Audio Anti-Spoofing Dataset (MLAAD). Results from both experiments demonstrate the robustness of the system to identify the different spoofing attributes of deepfake generation systems.
【3】 ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
标题: ElasticAST:适用于所有长度和分辨率的音频频谱图Transformer
作者:Jiu Feng,Mehmet Hamza Erol,Joon Son Chung,Arda Senocak
备注:Interspeech 2024. Code is available at this https URL
链接:点击下载PDF文件
摘要:Transformers已迅速取代基于CNN的架构,成为音频分类的新标准。基于变换器的模型,如音频频谱图Transformers(AST),也继承了CNN的固定大小输入范例。然而,当输入长度与训练不同时,这会导致AST在推理中的性能下降。本文介绍了一种在训练和推理过程中使用可变长度音频输入和AST模型的方法。通过采用序列打包,我们的方法ElasticAST在训练期间适应任何音频长度,从而在推理时提供所有长度和分辨率的灵活性。这种灵活性使ElasticAST能够在各种长度或分辨率下保持评估能力,并实现与在特定长度或分辨率下训练的标准AST相似的性能。此外,实验证明了ElasticAST在本机长度音频数据集上训练和评估时的更好性能。摘要:Transformers have rapidly overtaken CNN-based architectures as the new standard in audio classification. Transformer-based models, such as the Audio Spectrogram Transformers (AST), also inherit the fixed-size input paradigm from CNNs. However, this leads to performance degradation for ASTs in the inference when input lengths vary from the training. This paper introduces an approach that enables the use of variable-length audio inputs with AST models during both training and inference. By employing sequence packing, our method ElasticAST, accommodates any audio length during training, thereby offering flexibility across all lengths and resolutions at the inference. This flexibility allows ElasticAST to maintain evaluation capabilities at various lengths or resolutions and achieve similar performance to standard ASTs trained at specific lengths or resolutions. Moreover, experiments demonstrate ElasticAST's better performance when trained and evaluated on native-length audio datasets.
【4】 Speech dereverberation constrained on room impulse response characteristics
标题: 语音去回响受限于房间脉冲响应特性
作者:Louis Bahrman,Mathieu Fontaine,Jonathan Le Roux,Gaël Richard
Journal-ref:INTERSPEECH, Sep 2024, Kos Island, Greece
链接:点击下载PDF文件
摘要:单通道语音去混响的目的是从受室内声反射影响的录音中提取干语音信号。然而,目前大多数基于深度学习的语音去混响方法对于室内声学是不可解释的,并且在这方面可以被认为是黑箱系统。在这项工作中,我们解决了这个问题,通过正规化的训练损失使用一种新的物理相干损失,鼓励房间脉冲响应(RIR)引起的去混响输出的模型,以匹配的声学特性的房间中的信号被记录。我们的调查表明,保存的原始dereverberated信号旁边提供一个更物理相干的RIR。摘要:Single-channel speech dereverberation aims at extracting a dry speech signal from a recording affected by the acoustic reflections in a room. However, most current deep learning-based approaches for speech dereverberation are not interpretable for room acoustics, and can be considered as black-box systems in that regard. In this work, we address this problem by regularizing the training loss using a novel physical coherence loss which encourages the room impulse response (RIR) induced by the dereverberated output of the model to match the acoustic properties of the room in which the signal was recorded. Our investigation demonstrates the preservation of the original dereverberated signal alongside the provision of a more physically coherent RIR.
【5】 From Real to Cloned Singer Identification
标题: 从真实歌手到克隆歌手识别
作者:Dorian Desblancs,Gabriel Meseguer-Brocal,Romain Hennequin,Manuel Moussallam
备注:To be published at ISMIR 2024
链接:点击下载PDF文件
摘要:流行歌手的克隆声音听起来越来越逼真,在过去几年里越来越受欢迎。然而,由于人格权问题,他们对该行业构成了威胁。因此,需要识别合成语音中的原始歌手的方法。在本文中,我们探讨如何歌手识别方法可以用于这样的任务。我们提出了三个嵌入模型,使用歌手级对比学习方案进行训练,其中正对由来自相同歌手的声乐片段组成。这些片段可以是第一个模型的混合,第二个模型的人声,以及第三个模型的两者。我们证明,这三个模型是非常有能力识别真正的歌手。然而,他们的表现恶化时,分类克隆版本的歌手在我们的评估集。这对于使用混合物作为输入的模型尤其如此。这些发现强调了理解歌手识别系统中存在的偏见的必要性,以及它们如何影响音乐中声音深度伪造的识别。摘要:Cloned voices of popular singers sound increasingly realistic and have gained popularity over the past few years. They however pose a threat to the industry due to personality rights concerns. As such, methods to identify the original singer in synthetic voices are needed. In this paper, we investigate how singer identification methods could be used for such a task. We present three embedding models that are trained using a singer-level contrastive learning scheme, where positive pairs consist of segments with vocals from the same singers. These segments can be mixtures for the first model, vocals for the second, and both for the third. We demonstrate that all three models are highly capable of identifying real singers. However, their performance deteriorates when classifying cloned versions of singers in our evaluation set. This is especially true for models that use mixtures as an input. These findings highlight the need to understand the biases that exist within singer identification systems, and how they can influence the identification of voice deepfakes in music.
【6】 Autoregressive Speech Synthesis without Vector Quantization
标题: 无需量化的自回归语音合成
作者:Lingwei Meng,Long Zhou,Shujie Liu,Sanyuan Chen,Bing Han,Shujie Hu,Yanqing Liu,Jinyu Li,Sheng Zhao,Xixin Wu,Helen Meng,Furu Wei
链接:点击下载PDF文件
摘要:我们提出了MELLE,一种新的基于连续值标记的文本到语音合成(TTS)的语言建模方法。MELLE自回归直接从文本条件生成连续的梅尔频谱图帧,绕过了矢量量化的需要,矢量量化最初是为音频压缩而设计的,与梅尔频谱图相比牺牲了保真度。具体来说,(i)代替交叉熵损失,我们应用回归损失和提出的频谱图通量损失函数来模拟连续值令牌的概率分布。(ii)我们在MELLE中加入了变分推理,以简化采样机制,从而增强输出多样性和模型鲁棒性。实验表明,与两阶段编解码器语言模型VALL-E及其变体相比,单阶段MELLE通过避免采样离散代码的固有缺陷来减轻鲁棒性问题,在多个指标上实现卓越的性能,最重要的是,提供了一个更精简的范例。请访问https: aka.ms melle查看我们工作的演示。摘要:We present MELLE, a novel continuous-valued tokens based language modeling approach for text to speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need for vector quantization, which are originally designed for audio compression and sacrifice fidelity compared to mel-spectrograms. Specifically, (i) instead of cross-entropy loss, we apply regression loss with a proposed spectrogram flux loss function to model the probability distribution of the continuous-valued tokens. (ii) we have incorporated variational inference into MELLE to facilitate sampling mechanisms, thereby enhancing the output diversity and model robustness. Experiments demonstrate that, compared to the two-stage codec language models VALL-E and its variants, the single-stage MELLE mitigates robustness issues by avoiding the inherent flaws of sampling discrete codes, achieves superior performance across multiple metrics, and, most importantly, offers a more streamlined paradigm. See https: aka.ms melle for demos of our work.
【7】 Adversarial-MidiBERT: Symbolic Music Understanding Model Based on Unbias Pre-training and Mask Fine-tuning
标题: Adversarial-MidiBERT:基于Unbias预训练和Mass微调的象征性音乐理解模型
作者:Zijian Zhao
链接:点击下载PDF文件
摘要:作为音乐信息检索的重要组成部分,符号音乐理解(SMU)因其能够帮助音乐爱好者学习和创作音乐而受到广泛关注。近年来,由于符号音乐与自然语言具有巨大的相似性,预训练的语言模型在SMU中得到了广泛的应用,并且预训练的方式也有助于充分利用有限的音乐数据。然而,在预先训练的语言模型中观察到偏见问题,如性别歧视,年龄歧视和种族主义,这归因于训练数据的不平衡分布。它对下游任务的性能也有显著影响,这也发生在SMU中。为了解决这一挑战,我们提出了一个符号音乐理解模型的基础上双向编码器表示从Transformers(BERT)的Adversarial-MidiBERT。我们引入了一种基于对抗学习的无偏预训练方法,以最大限度地减少在训练过程中导致偏见的令牌的参与。此外,我们还提出了一种模板微调方法来缩小预训练和微调之间的数据差距,这可以帮助模型更快地收敛并表现得更好。我们评估我们的方法对四个音乐理解任务,我们的方法在所有这些都表现出出色的性能。我们模型的代码可在https: github.com RS2002 Adversarial-MidiBERT上公开获得。摘要:As an important part of Music Information Retrieval (MIR), Symbolic Music Understanding (SMU) has gained substantial attention, as it can assist musicians and amateurs in learning and creating music. Recently, pre-trained language models have been widely adopted in SMU because the symbolic music shares a huge similarity with natural language, and the pre-trained manner also helps make full use of limited music data. However, the issue of bias, such as sexism, ageism, and racism, has been observed in pre-trained language models, which is attributed to the imbalanced distribution of training data. It also has a significant influence on the performance of downstream tasks, which also happens in SMU. To address this challenge, we propose Adversarial-MidiBERT, a symbolic music understanding model based on Bidirectional Encoder Representations from Transformers (BERT). We introduce an unbiased pre-training method based on adversarial learning to minimize the participation of tokens that lead to biases during training. Furthermore, we propose a mask fine-tuning method to narrow the data gap between pre-training and fine-tuning, which can help the model converge faster and perform better. We evaluate our method on four music understanding tasks, and our approach demonstrates excellent performance in all of them. The code for our model is publicly available at https: github.com RS2002 Adversarial-MidiBERT.
【8】 An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
标题: 一种定位部分虚假音频中操纵区域的无监督域自适应方法
作者:Siding Zeng,Jiangyan Yi,Jianhua Tao,Yujie Chen,Shan Liang,Yong Ren,Xiaohui Zhang
链接:点击下载PDF文件
摘要:当在部分伪音频(PFA)中定位操作区域的任务涉及跨域数据集时,由于源域和目标域之间的偏移,深度学习模型的性能会显著下降。为了解决这个问题,现有的方法通常在训练之前使用数据增强。然而,它们忽略了源域中不存在的目标域的特征。受混合专家模型的启发,本文提出了一种无监督的基于多样性和熵的样本挖掘方法。我们的方法首先从一组不同的专家那里学习,这些专家从源域的不同角度获得了很好的性能,但在目标样本上存在模糊性。我们利用这些不同的专家,通过计算熵来选择信息量最大的样本。此外,我们还介绍了一种为这些选定的样本量身定制的标签生成方法,这些样本被纳入源域的训练过程中,并整合目标域信息。我们将我们的方法应用于跨域部分虚假音频检测数据集ADD2023Track2。通过从目标域引入10%的未知样本,我们获得了43.84%的F1分数,与第二好的方法相比,相对增加了77.2%。摘要:When the task of locating manipulation regions in partially-fake audio (PFA) involves cross-domain datasets, the performance of deep learning models drops significantly due to the shift between the source and target domains. To address this issue, existing approaches often employ data augmentation before training. However, they overlook the characteristics in target domain that are absent in source domain. Inspired by the mixture-of-experts model, we propose an unsupervised method named Samples mining with Diversity and Entropy (SDE). Our method first learns from a collection of diverse experts that achieve great performance from different perspectives in the source domain, but with ambiguity on target samples. We leverage these diverse experts to select the most informative samples by calculating their entropy. Furthermore, we introduced a label generation method tailored for these selected samples that are incorporated in the training process in source domain integrating the target domain information. We applied our method to a cross-domain partially fake audio detection dataset, ADD2023Track2. By introducing 10% of unknown samples from the target domain, we achieved an F1 score of 43.84%, which represents a relative increase of 77.2% compared to the second-best method.
【9】 Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
标题: 用于视听Zero-Shot学习的Spiking Tucker Fusion Transformer
作者:Wenrui Li,Penghong Wang,Ruiqin Xiong,Xiaopeng Fan
备注:Accepted by TIP
链接:点击下载PDF文件
摘要:脉冲神经网络(SNN),有效地编码时间序列已显示出巨大的潜力,在提取视听联合特征表示。然而,耦合SNN(二进制尖峰序列)和Transformers(浮点序列),以共同探索的时间语义信息仍然面临挑战。在本文中,我们介绍了一种新的尖峰塔克融合Transformer(STFT)的视听zero-shot学习(STNL)。STFT利用来自不同时间步长的时间和语义信息来生成鲁棒的表示。引入时间步长因子(TSF)动态合成后续推理信息。为了引导输入膜电位的形成并降低尖峰噪声,我们提出了一种全局-局部池化(GLP),它结合了最大和平均池化操作。此外,基于语义和时间线索动态调整尖峰神经元的阈值。由于在简单的双线性模型中增加了参数的数量,因此很难集成SNN和Transformers提取的时间和语义信息。为了解决这个问题,我们引入了一个时间语义Tucker融合模块,它实现了SNN和Transformer输出的多尺度融合,同时保持完整的二阶交互。我们的实验结果表明,所提出的方法在实现国家的最先进的性能在三个基准数据集的有效性。VGGSound、UCF 101和ActivityNet的谐波平均(HM)改善分别约为15.4%、3.9%和14.9%。摘要:The spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling SNNs (binary spike sequences) with transformers (float-point sequences) to jointly explore the temporal-semantic information still facing challenges. In this paper, we introduce a novel Spiking Tucker Fusion Transformer (STFT) for audio-visual zero-shot learning (ZSL). The STFT leverage the temporal and semantic information from different time steps to generate robust representations. The time-step factor (TSF) is introduced to dynamically synthesis the subsequent inference information. To guide the formation of input membrane potentials and reduce the spike noise, we propose a global-local pooling (GLP) which combines the max and average pooling operations. Furthermore, the thresholds of the spiking neurons are dynamically adjusted based on semantic and temporal cues. Integrating the temporal and semantic information extracted by SNNs and Transformers are difficult due to the increased number of parameters in a straightforward bilinear model. To address this, we introduce a temporal-semantic Tucker fusion module, which achieves multi-scale fusion of SNN and Transformer outputs while maintaining full second-order interactions. Our experimental results demonstrate the effectiveness of the proposed approach in achieving state-of-the-art performance in three benchmark datasets. The harmonic mean (HM) improvement of VGGSound, UCF101 and ActivityNet are around 15.4 %, 3.9 %, and 14.9 %, respectively.
机器翻译,仅供参考
