今日论文合集:cs.SD语音12篇,eess.AS音频处理13篇。本文经arXiv每日学术速递授权转载
【1】Certification of Speaker Recognition Models to Additive Perturbations
作者:Dmitrii Korzh,Elvir Karimov,Mikhail Pautov,Oleg Y. Rogov,Ivan Oseledets摘要:说话人识别技术应用于从个人虚拟助理到安全访问系统的各种任务中。然而,这些系统对对抗性攻击的鲁棒性,特别是对加性扰动的鲁棒性,仍然是一个重大挑战。在本文中,我们率先将鲁棒性认证技术应用于说话人识别,最初是为图像域开发的。在我们的工作中,我们通过转移和改进随机平滑认证技术对范数有界的附加扰动分类和Few-Shot学习任务的说话人识别来弥补这一差距。我们证明了这些方法在VoxCeleb 1和2数据集上的有效性。我们希望这项工作能够提高语音生物特征的鲁棒性,建立一个新的认证基准,并加快在音频领域的认证方法的研究。摘要:Speaker recognition technology is applied in various tasks ranging from personal virtual assistants to secure access systems. However, the robustness of these systems against adversarial attacks, particularly to additive perturbations, remains a significant challenge. In this paper, we pioneer applying robustness certification techniques to speaker recognition, originally developed for the image domain. In our work, we cover this gap by transferring and improving randomized smoothing certification techniques against norm-bounded additive perturbations for classification and few-shot learning tasks to speaker recognition. We demonstrate the effectiveness of these methods on VoxCeleb 1 and 2 datasets for several models. We expect this work to improve voice-biometry robustness, establish a new certification benchmark, and accelerate research of certification methods in the audio domain.
【2】 A Systematic Evaluation of Adversarial Attacks against Speech Emotion Recognition Models作者:Nicolas Facchinetti,Federico Simonetta,Stavros Ntalampiras摘要:近年来,语音情感识别(SER)由于其在不同领域的潜在应用以及深度学习技术提供的可能性而不断受到关注。然而,最近的研究表明,深度学习模型很容易受到对抗性攻击。在本文中,我们通过检查SER背景下各种对抗性白盒和黑盒攻击对不同语言和性别的影响,系统地评估了这个问题。我们首先提出了一种适合音频数据处理,特征提取和CNN—LSTM架构的方法。观察到的结果强调了CNN—LSTM模型对对抗性示例(AE)的显著脆弱性。事实上,所有考虑的对抗性攻击都能够显着降低构建模型的性能。此外,在评估攻击的效果时,注意到所分析的语言之间以及男性和女性语言之间的微小差异。总之,这项工作有助于理解CNN—LSTM模型的鲁棒性,特别是在SER场景中,以及AE的影响。有趣的是,我们的研究结果可以作为a)开发更强大的SER算法,b)设计更有效的攻击,c)调查可能的防御措施,d)提高对不同语言和性别之间声音差异的理解,e)总体而言,增强我们对SER任务的理解。摘要:Speech emotion recognition (SER) is constantly gaining attention in recent years due to its potential applications in diverse fields and thanks to the possibility offered by deep learning technologies. However, recent studies have shown that deep learning models can be vulnerable to adversarial attacks. In this paper, we systematically assess this problem by examining the impact of various adversarial white-box and black-box attacks on different languages and genders within the context of SER. We first propose a suitable methodology for audio data processing, feature extraction, and CNN-LSTM architecture. The observed outcomes highlighted the significant vulnerability of CNN-LSTM models to adversarial examples (AEs). In fact, all the considered adversarial attacks are able to significantly reduce the performance of the constructed models. Furthermore, when assessing the efficacy of the attacks, minor differences were noted between the languages analyzed as well as between male and female speech. In summary, this work contributes to the understanding of the robustness of CNN-LSTM models, particularly in SER scenarios, and the impact of AEs. Interestingly, our findings serve as a baseline for a) developing more robust algorithms for SER, b) designing more effective attacks, c) investigating possible defenses, d) improved understanding of the vocal differences between different languages and genders, and e) overall, enhancing our comprehension of the SER task.
【3】 Pièces de viole des Cinq Livres and their statistical signatures: the musical work of Marin Marais and Jordi Savall标题:Pièces de viole des Cinq Livres及其统计签名:Marin Marais和Jordi Savall的音乐作品作者:Igor Lugo,Martha G. Alatriste-Contreras摘要:本研究基于Marin Marais和Jordi Savall之间的合作工作,分析了与“Pieces de viole des Cinq Livres”相关的音频信号的频谱,以获取潜在的音乐信息。特别是,我们探讨可能的统计签名相关的音乐作品的识别。基于复杂系统方法,我们计算音频信号的频谱,分析和识别它们的最佳拟合统计分布,并使用科学音高表示法绘制它们的相对频率。研究结果表明,收集的频率分量相关的频谱的每一本书,形成这个音频工作显示高度倾斜和相关的统计分布。因此,最好地描述这些音频数据的集合并且可以与奇异统计签名相关联的最频繁的统计分布是指数分布。摘要:This study analyzes the spectrum of audio signals related to the work of "Pi\`eces de viole des Cinq Livres" based on the collaborative work between Marin Marais and Jordi Savall for the underlying musical information. In particular, we explore the identification of possible statistical signatures related to this musical work. Based on the complex systems approach, we compute the spectrum of audio signals, analyze and identify their best-fit statistical distributions, and plot their relative frequencies using the scientific pitch notation. Findings suggest that the collection of frequency components related to the spectrum of each of the books that form this audio work show highly skewed and associated statistical distributions. Therefore, the most frequent statistical distribution that best describes the collection of these audio data and may be associated with a singular statistical signature is the exponential.【4】 USAT: A Universal Speaker-Adaptive Text-to-Speech Approach作者:Wenbin Wang,Yang Song,Sanjay Jha摘要:传统的文语转换(TTS)研究主要集中在提高训练数据集中说话人的合成语音质量。为看不见的、数据集外的说话者(特别是那些参考数据有限的说话者)合成逼真的语音的挑战仍然是一个重大且未解决的问题。虽然已经探索了zero-shot或Few-Shot说话者自适应TTS方法,但是它们具有许多限制。Zero-shot方法往往遭受泛化性能不足,以再现说话人的声音与沉重的口音。虽然Few-Shot方法可以再现高度变化的口音,但它们带来了显著的存储负担以及过拟合和灾难性遗忘的风险。此外,现有方法仅提供zero-shot或Few-Shot自适应,限制了它们在具有不同需求的各种现实世界场景中的效用。此外,目前大多数的说话人自适应TTS的评估只进行本族语者的数据集,无意中忽略了很大一部分的非本族语者与不同的口音。我们提出的框架统一了zero-shot和Few-Shot的说话人自适应策略,我们称之为“即时”和“细粒度”的适应基于他们的优点。为了缓解在zero-shot说话人自适应中观察到的泛化性能不足,我们设计了两个创新的鉴别器,并为语音解码器引入了记忆机制。为了防止灾难性的遗忘和减少存储的影响,Few-Shot扬声器适应,我们设计了两个适配器和一个独特的适应过程。摘要:Conventional text-to-speech (TTS) research has predominantly focused on enhancing the quality of synthesized speech for speakers in the training dataset. The challenge of synthesizing lifelike speech for unseen, out-of-dataset speakers, especially those with limited reference data, remains a significant and unresolved problem. While zero-shot or few-shot speaker-adaptive TTS approaches have been explored, they have many limitations. Zero-shot approaches tend to suffer from insufficient generalization performance to reproduce the voice of speakers with heavy accents. While few-shot methods can reproduce highly varying accents, they bring a significant storage burden and the risk of overfitting and catastrophic forgetting. In addition, prior approaches only provide either zero-shot or few-shot adaptation, constraining their utility across varied real-world scenarios with different demands. Besides, most current evaluations of speaker-adaptive TTS are conducted only on datasets of native speakers, inadvertently neglecting a vast portion of non-native speakers with diverse accents. Our proposed framework unifies both zero-shot and few-shot speaker adaptation strategies, which we term as "instant" and "fine-grained" adaptations based on their merits. To alleviate the insufficient generalization performance observed in zero-shot speaker adaptation, we designed two innovative discriminators and introduced a memory mechanism for the speech decoder. To prevent catastrophic forgetting and reduce storage implications for few-shot speaker adaptation, we designed two adapters and a unique adaptation procedure.【5】 ComposerX: Multi-Agent Symbolic Music Composition with LLMs标题:ComposerX:使用LLM的多智能体符号音乐创作作者:Qixin Deng,Qikai Yang,Ruibin Yuan,Yipeng Huang,Yi Wang,Xubo Liu,Zeyue Tian,Jiahao Pan,Ge Zhang,Hanfeng Lin,Yizhi Li,Yinghao Ma,Jie Fu,Chenghua Lin,Emmanouil Benetos,Wenwu Wang,Guangyu Xia,Wei Xue,Yike Guo摘要:音乐创作代表了人类创造性的一面,它本身是一项复杂的任务,需要理解和生成具有长期依赖性和和声约束的信息的能力。虽然在STEM科目中表现出令人印象深刻的能力,但目前的LLM很容易在这项任务中失败,即使配备了上下文学习和思想链等现代技术,也会产生写得很差的音乐。为了进一步探索和提高LLM的音乐创作潜力,利用他们的推理能力和音乐历史和理论的大型知识库,我们提出了ComposerX,一个基于代理的符号音乐生成框架。我们发现,应用多代理的方法显着提高了GPT-4的音乐创作质量。结果表明,ComposerX能够制作具有迷人旋律的连贯复调音乐作品,同时遵守用户指令。摘要:Music composition represents the creative side of humanity, and itself is a complex task that requires abilities to understand and generate information with long dependency and harmony constraints. While demonstrating impressive capabilities in STEM subjects, current LLMs easily fail in this task, generating ill-written music even when equipped with modern techniques like In-Context-Learning and Chain-of-Thoughts. To further explore and enhance LLMs' potential in music composition by leveraging their reasoning ability and the large knowledge base in music history and theory, we propose ComposerX, an agent-based symbolic music generation framework. We find that applying a multi-agent approach significantly improves the music composition quality of GPT-4. The results demonstrate that ComposerX is capable of producing coherent polyphonic music compositions with captivating melodies, while adhering to user instructions.
【6】 Towards Privacy-Preserving Audio Classification Systems作者:Bhawana Chhaglani,Jeremy Gummeson,Prashant Shenoy摘要:音频信号可以揭示一个人生活的私密细节,包括他们的谈话、健康状况、情绪、位置和个人偏好。未经授权访问或滥用这些信息可能会产生深刻的个人和社会影响。在一个能够录音的设备越来越多的时代,保护用户隐私是一项关键义务。这项工作研究了当前音频分类系统中的道德和隐私问题。我们讨论了设计隐私保护音频传感系统的挑战和研究方向。我们提出了隐私保护的音频功能,可用于分类范围广泛的音频类,而隐私保护。摘要:Audio signals can reveal intimate details about a person's life, including their conversations, health status, emotions, location, and personal preferences. Unauthorized access or misuse of this information can have profound personal and social implications. In an era increasingly populated by devices capable of audio recording, safeguarding user privacy is a critical obligation. This work studies the ethical and privacy concerns in current audio classification systems. We discuss the challenges and research directions in designing privacy-preserving audio sensing systems. We propose privacy-preserving audio features that can be used to classify wide range of audio classes, while being privacy preserving.
【7】 TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality标题:TI-ASO:通过针对缺失语音情态的文本到语音插补实现稳健的自动语音理解作者:Tiantian Feng,Xuan Shi,Rahul Gupta,Shrikanth S. Narayanan摘要:自动语音理解(ASU)的目标是像人类一样的语音解释,从语音和语音中传达的语言(文本)内容中提供细致入微的意图,情感,情感和内容理解。通常,训练鲁棒ASU模型在很大程度上依赖于获取大规模、高质量的语音和相关的传输。然而,由于隐私等问题,收集或使用语音数据来训练ASU通常具有挑战性。为了在语音(音频)模态缺失时实现ASU的设置,我们提出了TI-ASU,使用预先训练的文本到语音模型来估算缺失的语音。我们报告了大量的实验评估TI-ASU的各种缺失的尺度,多模态和单模态设置,并使用LLM。我们的研究结果表明,TI-ASU在甚至高达95%的训练语音缺失的情况下,对提高ASU产生了实质性的好处。此外,我们表明,TI-ASU是自适应的辍学训练,提高模型的鲁棒性,在解决丢失的语音在推理过程中。摘要:Automatic Speech Understanding (ASU) aims at human-like speech interpretation, providing nuanced intent, emotion, sentiment, and content understanding from speech and language (text) content conveyed in speech. Typically, training a robust ASU model relies heavily on acquiring large-scale, high-quality speech and associated transcriptions. However, it is often challenging to collect or use speech data for training ASU due to concerns such as privacy. To approach this setting of enabling ASU when speech (audio) modality is missing, we propose TI-ASU, using a pre-trained text-to-speech model to impute the missing speech. We report extensive experiments evaluating TI-ASU on various missing scales, both multi- and single-modality settings, and the use of LLMs. Our findings show that TI-ASU yields substantial benefits to improve ASU in scenarios where even up to 95% of training speech is missing. Moreover, we show that TI-ASU is adaptive to dropout training, improving model robustness in addressing missing speech during inference.【8】 Usefulness of Emotional Prosody in Neural Machine Translation作者:Charles Brazier,Jean-Luc Rouas备注:5 pages, In Proceedings of the 11th International Conference on Speech Prosody (SP), Leiden, The Netherlands, 2024摘要:神经机器翻译(NMT)是使用经过训练的神经网络将文本从一种语言翻译成另一种语言的任务。一些现有的工作旨在将外部信息纳入NMT模型,以改善或控制预测的翻译(例如情感,礼貌,性别)。在这项工作中,我们建议通过添加另一个外部信息源来提高翻译质量:语音中自动识别的情感。这项工作的动机是假设每种情绪都与一个特定的词汇,可以重叠的情绪。我们提出的方法遵循两个阶段的程序。首先,我们选择一个最先进的语音情感识别(SER)模型来预测维度情感值从数据集中的所有输入音频。然后,我们使用这些预测的情感作为源标记添加在输入文本的开头训练我们的NMT模型。我们表明,将情感信息,特别是唤醒,到NMT系统,导致更好的翻译。摘要:Neural Machine Translation (NMT) is the task of translating a text from one language to another with the use of a trained neural network. Several existing works aim at incorporating external information into NMT models to improve or control predicted translations (e.g. sentiment, politeness, gender). In this work, we propose to improve translation quality by adding another external source of information: the automatically recognized emotion in the voice. This work is motivated by the assumption that each emotion is associated with a specific lexicon that can overlap between emotions. Our proposed method follows a two-stage procedure. At first, we select a state-of-the-art Speech Emotion Recognition (SER) model to predict dimensional emotion values from all input audio in the dataset. Then, we use these predicted emotions as source tokens added at the beginning of input texts to train our NMT model. We show that integrating emotion information, especially arousal, into NMT systems leads to better translations.【9】 An automatic mixing speech enhancement system for multi-track audio作者:Xiaojing Liu,Angeliki Mourgela,Hongwei Ai,Joshua D. Reiss摘要:提出了一种多声道语音增强系统。该系统将最大限度地减少听觉掩蔽,同时允许一个人听到多个同时发言者。该系统可以用于多种通信场景,电话会议、发票游戏和直播流媒体。ITU—R BS. 1387音频质量感知评估(PEAQ)模型用于评估音频信号中的掩蔽量。不同的音频效果,例如,经由旨在最小化掩蔽的迭代和声搜索算法来应用电平平衡、均衡、动态范围压缩和空间化。在主观听觉测试中,所设计的系统可以与专业音响工程师的混音相媲美,并且优于现有的自动混音系统。摘要:We propose a speech enhancement system for multitrack audio. The system will minimize auditory masking while allowing one to hear multiple simultaneous speakers. The system can be used in multiple communication scenarios e.g., teleconferencing, invoice gaming, and live streaming. The ITU-R BS.1387 Perceptual Evaluation of Audio Quality (PEAQ) model is used to evaluate the amount of masking in the audio signals. Different audio effects e.g., level balance, equalization, dynamic range compression, and spatialization are applied via an iterative Harmony searching algorithm that aims to minimize the masking. In the subjective listening test, the designed system can compete with mixes by professional sound engineers and outperforms mixes by existing auto-mixing systems.
【10】 T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining作者:Yi Yuan,Zhuo Chen,Xubo Liu,Haohe Liu,Xuenan Xu,Dongya Jia,Yuanzhe Chen,Mark D. Plumbley,Wenwu Wang备注:Preprint submitted to IEEE MLSP 2024摘要:对比语言—音频预训练(CLAP)是一种将音频和语言的表征结合起来的方法,在检索和分类任务中取得了显著的效果。然而,目前的CLAP的斗争,以捕捉音频和文本功能的时间信息,提出了大量的限制,如音频检索和生成的任务。为了解决这个问题,我们引入了T—CLAP,一个时间增强的CLAP模型。我们使用大语言模型(LLM)和混合策略,从大量的音频文本数据集生成音频片段的时间对比字幕。随后,一个新的时间聚焦对比损失的设计,通过将这些合成数据的CLAP模型进行微调。我们在多个下游任务中进行全面的实验和分析。T—CLAP在捕获声音事件的时间关系方面表现出更好的能力,并且显著优于最先进的模型。摘要:Contrastive language-audio pretraining~(CLAP) has been developed to align the representations of audio and language, achieving remarkable performance in retrieval and classification tasks. However, current CLAP struggles to capture temporal information within audio and text features, presenting substantial limitations for tasks such as audio retrieval and generation. To address this gap, we introduce T-CLAP, a temporal-enhanced CLAP model. We use Large Language Models~(LLMs) and mixed-up strategies to generate temporal-contrastive captions for audio clips from extensive audio-text datasets. Subsequently, a new temporal-focused contrastive loss is designed to fine-tune the CLAP model by incorporating these synthetic data. We conduct comprehensive experiments and analysis in multiple downstream tasks. T-CLAP shows improved capability in capturing the temporal relationship of sound events and outperforms state-of-the-art models by a significant margin.【11】 An RFP dataset for Real, Fake, and Partially fake audio detection标题:用于真实、虚假和部分虚假音频检测的RFP数据集作者:Abdulazeez AlAli,George Theodorakopoulos摘要:深度学习的最新进展使人们能够创建听起来自然的合成语音。然而,攻击者也利用这些技术进行网络钓鱼等攻击。已经创建了许多公共数据集,以促进有效检测模型的开发。然而,可用的数据集只包含完全虚假的音频;因此,检测模型可能会错过用虚假音频替换一小部分真实音频的攻击。为了解决这个问题,本文提出了RFP数据集,它包括五种不同的音频类型:部分假音频(PF),带噪声的音频,语音转换(VC),文本到语音(TTS)和真实的。然后使用这些数据来评估几种检测模型,揭示了可用的检测模型在检测PF音频而不是完全虚假的音频时会产生明显更高的等错误率(EER)。最低EER为25.42%。因此,我们认为检测模型的创建者必须认真考虑使用像RFP这样的数据集,其中包括PF和其他类型的假音频。摘要:Recent advances in deep learning have enabled the creation of natural-sounding synthesised speech. However, attackers have also utilised these tech-nologies to conduct attacks such as phishing. Numerous public datasets have been created to facilitate the development of effective detection models. How-ever, available datasets contain only entirely fake audio; therefore, detection models may miss attacks that replace a short section of the real audio with fake audio. In recognition of this problem, the current paper presents the RFP da-taset, which comprises five distinct audio types: partial fake (PF), audio with noise, voice conversion (VC), text-to-speech (TTS), and real. The data are then used to evaluate several detection models, revealing that the available detec-tion models incur a markedly higher equal error rate (EER) when detecting PF audio instead of entirely fake audio. The lowest EER recorded was 25.42%. Therefore, we believe that creators of detection models must seriously consid-er using datasets like RFP that include PF and other types of fake audio.【12】 Synthesizing Audio from Silent Video using Sequence to Sequence Modeling作者:Hugo Garrido-Lestache Belinchon,Helina Mulugeta,Adam Haile摘要:从视频的视觉环境中生成音频在改善我们与视听媒体的交互方式方面具有多种实际应用—例如,增强CCTV镜头分析,恢复历史视频(例如,无声电影),并改进视频生成模型。我们提出了一种使用序列到序列模型从视频生成音频的新方法,改进了使用CNN和WaveNet的先前工作,并面临声音多样性和泛化挑战。我们的方法采用3D矢量量化变分自动编码器(VQ—VAE)来捕获视频的空间和时间结构,使用自定义音频解码器进行解码,以获得更广泛的声音。我们的模型在Youtube8M数据集上进行了训练,专注于特定领域,旨在增强CCTV镜头分析,无声电影恢复和视频生成模型等应用程序。摘要:Generating audio from a video's visual context has multiple practical applications in improving how we interact with audio-visual media - for example, enhancing CCTV footage analysis, restoring historical videos (e.g., silent movies), and improving video generation models. We propose a novel method to generate audio from video using a sequence-to-sequence model, improving on prior work that used CNNs and WaveNet and faced sound diversity and generalization challenges. Our approach employs a 3D Vector Quantized Variational Autoencoder (VQ-VAE) to capture the video's spatial and temporal structures, decoding with a custom audio decoder for a broader range of sounds. Trained on the Youtube8M dataset segment, focusing on specific domains, our model aims to enhance applications like CCTV footage analysis, silent movie restoration, and video generation models.
【1】 Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention作者:Ruijie Tao,Xinyuan Qian,Yidi Jiang,Junjie Li,Jiadong Wang,Haizhou Li摘要:视听目标说话人提取(AV-TSE)的目的是在给定辅助视觉线索的情况下,从混合音频中提取特定人的语音。以前的方法通常通过语音唇同步来搜索目标语音。然而,这种策略主要关注目标语音的存在,而忽略了噪声特性的变化。这可能导致在具有挑战性的声学情况下从不正确的声源中提取噪声信号。为此,我们提出了一种新的反向选择性听觉注意机制,它可以抑制干扰说话人和非语音信号,以避免不正确的说话人提取。通过这种机制估计和利用不需要的噪声信号,我们设计了一个AV-TSE框架称为减法和提取网络(SEANet)来抑制噪声信号。我们通过重新实现三种流行的AV-TSE方法作为基线并涉及九个评估指标进行了大量的实验。实验结果表明,我们提出的SEANet达到了最先进的结果,并表现良好的所有五个数据集。我们会公布代码,模型和数据日志。摘要:Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization. However, this strategy mainly focuses on the existence of target speech, while ignoring the variations of the noise characteristics. That may result in extracting noisy signals from the incorrect sound source in challenging acoustic situations. To this end, we propose a novel reverse selective auditory attention mechanism, which can suppress interference speakers and non-speech signals to avoid incorrect speaker extraction. By estimating and utilizing the undesired noisy signal through this mechanism, we design an AV-TSE framework named Subtraction-and-ExtrAction network (SEANet) to suppress the noisy signals. We conduct abundant experiments by re-implementing three popular AV-TSE methods as the baselines and involving nine metrics for evaluation. The experimental results show that our proposed SEANet achieves state-of-the-art results and performs well for all five datasets. We will release the codes, the models and data logs.【2】 A Comparison of Differential Performance Metrics for the Evaluation of Automatic Speaker Verification Fairness作者:Oubaida Chouchane,Christoph Busch,Chiara Galdi,Nicholas Evans,Massimiliano Todisco摘要:在作出决定和通过自动化程序处理个人数据时,人们期望公平-不同人口群体的成员得到公平对待。这种期望适用于生物识别系统,如自动说话人验证(ASV)。 我们提出了三个候选人的公平性指标的比较,并扩展以前的工作进行人脸识别,通过检查在一系列不同的ASV操作点的差异性能。结果表明,基尼聚合率的生物统计公平性(GARBE)是唯一一个满足三个功能的公平性衡量标准。 此外,还提出了五个国家的最先进的ASV系统的公平性和验证性能的综合评价。我们的研究结果揭示了公平性和验证准确性之间的微妙权衡,强调了系统设计,人口包容性和验证可靠性之间的复杂相互作用。摘要:When decisions are made and when personal data is treated by automated processes, there is an expectation of fairness -- that members of different demographic groups receive equitable treatment. This expectation applies to biometric systems such as automatic speaker verification (ASV). We present a comparison of three candidate fairness metrics and extend previous work performed for face recognition, by examining differential performance across a range of different ASV operating points. Results show that the Gini Aggregation Rate for Biometric Equitability (GARBE) is the only one which meets three functional fairness measure criteria. Furthermore, a comprehensive evaluation of the fairness and verification performance of five state-of-the-art ASV systems is also presented. Our findings reveal a nuanced trade-off between fairness and verification accuracy underscoring the complex interplay between system design, demographic inclusiveness, and verification reliability.【3】 Certification of Speaker Recognition Models to Additive Perturbations作者:Dmitrii Korzh,Elvir Karimov,Mikhail Pautov,Oleg Y. Rogov,Ivan Oseledets摘要:说话人识别技术应用于从个人虚拟助理到安全访问系统的各种任务中。然而,这些系统对对抗性攻击的鲁棒性,特别是对加性扰动的鲁棒性,仍然是一个重大挑战。在本文中,我们率先将鲁棒性认证技术应用于说话人识别,最初是为图像域开发的。在我们的工作中,我们通过转移和改进随机平滑认证技术对范数有界的附加扰动分类和Few-Shot学习任务的说话人识别来弥补这一差距。我们证明了这些方法在VoxCeleb 1和2数据集上的有效性。我们希望这项工作能够提高语音生物特征的鲁棒性,建立一个新的认证基准,并加快在音频领域的认证方法的研究。摘要:Speaker recognition technology is applied in various tasks ranging from personal virtual assistants to secure access systems. However, the robustness of these systems against adversarial attacks, particularly to additive perturbations, remains a significant challenge. In this paper, we pioneer applying robustness certification techniques to speaker recognition, originally developed for the image domain. In our work, we cover this gap by transferring and improving randomized smoothing certification techniques against norm-bounded additive perturbations for classification and few-shot learning tasks to speaker recognition. We demonstrate the effectiveness of these methods on VoxCeleb 1 and 2 datasets for several models. We expect this work to improve voice-biometry robustness, establish a new certification benchmark, and accelerate research of certification methods in the audio domain.
【4】 A Systematic Evaluation of Adversarial Attacks against Speech Emotion Recognition Models作者:Nicolas Facchinetti,Federico Simonetta,Stavros Ntalampiras摘要:近年来,语音情感识别(SER)由于其在不同领域的潜在应用以及深度学习技术提供的可能性而不断受到关注。然而,最近的研究表明,深度学习模型很容易受到对抗性攻击。在本文中,我们通过检查SER背景下各种对抗性白盒和黑盒攻击对不同语言和性别的影响,系统地评估了这个问题。我们首先提出了一种适合音频数据处理,特征提取和CNN-LSTM架构的方法。观察到的结果强调了CNN-LSTM模型对对抗性示例(AE)的显著脆弱性。事实上,所有考虑的对抗性攻击都能够显着降低构建模型的性能。此外,在评估攻击的效果时,注意到所分析的语言之间以及男性和女性语言之间的微小差异。总之,这项工作有助于理解CNN-LSTM模型的鲁棒性,特别是在SER场景中,以及AE的影响。有趣的是,我们的研究结果可以作为a)开发更强大的SER算法,b)设计更有效的攻击,c)调查可能的防御措施,d)提高对不同语言和性别之间声音差异的理解,e)总体而言,增强我们对SER任务的理解。摘要:Speech emotion recognition (SER) is constantly gaining attention in recent years due to its potential applications in diverse fields and thanks to the possibility offered by deep learning technologies. However, recent studies have shown that deep learning models can be vulnerable to adversarial attacks. In this paper, we systematically assess this problem by examining the impact of various adversarial white-box and black-box attacks on different languages and genders within the context of SER. We first propose a suitable methodology for audio data processing, feature extraction, and CNN-LSTM architecture. The observed outcomes highlighted the significant vulnerability of CNN-LSTM models to adversarial examples (AEs). In fact, all the considered adversarial attacks are able to significantly reduce the performance of the constructed models. Furthermore, when assessing the efficacy of the attacks, minor differences were noted between the languages analyzed as well as between male and female speech. In summary, this work contributes to the understanding of the robustness of CNN-LSTM models, particularly in SER scenarios, and the impact of AEs. Interestingly, our findings serve as a baseline for a) developing more robust algorithms for SER, b) designing more effective attacks, c) investigating possible defenses, d) improved understanding of the vocal differences between different languages and genders, and e) overall, enhancing our comprehension of the SER task.
【5】 Pièces de viole des Cinq Livres and their statistical signatures: the musical work of Marin Marais and Jordi Savall标题:Pièces de viole des Cinq Livres及其统计签名:Marin Marais和Jordi Savall的音乐作品作者:Igor Lugo,Martha G. Alatriste-Contreras摘要:本研究基于Marin Marais和Jordi Savall之间的合作工作,分析了与"Pieces de viole des Cinq Livres"相关的音频信号的频谱,以获取潜在的音乐信息。特别是,我们探讨可能的统计签名相关的音乐作品的识别。基于复杂系统方法,我们计算音频信号的频谱,分析和识别它们的最佳拟合统计分布,并使用科学音高表示法绘制它们的相对频率。研究结果表明,收集的频率分量相关的频谱的每一本书,形成这个音频工作显示高度倾斜和相关的统计分布。因此,最好地描述这些音频数据的集合并且可以与奇异统计签名相关联的最频繁的统计分布是指数分布。摘要:This study analyzes the spectrum of audio signals related to the work of "Pi\`eces de viole des Cinq Livres" based on the collaborative work between Marin Marais and Jordi Savall for the underlying musical information. In particular, we explore the identification of possible statistical signatures related to this musical work. Based on the complex systems approach, we compute the spectrum of audio signals, analyze and identify their best-fit statistical distributions, and plot their relative frequencies using the scientific pitch notation. Findings suggest that the collection of frequency components related to the spectrum of each of the books that form this audio work show highly skewed and associated statistical distributions. Therefore, the most frequent statistical distribution that best describes the collection of these audio data and may be associated with a singular statistical signature is the exponential.【6】 ComposerX: Multi-Agent Symbolic Music Composition with LLMs标题:ComposerX:使用LLM的多智能体符号音乐创作作者:Qixin Deng,Qikai Yang,Ruibin Yuan,Yipeng Huang,Yi Wang,Xubo Liu,Zeyue Tian,Jiahao Pan,Ge Zhang,Hanfeng Lin,Yizhi Li,Yinghao Ma,Jie Fu,Chenghua Lin,Emmanouil Benetos,Wenwu Wang,Guangyu Xia,Wei Xue,Yike Guo摘要:音乐创作代表了人类创造性的一面,它本身是一项复杂的任务,需要理解和生成具有长期依赖性和和声约束的信息的能力。虽然在STEM科目中表现出令人印象深刻的能力,但目前的LLM很容易在这项任务中失败,即使配备了上下文学习和思想链等现代技术,也会产生写得很差的音乐。为了进一步探索和提高LLM的音乐创作潜力,利用他们的推理能力和音乐历史和理论的大型知识库,我们提出了ComposerX,一个基于代理的符号音乐生成框架。我们发现,应用多代理的方法显着提高了GPT-4的音乐创作质量。结果表明,ComposerX能够制作具有迷人旋律的连贯复调音乐作品,同时遵守用户指令。摘要:Music composition represents the creative side of humanity, and itself is a complex task that requires abilities to understand and generate information with long dependency and harmony constraints. While demonstrating impressive capabilities in STEM subjects, current LLMs easily fail in this task, generating ill-written music even when equipped with modern techniques like In-Context-Learning and Chain-of-Thoughts. To further explore and enhance LLMs' potential in music composition by leveraging their reasoning ability and the large knowledge base in music history and theory, we propose ComposerX, an agent-based symbolic music generation framework. We find that applying a multi-agent approach significantly improves the music composition quality of GPT-4. The results demonstrate that ComposerX is capable of producing coherent polyphonic music compositions with captivating melodies, while adhering to user instructions.
【7】 Towards Privacy-Preserving Audio Classification Systems作者:Bhawana Chhaglani,Jeremy Gummeson,Prashant Shenoy摘要:音频信号可以揭示一个人生活的私密细节,包括他们的谈话、健康状况、情绪、位置和个人偏好。未经授权访问或滥用这些信息可能会产生深刻的个人和社会影响。在一个能够录音的设备越来越多的时代,保护用户隐私是一项关键义务。这项工作研究了当前音频分类系统中的道德和隐私问题。我们讨论了设计隐私保护音频传感系统的挑战和研究方向。我们提出了隐私保护的音频功能,可用于分类范围广泛的音频类,而隐私保护。摘要:Audio signals can reveal intimate details about a person's life, including their conversations, health status, emotions, location, and personal preferences. Unauthorized access or misuse of this information can have profound personal and social implications. In an era increasingly populated by devices capable of audio recording, safeguarding user privacy is a critical obligation. This work studies the ethical and privacy concerns in current audio classification systems. We discuss the challenges and research directions in designing privacy-preserving audio sensing systems. We propose privacy-preserving audio features that can be used to classify wide range of audio classes, while being privacy preserving.
【8】 TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality标题:TI-ASO:通过针对缺失语音情态的文本到语音插补实现稳健的自动语音理解作者:Tiantian Feng,Xuan Shi,Rahul Gupta,Shrikanth S. Narayanan摘要:自动语音理解(ASU)的目标是像人类一样的语音解释,从语音和语音中传达的语言(文本)内容中提供细致入微的意图,情感,情感和内容理解。通常,训练鲁棒ASU模型在很大程度上依赖于获取大规模、高质量的语音和相关的传输。然而,由于隐私等问题,收集或使用语音数据来训练ASU通常具有挑战性。为了在语音(音频)模态缺失时实现ASU的设置,我们提出了TI-ASU,使用预先训练的文本到语音模型来估算缺失的语音。我们报告了大量的实验评估TI-ASU的各种缺失的尺度,多模态和单模态设置,并使用LLM。我们的研究结果表明,TI-ASU在甚至高达95%的训练语音缺失的情况下,对提高ASU产生了实质性的好处。此外,我们表明,TI-ASU是自适应的辍学训练,提高模型的鲁棒性,在解决丢失的语音在推理过程中。摘要:Automatic Speech Understanding (ASU) aims at human-like speech interpretation, providing nuanced intent, emotion, sentiment, and content understanding from speech and language (text) content conveyed in speech. Typically, training a robust ASU model relies heavily on acquiring large-scale, high-quality speech and associated transcriptions. However, it is often challenging to collect or use speech data for training ASU due to concerns such as privacy. To approach this setting of enabling ASU when speech (audio) modality is missing, we propose TI-ASU, using a pre-trained text-to-speech model to impute the missing speech. We report extensive experiments evaluating TI-ASU on various missing scales, both multi- and single-modality settings, and the use of LLMs. Our findings show that TI-ASU yields substantial benefits to improve ASU in scenarios where even up to 95% of training speech is missing. Moreover, we show that TI-ASU is adaptive to dropout training, improving model robustness in addressing missing speech during inference.
【9】 Usefulness of Emotional Prosody in Neural Machine Translation作者:Charles Brazier,Jean-Luc Rouas备注:5 pages, In Proceedings of the 11th International Conference on Speech Prosody (SP), Leiden, The Netherlands, 2024摘要:神经机器翻译(NMT)是使用经过训练的神经网络将文本从一种语言翻译成另一种语言的任务。一些现有的工作旨在将外部信息纳入NMT模型,以改善或控制预测的翻译(例如情感,礼貌,性别)。在这项工作中,我们建议通过添加另一个外部信息源来提高翻译质量:语音中自动识别的情感。这项工作的动机是假设每种情绪都与一个特定的词汇,可以重叠的情绪。我们提出的方法遵循两个阶段的程序。首先,我们选择一个最先进的语音情感识别(SER)模型来预测维度情感值从数据集中的所有输入音频。然后,我们使用这些预测的情感作为源标记添加在输入文本的开头训练我们的NMT模型。我们表明,将情感信息,特别是唤醒,到NMT系统,导致更好的翻译。摘要:Neural Machine Translation (NMT) is the task of translating a text from one language to another with the use of a trained neural network. Several existing works aim at incorporating external information into NMT models to improve or control predicted translations (e.g. sentiment, politeness, gender). In this work, we propose to improve translation quality by adding another external source of information: the automatically recognized emotion in the voice. This work is motivated by the assumption that each emotion is associated with a specific lexicon that can overlap between emotions. Our proposed method follows a two-stage procedure. At first, we select a state-of-the-art Speech Emotion Recognition (SER) model to predict dimensional emotion values from all input audio in the dataset. Then, we use these predicted emotions as source tokens added at the beginning of input texts to train our NMT model. We show that integrating emotion information, especially arousal, into NMT systems leads to better translations.
【10】 An automatic mixing speech enhancement system for multi-track audio作者:Xiaojing Liu,Angeliki Mourgela,Hongwei Ai,Joshua D. Reiss摘要:提出了一种多声道语音增强系统。该系统将最大限度地减少听觉掩蔽,同时允许一个人听到多个同时发言者。该系统可以用于多种通信场景,电话会议、发票游戏和直播流媒体。ITU-R BS. 1387音频质量感知评估(PEAQ)模型用于评估音频信号中的掩蔽量。不同的音频效果,例如,经由旨在最小化掩蔽的迭代和声搜索算法来应用电平平衡、均衡、动态范围压缩和空间化。在主观听觉测试中,所设计的系统可以与专业音响工程师的混音相媲美,并且优于现有的自动混音系统。摘要:We propose a speech enhancement system for multitrack audio. The system will minimize auditory masking while allowing one to hear multiple simultaneous speakers. The system can be used in multiple communication scenarios e.g., teleconferencing, invoice gaming, and live streaming. The ITU-R BS.1387 Perceptual Evaluation of Audio Quality (PEAQ) model is used to evaluate the amount of masking in the audio signals. Different audio effects e.g., level balance, equalization, dynamic range compression, and spatialization are applied via an iterative Harmony searching algorithm that aims to minimize the masking. In the subjective listening test, the designed system can compete with mixes by professional sound engineers and outperforms mixes by existing auto-mixing systems.
【11】 T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining作者:Yi Yuan,Zhuo Chen,Xubo Liu,Haohe Liu,Xuenan Xu,Dongya Jia,Yuanzhe Chen,Mark D. Plumbley,Wenwu Wang备注:Preprint submitted to IEEE MLSP 2024摘要:对比语言-音频预训练(CLAP)是一种将音频和语言的表征结合起来的方法,在检索和分类任务中取得了显著的效果。然而,目前的CLAP的斗争,以捕捉音频和文本功能的时间信息,提出了大量的限制,如音频检索和生成的任务。为了解决这个问题,我们引入了T-CLAP,一个时间增强的CLAP模型。我们使用大语言模型(LLM)和混合策略,从大量的音频文本数据集生成音频片段的时间对比字幕。随后,一个新的时间聚焦对比损失的设计,通过将这些合成数据的CLAP模型进行微调。我们在多个下游任务中进行全面的实验和分析。T-CLAP在捕获声音事件的时间关系方面表现出更好的能力,并且显著优于最先进的模型。摘要:Contrastive language-audio pretraining~(CLAP) has been developed to align the representations of audio and language, achieving remarkable performance in retrieval and classification tasks. However, current CLAP struggles to capture temporal information within audio and text features, presenting substantial limitations for tasks such as audio retrieval and generation. To address this gap, we introduce T-CLAP, a temporal-enhanced CLAP model. We use Large Language Models~(LLMs) and mixed-up strategies to generate temporal-contrastive captions for audio clips from extensive audio-text datasets. Subsequently, a new temporal-focused contrastive loss is designed to fine-tune the CLAP model by incorporating these synthetic data. We conduct comprehensive experiments and analysis in multiple downstream tasks. T-CLAP shows improved capability in capturing the temporal relationship of sound events and outperforms state-of-the-art models by a significant margin.
【12】 An RFP dataset for Real, Fake, and Partially fake audio detection标题:用于真实、虚假和部分虚假音频检测的RFP数据集作者:Abdulazeez AlAli,George Theodorakopoulos摘要:深度学习的最新进展使人们能够创建听起来自然的合成语音。然而,攻击者也利用这些技术进行网络钓鱼等攻击。已经创建了许多公共数据集,以促进有效检测模型的开发。然而,可用的数据集只包含完全虚假的音频;因此,检测模型可能会错过用虚假音频替换一小部分真实音频的攻击。为了解决这个问题,本文提出了RFP数据集,它包括五种不同的音频类型:部分假音频(PF),带噪声的音频,语音转换(VC),文本到语音(TTS)和真实的。然后使用这些数据来评估几种检测模型,揭示了可用的检测模型在检测PF音频而不是完全虚假的音频时会产生明显更高的等错误率(EER)。最低EER为25.42%。因此,我们认为检测模型的创建者必须认真考虑使用像RFP这样的数据集,其中包括PF和其他类型的假音频。摘要:Recent advances in deep learning have enabled the creation of natural-sounding synthesised speech. However, attackers have also utilised these tech-nologies to conduct attacks such as phishing. Numerous public datasets have been created to facilitate the development of effective detection models. How-ever, available datasets contain only entirely fake audio; therefore, detection models may miss attacks that replace a short section of the real audio with fake audio. In recognition of this problem, the current paper presents the RFP da-taset, which comprises five distinct audio types: partial fake (PF), audio with noise, voice conversion (VC), text-to-speech (TTS), and real. The data are then used to evaluate several detection models, revealing that the available detec-tion models incur a markedly higher equal error rate (EER) when detecting PF audio instead of entirely fake audio. The lowest EER recorded was 25.42%. Therefore, we believe that creators of detection models must seriously consid-er using datasets like RFP that include PF and other types of fake audio.
【13】 Synthesizing Audio from Silent Video using Sequence to Sequence Modeling作者:Hugo Garrido-Lestache Belinchon,Helina Mulugeta,Adam Haile摘要:从视频的视觉环境中生成音频在改善我们与视听媒体的交互方式方面具有多种实际应用-例如,增强CCTV镜头分析,恢复历史视频(例如,无声电影),并改进视频生成模型。我们提出了一种使用序列到序列模型从视频生成音频的新方法,改进了使用CNN和WaveNet的先前工作,并面临声音多样性和泛化挑战。我们的方法采用3D矢量量化变分自动编码器(VQ-VAE)来捕获视频的空间和时间结构,使用自定义音频解码器进行解码,以获得更广泛的声音。我们的模型在Youtube 8 M数据集上进行了训练,专注于特定领域,旨在增强CCTV镜头分析,无声电影恢复和视频生成模型等应用程序。摘要:Generating audio from a video's visual context has multiple practical applications in improving how we interact with audio-visual media - for example, enhancing CCTV footage analysis, restoring historical videos (e.g., silent movies), and improving video generation models. We propose a novel method to generate audio from video using a sequence-to-sequence model, improving on prior work that used CNNs and WaveNet and faced sound diversity and generalization challenges. Our approach employs a 3D Vector Quantized Variational Autoencoder (VQ-VAE) to capture the video's spatial and temporal structures, decoding with a custom audio decoder for a broader range of sounds. Trained on the Youtube8M dataset segment, focusing on specific domains, our model aims to enhance applications like CCTV footage analysis, silent movie restoration, and video generation models.