今日论文合集:cs.SD语音10篇,eess.AS音频处理10篇。本文经arXiv每日学术速递授权转载
【1】Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning
标题:研究通用音频表示学习的联合嵌入预测架构中的设计选择作者:Alain Riou,Stefan Lattner,Gaëtan Hadjeres,Geoffroy Peeters备注:Self-supervision in Audio, Speech and Beyond workshop, IEEE International Conference on Acoustics, Speech, and Signal Processing, 2024摘要:本文讨论了自监督通用音频表示学习问题。我们探索使用联合嵌入预测架构(JEPA)来完成这项任务,该任务包括将输入mel频谱图分为两部分(上下文和目标),计算每个部分的神经表示,并训练神经网络从上下文表示中预测目标表示。我们研究了这个框架内的几个设计选择,并通过广泛的实验研究它们的影响,通过评估我们的模型对各种音频分类基准,包括环境声音,语音和音乐下游任务。我们特别关注输入数据的哪一部分被用作上下文或目标,并通过实验表明它会显著影响模型的质量。特别是,我们注意到,在图像域中的一些有效的设计选择导致音频性能不佳,从而突出了这两种模式之间的主要差异。摘要:This paper addresses the problem of self-supervised general-purpose audio representation learning. We explore the use of Joint-Embedding Predictive Architectures (JEPA) for this task, which consists of splitting an input mel-spectrogram into two parts (context and target), computing neural representations for each, and training the neural network to predict the target representations from the context representations. We investigate several design choices within this framework and study their influence through extensive experiments by evaluating our models on various audio classification benchmarks, including environmental sounds, speech and music downstream tasks. We focus notably on which part of the input data is used as context or target and show experimentally that it significantly impacts the model's quality. In particular, we notice that some effective design choices in the image domain lead to poor performance on audio, thus highlighting major differences between these two modalities.【2】 EVDA: Evolving Deepfake Audio Detection Continual Learning Benchmark标题:EVDA:不断发展的Deepfake音频检测持续学习基准作者:Xiaohui Zhang,Jiangyan Yi,Jianhua Tao摘要:GPT-4、GPT-4 o和Claude家族等高级大型语言模型的兴起,使得虚假音频检测变得越来越具有挑战性。传统的微调方法很难跟上合成语音不断发展的步伐,需要不断学习的方法来适应新的音频,同时保留检测旧类型的能力。持续学习是检测新出现的deepfake音频的有效工具,同时保持旧类型的性能,缺乏一个结构良好且用户友好的评估框架。为了解决这一差距,我们引入了EVDA,这是一个用于评估deepfake音频检测中持续学习方法的基准。EVDA包括来自反欺骗语音系列、中文虚假音频检测系列的经典数据集,以及来自GPT-4和GPT-4 o等模型的新生成的deepfake音频。它支持各种持续学习技术,如弹性权重合并(EWC),学习不遗忘(LwF),以及最近的方法,如正则化自适应权重修改(RAWM)和弧度权重修改(RWM)。此外,EVDA通过提供用于集成新的持续学习方法的开放接口来促进鲁棒算法的开发摘要:The rise of advanced large language models such as GPT-4, GPT-4o, and the Claude family has made fake audio detection increasingly challenging. Traditional fine-tuning methods struggle to keep pace with the evolving landscape of synthetic speech, necessitating continual learning approaches that can adapt to new audio while retaining the ability to detect older types. Continual learning, which acts as an effective tool for detecting newly emerged deepfake audio while maintaining performance on older types, lacks a well-constructed and user-friendly evaluation framework. To address this gap, we introduce EVDA, a benchmark for evaluating continual learning methods in deepfake audio detection. EVDA includes classic datasets from the Anti-Spoofing Voice series, Chinese fake audio detection series, and newly generated deepfake audio from models like GPT-4 and GPT-4o. It supports various continual learning techniques, such as Elastic Weight Consolidation (EWC), Learning without Forgetting (LwF), and recent methods like Regularized Adaptive Weight Modification (RAWM) and Radian Weight Modification (RWM). Additionally, EVDA facilitates the development of robust algorithms by providing an open interface for integrating new continual learning methods
【3】 Abnormal Respiratory Sound Identification Using Audio-Spectrogram Vision Transformer标题:使用音频频谱图视觉Transformer识别异常呼吸音作者:Whenty Ariyanti,Kai-Chun Liu,Kuan-Yu Chen,Yu Tsao摘要:呼吸道疾病是全球第三大死亡原因,被认为是一种高度优先的疾病,需要对识别和治疗进行大量研究。听诊器记录的肺部声音和人工智能驱动的设备已被用于识别肺部疾病,并帮助专家做出准确的诊断。在这项研究中,音频频谱图Vision Transformer(AS—ViT),一种新的方法来识别异常呼吸音,开发。肺部的声音被转换成视觉表示称为频谱图使用一种技术称为短时傅立叶变换(STFT)。然后使用称为Vision Transformer的模型分析这些图像,以识别不同类型的呼吸声。该分类使用ICBHI 2017数据库进行,该数据库包括具有不同频率,噪声水平和背景的各种类型的肺音。提出的AS—ViT方法使用三个指标进行评估,在呼吸音检测的未加权平均召回率和总得分方面,60:40分割比分别达到79.1%和59.8%,80:20分割比分别达到86.4%和69.3%,超过了以前的最先进的结果。摘要:Respiratory disease, the third leading cause of deaths globally, is considered a high-priority ailment requiring significant research on identification and treatment. Stethoscope-recorded lung sounds and artificial intelligence-powered devices have been used to identify lung disorders and aid specialists in making accurate diagnoses. In this study, audio-spectrogram vision transformer (AS-ViT), a new approach for identifying abnormal respiration sounds, was developed. The sounds of the lungs are converted into visual representations called spectrograms using a technique called short-time Fourier transform (STFT). These images are then analyzed using a model called vision transformer to identify different types of respiratory sounds. The classification was carried out using the ICBHI 2017 database, which includes various types of lung sounds with different frequencies, noise levels, and backgrounds. The proposed AS-ViT method was evaluated using three metrics and achieved 79.1% and 59.8% for 60:40 split ratio and 86.4% and 69.3% for 80:20 split ratio in terms of unweighted average recall and overall scores respectively for respiratory sound detection, surpassing previous state-of-the-art results.【4】 SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models标题:SpeechGuard:探索多模式大型语言模型的对抗鲁棒性作者:Raghuveer Peri,Sai Muralidhar Jayanthi,Srikanth Ronanki,Anshu Bhatia,Karel Mundnich,Saket Dingliwal,Nilaksh Das,Zejiang Hou,Goeric Huybrechts,Srikanth Vishnubhotla,Daniel Garcia-Romero,Sundararajan Srinivasan,Kyu J Han,Katrin Kirchhoff备注:9+6 pages, Submitted to ACL 2024摘要:集成语音和大语言模型(SLM),可以遵循语音指令,并生成相关的文本响应最近得到了普及。然而,这些模型的安全性和稳健性在很大程度上仍不清楚。在这项工作中,我们研究了这种防御性语音语言模型对对抗性攻击和越狱的潜在脆弱性。具体来说,我们设计的算法可以生成对抗性的例子,在没有人类参与的情况下,在白盒和黑盒攻击设置中越狱SLM。此外,我们提出了一些对策来阻止这种越狱攻击。我们的模型在带有语音指令的对话数据上进行了训练,在口语问答任务上实现了最先进的性能,在安全性和有用性指标上得分超过80%。尽管有安全护栏,但越狱实验证明了SLM对对抗性扰动和传输攻击的脆弱性,当对跨越12个不同有毒类别的精心设计的有害问题数据集进行评估时,平均攻击成功率分别为90%和10%。然而,我们证明,我们提出的对策,减少攻击成功显着。摘要:Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such instruction-following speech-language models to adversarial attacks and jailbreaking. Specifically, we design algorithms that can generate adversarial examples to jailbreak SLMs in both white-box and black-box attack settings without human involvement. Additionally, we propose countermeasures to thwart such jailbreaking attacks. Our models, trained on dialog data with speech instructions, achieve state-of-the-art performance on spoken question-answering task, scoring over 80% on both safety and helpfulness metrics. Despite safety guardrails, experiments on jailbreaking demonstrate the vulnerability of SLMs to adversarial perturbations and transfer attacks, with average attack success rates of 90% and 10% respectively when evaluated on a dataset of carefully designed harmful questions spanning 12 different toxic categories. However, we demonstrate that our proposed countermeasures reduce the attack success significantly.【5】 SpeechVerse: A Large-scale Generalizable Audio Language Model标题:SpeechVerse:一种大规模可推广的音频语言模型作者:Nilaksh Das,Saket Dingliwal,Srikanth Ronanki,Rohit Paturi,David Huang,Prashant Mathur,Jie Yuan,Dhanush Bekal,Xing Niu,Sai Muralidhar Jayanthi,Xilai Li,Karel Mundnich,Monica Sunkara,Sundararajan Srinivasan,Kyu J Han,Katrin Kirchhoff备注:Single Column, 13 page摘要:大型语言模型(LLM)在执行需要对自然语言指令进行语义理解的任务方面表现出令人难以置信的熟练程度。最近,许多作品进一步扩展了这种感知多模态音频和文本输入的能力,但它们的能力往往局限于特定的微调任务,如自动语音识别和翻译。因此,我们开发了SpeechVerse,这是一个强大的多任务训练和课程学习框架,它通过一小组可学习的参数将预训练的语音和文本基础模型结合起来,同时在训练过程中保持预训练的模型冻结。使用从语音基础模型提取的连续潜在表示对模型进行指令微调,以使用自然语言指令在各种语音处理任务上实现最佳zero-shot性能。我们执行广泛的基准测试,包括将我们的模型性能与多个数据集和任务的传统基线进行比较。此外,我们评估模型的能力,通过测试域外数据集,新的提示,和看不见的任务,广义指令。我们的实证实验表明,我们的多任务SpeechVerse模型在11个任务中的9个任务上甚至优于传统的特定任务基线。摘要:Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have further expanded this capability to perceive multimodal audio and text inputs, but their capabilities are often limited to specific fine-tuned tasks such as automatic speech recognition and translation. We therefore develop SpeechVerse, a robust multi-task training and curriculum learning framework that combines pre-trained speech and text foundation models via a small set of learnable parameters, while keeping the pre-trained models frozen during training. The models are instruction finetuned using continuous latent representations extracted from the speech foundation model to achieve optimal zero-shot performance on a diverse range of speech processing tasks using natural language instructions. We perform extensive benchmarking that includes comparing our model performance against traditional baselines across several datasets and tasks. Furthermore, we evaluate the model's capability for generalized instruction following by testing on out-of-domain datasets, novel prompts, and unseen tasks. Our empirical experiments reveal that our multi-task SpeechVerse model is even superior to conventional task-specific baselines on 9 out of the 11 tasks.【6】 A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech标题:预测学习模型可以模拟连续语音的神经表示中发现的时间动态和上下文效应作者:Oli Danyi Liu,Hao Tang,Naomi Feldman,Sharon Goldwater备注:Accepted to CogSci 2024摘要:言语感知涉及存储和整合顺序呈现的项目。认知神经科学的最新研究已经确定了人类语言神经编码中的时间和语境特征,这些特征可能有助于这种时间处理。在这项研究中,我们模拟了类似的分析,从一个计算模型中提取的表示,该模型是在未标记的语音上训练的,其学习目标是预测即将到来的声学。我们的模拟揭示了类似于大脑信号的时间动态,这意味着这些特性可以在没有语言知识的情况下出现。大脑和模型之间的另一个共同属性是音素的编码模式支持一定程度的跨上下文泛化。然而,我们发现的证据表明,这些概括的有效性取决于具体的上下文,这表明,这种分析本身是不够的,以支持上下文不变编码的存在。摘要:Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may facilitate this temporal processing. In this study, we simulated similar analyses with representations extracted from a computational model that was trained on unlabelled speech with the learning objective of predicting upcoming acoustics. Our simulations revealed temporal dynamics similar to those in brain signals, implying that these properties can arise without linguistic knowledge. Another property shared between brains and the model is that the encoding patterns of phonemes support some degree of cross-context generalization. However, we found evidence that the effectiveness of these generalizations depends on the specific contexts, which suggests that this analysis alone is insufficient to support the presence of context-invariant encoding.【7】 Diff-ETS: Learning a Diffusion Probabilistic Model for Electromyography-to-Speech Conversion标题:Dist-ETS:学习肌电到语音转换的扩散概率模型作者:Zhao Ren,Kevin Scheck,Qinhan Hou,Stefano van Gogh,Michael Wand,Tanja Schultz摘要:肌电图到语音(ETS)的转换已被证明是无声的语音接口的潜力,从肌电图(EMG)信号在无声的发音过程中产生可听的语音。ETS模型通常由一个EMG编码器和一个声码器组成,EMG编码器将EMG信号转换为声学语音特征,声码器然后合成语音信号。由于可用数据量不足和噪声信号,合成语音通常表现出低水平的自然度。在这项工作中,我们提出了Diff-ETS,ETS模型,它使用基于分数的扩散概率模型,以提高合成语音的自然度。扩散模型被应用于改善由EMG编码器预测的声学特征的质量。在我们的实验中,我们评估了对预先训练的EMG编码器的预测进行微调的扩散模型,并以端到端的方式训练这两个模型。我们比较了Diff-ETS与基线ETS模型没有扩散使用客观指标和听力测试。实验结果表明,提出的Diff-ETS显著提高了语音的自然度。摘要:Electromyography-to-Speech (ETS) conversion has demonstrated its potential for silent speech interfaces by generating audible speech from Electromyography (EMG) signals during silent articulations. ETS models usually consist of an EMG encoder which converts EMG signals to acoustic speech features, and a vocoder which then synthesises the speech signals. Due to an inadequate amount of available data and noisy signals, the synthesised speech often exhibits a low level of naturalness. In this work, we propose Diff-ETS, an ETS model which uses a score-based diffusion probabilistic model to enhance the naturalness of synthesised speech. The diffusion model is applied to improve the quality of the acoustic features predicted by an EMG encoder. In our experiments, we evaluated fine-tuning the diffusion model on predictions of a pre-trained EMG encoder, and training both models in an end-to-end fashion. We compared Diff-ETS with a baseline ETS model without diffusion using objective metrics and a listening test. The results indicated the proposed Diff-ETS significantly improved speech naturalness over the baseline.
【8】 A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes标题:能够平衡沉浸式和增强模式的可调双耳音频遥现系统作者:Yicheng Hsu,Mingsian R. Bai摘要:双耳音频临场感(BAT)旨在将远端的声学场景编码为近端用户的双耳信号。BAT涵盖了广泛的应用,可以在沉浸式BAT(I—BAT)和增强型BAT(E—BAT)两种极端模式之间变化。对于I—BAT,我们的目标是保持完整的氛围,就像我们在远端一样,而对于E—BAT,我们的目标是通过显著提高的语音质量和可懂度来增强远端对话。为此,本文提出了一种可调BAT系统,在这两种AT模式之间变化,具有所需的特定于应用的平衡。麦克风信号被转换成具有规定的环境因子的双耳信号。提出了一种新的空间相干表示(SCORE)作为模型训练的输入特征,使网络对不同的阵列设置保持鲁棒性。实验结果表明,所提出的BAT的优越性能,即使在阵列配置不包括在训练阶段。摘要:Binaural Audio Telepresence (BAT) aims to encode the acoustic scene at the far end into binaural signals for the user at the near end. BAT encompasses an immense range of applications that can vary between two extreme modes of Immersive BAT (I-BAT) and Enhanced BAT (E-BAT). With I-BAT, our goal is to preserve the full ambience as if we were at the far end, while with E-BAT, our goal is to enhance the far-end conversation with significantly improved speech quality and intelligibility. To this end, this paper presents a tunable BAT system to vary between these two AT modes with a desired application-specific balance. Microphone signals are converted into binaural signals with prescribed ambience factor. A novel Spatial COherence REpresentation (SCORE) is proposed as an input feature for model training so that the network remains robust to different array setups. Experimental results demonstrated the superior performance of the proposed BAT, even when the array configurations were not included in the training phase.
【9】 Simple and Efficient Quantization Techniques for Neural Speech Coding作者:Andreas Brendel,Nicola Pia,Kishan Gupta,Guillaume Fuchs,Markus Multrus摘要:神经音频编码已经成为一个生动的研究方向,有望在非常低的比特率,经典编码技术无法实现良好的音频质量。这里,端到端可训练的自动编码器类模型代表了现有技术,其中必须学习自动编码器瓶颈中的离散表示,以允许输入音频信号的有效传输。这种离散表示通常通过将量化器应用于神经编码器的输出来生成。在几乎所有现有技术的神经音频编码方法中,该量化器被实现为矢量量化器(VQ),并且当与神经音频编码器一起使用时,已经花费了大量努力来减轻该量化技术的缺点。在本文中,我们提出了简单的替代VQ,这是基于投影标量量化(SQ)。这些量化技术不需要任何额外的损失、调度参数或码本存储,从而简化了神经音频编解码器的训练。此外,我们提出了一个新的因果网络架构的神经语音编码,表现出良好的性能,在非常低的计算复杂度。摘要:Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end trainable autoencoder-like models represent the state of the art, where a discrete representation in the bottleneck of the autoencoder has to be learned that allows for efficient transmission of the input audio signal. This discrete representation is typically generated by applying a quantizer to the output of the neural encoder. In almost all state-of-the-art neural audio coding approaches, this quantizer is realized as a Vector Quantizer (VQ) and a lot of effort has been spent to alleviate drawbacks of this quantization technique when used together with a neural audio coder. In this paper, we propose simple alternatives to VQ, which are based on projected Scalar Quantization (SQ). These quantization techniques do not need any additional losses, scheduling parameters or codebook storage thereby simplifying the training of neural audio codecs. Furthermore, we propose a new causal network architecture for neural speech coding that shows good performance at very low computational complexity.
【10】 Semantic MIMO Systems for Speech-to-Text Transmission作者:Zhenzi Weng,Zhijin Qin,Huiqiang Xie,Xiaoming Tao,Khaled B. Letaief摘要:语义通信已被用于通过传输任务相关的语义信息而不是比特来执行许多智能任务。本文针对单用户多输入多输出(MIMO)和多用户MIMO通信场景,提出了一种语义感知的语音到文本传输系统SAC-ST。首先设计了一个语义通信系统,用于在接收端完成语音到文本的转换任务,该系统利用Transformer模块压缩语义信息,生成低维语义特征。此外,提出了一种新的语义感知网络,以促进高语义保真度的传输,以识别关键的语义信息,并保证它被准确地恢复。此外,我们扩展了SAC-ST与神经网络使能的信道估计网络,以减轻对准确的信道状态信息的依赖,并验证SAC-ST在实际通信环境中的可行性。仿真结果将表明,建议的SAC-ST优于通信框架没有语义感知网络的语音到文本传输的MIMO信道的语音到文本的指标,特别是在低信号噪声制度。此外,SAC-ST与发达的信道估计网络是可比的SAC-ST与完善的信道状态信息。摘要:Semantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the single-user multiple-input multiple-output (MIMO) and multi-user MIMO communication scenarios, named SAC-ST. Particularly, a semantic communication system to serve the speech-to-text task at the receiver is first designed, which compresses the semantic information and generates the low-dimensional semantic features by leveraging the transformer module. In addition, a novel semantic-aware network is proposed to facilitate the transmission with high semantic fidelity to identify the critical semantic information and guarantee it is recovered accurately. Furthermore, we extend the SAC-ST with a neural network-enabled channel estimation network to mitigate the dependence on accurate channel state information and validate the feasibility of SAC-ST in practical communication environments. Simulation results will show that the proposed SAC-ST outperforms the communication framework without the semantic-aware network for speech-to-text transmission over the MIMO channels in terms of the speech-to-text metrics, especially in the low signal-to-noise regime. Moreover, the SAC-ST with the developed channel estimation network is comparable to the SAC-ST with perfect channel state information.【1】 A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes标题:能够平衡沉浸式和增强模式的可调双耳音频遥现系统作者:Yicheng Hsu,Mingsian R. Bai摘要:双耳音频临场感(BAT)旨在将远端的声学场景编码为近端用户的双耳信号。BAT涵盖了广泛的应用,可以在沉浸式BAT(I—BAT)和增强型BAT(E—BAT)两种极端模式之间变化。对于I—BAT,我们的目标是保持完整的氛围,就像我们在远端一样,而对于E—BAT,我们的目标是通过显著提高的语音质量和可懂度来增强远端对话。为此,本文提出了一种可调BAT系统,在这两种AT模式之间变化,具有所需的特定于应用的平衡。麦克风信号被转换成具有规定的环境因子的双耳信号。提出了一种新的空间相干表示(SCORE)作为模型训练的输入特征,使网络对不同的阵列设置保持鲁棒性。实验结果表明,所提出的BAT的优越性能,即使在阵列配置不包括在训练阶段。摘要:Binaural Audio Telepresence (BAT) aims to encode the acoustic scene at the far end into binaural signals for the user at the near end. BAT encompasses an immense range of applications that can vary between two extreme modes of Immersive BAT (I-BAT) and Enhanced BAT (E-BAT). With I-BAT, our goal is to preserve the full ambience as if we were at the far end, while with E-BAT, our goal is to enhance the far-end conversation with significantly improved speech quality and intelligibility. To this end, this paper presents a tunable BAT system to vary between these two AT modes with a desired application-specific balance. Microphone signals are converted into binaural signals with prescribed ambience factor. A novel Spatial COherence REpresentation (SCORE) is proposed as an input feature for model training so that the network remains robust to different array setups. Experimental results demonstrated the superior performance of the proposed BAT, even when the array configurations were not included in the training phase.【2】 Simple and Efficient Quantization Techniques for Neural Speech Coding作者:Andreas Brendel,Nicola Pia,Kishan Gupta,Guillaume Fuchs,Markus Multrus摘要:神经音频编码已经成为一个生动的研究方向,有望在非常低的比特率,经典编码技术无法实现良好的音频质量。这里,端到端可训练的自动编码器类模型代表了现有技术,其中必须学习自动编码器瓶颈中的离散表示,以允许输入音频信号的有效传输。这种离散表示通常通过将量化器应用于神经编码器的输出来生成。在几乎所有现有技术的神经音频编码方法中,该量化器被实现为矢量量化器(VQ),并且当与神经音频编码器一起使用时,已经花费了大量努力来减轻该量化技术的缺点。在本文中,我们提出了简单的替代VQ,这是基于投影标量量化(SQ)。这些量化技术不需要任何额外的损失、调度参数或码本存储,从而简化了神经音频编解码器的训练。此外,我们提出了一个新的因果网络架构的神经语音编码,表现出良好的性能,在非常低的计算复杂度。摘要:Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end trainable autoencoder-like models represent the state of the art, where a discrete representation in the bottleneck of the autoencoder has to be learned that allows for efficient transmission of the input audio signal. This discrete representation is typically generated by applying a quantizer to the output of the neural encoder. In almost all state-of-the-art neural audio coding approaches, this quantizer is realized as a Vector Quantizer (VQ) and a lot of effort has been spent to alleviate drawbacks of this quantization technique when used together with a neural audio coder. In this paper, we propose simple alternatives to VQ, which are based on projected Scalar Quantization (SQ). These quantization techniques do not need any additional losses, scheduling parameters or codebook storage thereby simplifying the training of neural audio codecs. Furthermore, we propose a new causal network architecture for neural speech coding that shows good performance at very low computational complexity.【3】 Semantic MIMO Systems for Speech-to-Text Transmission作者:Zhenzi Weng,Zhijin Qin,Huiqiang Xie,Xiaoming Tao,Khaled B. Letaief摘要:语义通信已被用于通过传输任务相关的语义信息而不是比特来执行许多智能任务。本文针对单用户多输入多输出(MIMO)和多用户MIMO通信场景,提出了一种语义感知的语音到文本传输系统SAC-ST。首先设计了一个语义通信系统,用于在接收端完成语音到文本的转换任务,该系统利用Transformer模块压缩语义信息,生成低维语义特征。此外,提出了一种新的语义感知网络,以促进高语义保真度的传输,以识别关键的语义信息,并保证它被准确地恢复。此外,我们扩展了SAC-ST与神经网络使能的信道估计网络,以减轻对准确的信道状态信息的依赖,并验证SAC-ST在实际通信环境中的可行性。仿真结果将表明,建议的SAC-ST优于通信框架没有语义感知网络的语音到文本传输的MIMO信道的语音到文本的指标,特别是在低信号噪声制度。此外,SAC-ST与发达的信道估计网络是可比的SAC-ST与完善的信道状态信息。摘要:Semantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the single-user multiple-input multiple-output (MIMO) and multi-user MIMO communication scenarios, named SAC-ST. Particularly, a semantic communication system to serve the speech-to-text task at the receiver is first designed, which compresses the semantic information and generates the low-dimensional semantic features by leveraging the transformer module. In addition, a novel semantic-aware network is proposed to facilitate the transmission with high semantic fidelity to identify the critical semantic information and guarantee it is recovered accurately. Furthermore, we extend the SAC-ST with a neural network-enabled channel estimation network to mitigate the dependence on accurate channel state information and validate the feasibility of SAC-ST in practical communication environments. Simulation results will show that the proposed SAC-ST outperforms the communication framework without the semantic-aware network for speech-to-text transmission over the MIMO channels in terms of the speech-to-text metrics, especially in the low signal-to-noise regime. Moreover, the SAC-ST with the developed channel estimation network is comparable to the SAC-ST with perfect channel state information.【4】 Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning标题:研究通用音频表示学习的联合嵌入预测架构中的设计选择作者:Alain Riou,Stefan Lattner,Gaëtan Hadjeres,Geoffroy Peeters备注:Self-supervision in Audio, Speech and Beyond workshop, IEEE International Conference on Acoustics, Speech, and Signal Processing, 2024摘要:本文讨论了自监督通用音频表示学习问题。我们探索使用联合嵌入预测架构(JEPA)来完成这项任务,该任务包括将输入mel频谱图分为两部分(上下文和目标),计算每个部分的神经表示,并训练神经网络从上下文表示中预测目标表示。我们在这个框架内研究了几个设计选择,并通过广泛的实验研究它们的影响,通过评估我们的模型对各种音频分类基准,包括环境声音,语音和音乐下游任务。我们特别关注输入数据的哪一部分被用作上下文或目标,并通过实验表明它会显著影响模型的质量。特别是,我们注意到,在图像域中的一些有效的设计选择导致音频性能不佳,从而突出了这两种模式之间的主要差异。摘要:This paper addresses the problem of self-supervised general-purpose audio representation learning. We explore the use of Joint-Embedding Predictive Architectures (JEPA) for this task, which consists of splitting an input mel-spectrogram into two parts (context and target), computing neural representations for each, and training the neural network to predict the target representations from the context representations. We investigate several design choices within this framework and study their influence through extensive experiments by evaluating our models on various audio classification benchmarks, including environmental sounds, speech and music downstream tasks. We focus notably on which part of the input data is used as context or target and show experimentally that it significantly impacts the model's quality. In particular, we notice that some effective design choices in the image domain lead to poor performance on audio, thus highlighting major differences between these two modalities.
【5】 EVDA: Evolving Deepfake Audio Detection Continual Learning Benchmark标题:EVDA:不断发展的Deepfake音频检测持续学习基准作者:Xiaohui Zhang,Jiangyan Yi,Jianhua Tao摘要:GPT—4、GPT—4o和Claude家族等高级大型语言模型的兴起,使得虚假音频检测变得越来越具有挑战性。传统的微调方法很难跟上合成语音不断发展的步伐,需要不断学习的方法来适应新的音频,同时保留检测旧类型的能力。持续学习是检测新出现的deepfake音频的有效工具,同时保持旧类型的性能,缺乏一个结构良好且用户友好的评估框架。为了解决这一差距,我们引入了EVDA,这是一个用于评估deepfake音频检测中持续学习方法的基准。EVDA包括来自反欺骗语音系列、中文虚假音频检测系列的经典数据集,以及来自GPT—4和GPT—4o等模型的新生成的deepfake音频。它支持各种持续学习技术,如弹性权重合并(EWC),学习不遗忘(LwF),以及最近的方法,如正则化自适应权重修改(RAWM)和弧度权重修改(RWM)。此外,EVDA通过提供用于集成新的持续学习方法的开放接口来促进鲁棒算法的开发摘要:The rise of advanced large language models such as GPT-4, GPT-4o, and the Claude family has made fake audio detection increasingly challenging. Traditional fine-tuning methods struggle to keep pace with the evolving landscape of synthetic speech, necessitating continual learning approaches that can adapt to new audio while retaining the ability to detect older types. Continual learning, which acts as an effective tool for detecting newly emerged deepfake audio while maintaining performance on older types, lacks a well-constructed and user-friendly evaluation framework. To address this gap, we introduce EVDA, a benchmark for evaluating continual learning methods in deepfake audio detection. EVDA includes classic datasets from the Anti-Spoofing Voice series, Chinese fake audio detection series, and newly generated deepfake audio from models like GPT-4 and GPT-4o. It supports various continual learning techniques, such as Elastic Weight Consolidation (EWC), Learning without Forgetting (LwF), and recent methods like Regularized Adaptive Weight Modification (RAWM) and Radian Weight Modification (RWM). Additionally, EVDA facilitates the development of robust algorithms by providing an open interface for integrating new continual learning methods
【6】 Abnormal Respiratory Sound Identification Using Audio-Spectrogram Vision Transformer标题:使用音频频谱图视觉Transformer识别异常呼吸音作者:Whenty Ariyanti,Kai-Chun Liu,Kuan-Yu Chen,Yu Tsao摘要:呼吸道疾病是全球第三大死亡原因,被认为是一种高度优先的疾病,需要对识别和治疗进行大量研究。听诊器记录的肺部声音和人工智能驱动的设备已被用于识别肺部疾病,并帮助专家做出准确的诊断。在这项研究中,音频频谱图Vision Transformer(AS—ViT),一种新的方法来识别异常呼吸音,开发。肺部的声音被转换成视觉表示称为频谱图使用一种技术称为短时傅立叶变换(STFT)。然后使用称为Vision Transformer的模型分析这些图像,以识别不同类型的呼吸声。该分类使用ICBHI 2017数据库进行,该数据库包括具有不同频率,噪声水平和背景的各种类型的肺音。提出的AS—ViT方法使用三个指标进行评估,在呼吸音检测的未加权平均召回率和总得分方面,60:40分割比分别达到79.1%和59.8%,80:20分割比分别达到86.4%和69.3%,超过了以前的最先进的结果。摘要:Respiratory disease, the third leading cause of deaths globally, is considered a high-priority ailment requiring significant research on identification and treatment. Stethoscope-recorded lung sounds and artificial intelligence-powered devices have been used to identify lung disorders and aid specialists in making accurate diagnoses. In this study, audio-spectrogram vision transformer (AS-ViT), a new approach for identifying abnormal respiration sounds, was developed. The sounds of the lungs are converted into visual representations called spectrograms using a technique called short-time Fourier transform (STFT). These images are then analyzed using a model called vision transformer to identify different types of respiratory sounds. The classification was carried out using the ICBHI 2017 database, which includes various types of lung sounds with different frequencies, noise levels, and backgrounds. The proposed AS-ViT method was evaluated using three metrics and achieved 79.1% and 59.8% for 60:40 split ratio and 86.4% and 69.3% for 80:20 split ratio in terms of unweighted average recall and overall scores respectively for respiratory sound detection, surpassing previous state-of-the-art results.【7】 SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models标题:SpeechGuard:探索多模式大型语言模型的对抗鲁棒性作者:Raghuveer Peri,Sai Muralidhar Jayanthi,Srikanth Ronanki,Anshu Bhatia,Karel Mundnich,Saket Dingliwal,Nilaksh Das,Zejiang Hou,Goeric Huybrechts,Srikanth Vishnubhotla,Daniel Garcia-Romero,Sundararajan Srinivasan,Kyu J Han,Katrin Kirchhoff备注:9+6 pages, Submitted to ACL 2024摘要:集成语音和大语言模型(SLM),可以遵循语音指令,并生成相关的文本响应最近得到了普及。然而,这些模型的安全性和稳健性在很大程度上仍不清楚。在这项工作中,我们研究了这种防御性语音语言模型对对抗性攻击和越狱的潜在脆弱性。具体来说,我们设计的算法可以生成对抗性的例子,在没有人类参与的情况下,在白盒和黑盒攻击设置中越狱SLM。此外,我们提出了一些对策来阻止这种越狱攻击。我们的模型在带有语音指令的对话数据上进行了训练,在口语问答任务上实现了最先进的性能,在安全性和有用性指标上得分超过80%。尽管有安全护栏,但越狱实验证明了SLM对对抗性扰动和传输攻击的脆弱性,当对跨越12个不同有毒类别的精心设计的有害问题数据集进行评估时,平均攻击成功率分别为90%和10%。然而,我们证明,我们提出的对策,减少攻击成功显着。摘要:Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such instruction-following speech-language models to adversarial attacks and jailbreaking. Specifically, we design algorithms that can generate adversarial examples to jailbreak SLMs in both white-box and black-box attack settings without human involvement. Additionally, we propose countermeasures to thwart such jailbreaking attacks. Our models, trained on dialog data with speech instructions, achieve state-of-the-art performance on spoken question-answering task, scoring over 80% on both safety and helpfulness metrics. Despite safety guardrails, experiments on jailbreaking demonstrate the vulnerability of SLMs to adversarial perturbations and transfer attacks, with average attack success rates of 90% and 10% respectively when evaluated on a dataset of carefully designed harmful questions spanning 12 different toxic categories. However, we demonstrate that our proposed countermeasures reduce the attack success significantly.【8】 SpeechVerse: A Large-scale Generalizable Audio Language Model标题:SpeechVerse:一种大规模可推广的音频语言模型作者:Nilaksh Das,Saket Dingliwal,Srikanth Ronanki,Rohit Paturi,David Huang,Prashant Mathur,Jie Yuan,Dhanush Bekal,Xing Niu,Sai Muralidhar Jayanthi,Xilai Li,Karel Mundnich,Monica Sunkara,Sundararajan Srinivasan,Kyu J Han,Katrin Kirchhoff备注:Single Column, 13 page摘要:大型语言模型(LLM)在执行需要对自然语言指令进行语义理解的任务方面表现出令人难以置信的熟练程度。最近,许多作品进一步扩展了这种感知多模态音频和文本输入的能力,但它们的能力往往局限于特定的微调任务,如自动语音识别和翻译。因此,我们开发了SpeechVerse,这是一个强大的多任务训练和课程学习框架,它通过一小组可学习的参数将预训练的语音和文本基础模型结合起来,同时在训练过程中保持预训练的模型冻结。使用从语音基础模型提取的连续潜在表示对模型进行指令微调,以使用自然语言指令在各种语音处理任务上实现最佳zero-shot性能。我们执行广泛的基准测试,包括将我们的模型性能与多个数据集和任务的传统基线进行比较。此外,我们评估模型的能力,通过测试域外数据集,新的提示,和看不见的任务,广义指令。我们的实证实验表明,我们的多任务SpeechVerse模型在11个任务中的9个任务上甚至优于传统的特定任务基线。摘要:Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have further expanded this capability to perceive multimodal audio and text inputs, but their capabilities are often limited to specific fine-tuned tasks such as automatic speech recognition and translation. We therefore develop SpeechVerse, a robust multi-task training and curriculum learning framework that combines pre-trained speech and text foundation models via a small set of learnable parameters, while keeping the pre-trained models frozen during training. The models are instruction finetuned using continuous latent representations extracted from the speech foundation model to achieve optimal zero-shot performance on a diverse range of speech processing tasks using natural language instructions. We perform extensive benchmarking that includes comparing our model performance against traditional baselines across several datasets and tasks. Furthermore, we evaluate the model's capability for generalized instruction following by testing on out-of-domain datasets, novel prompts, and unseen tasks. Our empirical experiments reveal that our multi-task SpeechVerse model is even superior to conventional task-specific baselines on 9 out of the 11 tasks.
【9】 A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech标题:预测学习模型可以模拟连续语音的神经表示中发现的时间动态和上下文效应作者:Oli Danyi Liu,Hao Tang,Naomi Feldman,Sharon Goldwater备注:Accepted to CogSci 2024摘要:言语感知涉及存储和整合顺序呈现的项目。认知神经科学的最新研究已经确定了人类语言神经编码中的时间和语境特征,这些特征可能有助于这种时间处理。在这项研究中,我们模拟了类似的分析,从一个计算模型中提取的表示,该模型是在未标记的语音上训练的,其学习目标是预测即将到来的声学。我们的模拟揭示了类似于大脑信号的时间动态,这意味着这些特性可以在没有语言知识的情况下出现。大脑和模型之间的另一个共同属性是音素的编码模式支持一定程度的跨上下文泛化。然而,我们发现的证据表明,这些概括的有效性取决于具体的上下文,这表明,这种分析本身是不够的,以支持上下文不变编码的存在。摘要:Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may facilitate this temporal processing. In this study, we simulated similar analyses with representations extracted from a computational model that was trained on unlabelled speech with the learning objective of predicting upcoming acoustics. Our simulations revealed temporal dynamics similar to those in brain signals, implying that these properties can arise without linguistic knowledge. Another property shared between brains and the model is that the encoding patterns of phonemes support some degree of cross-context generalization. However, we found evidence that the effectiveness of these generalizations depends on the specific contexts, which suggests that this analysis alone is insufficient to support the presence of context-invariant encoding.【10】 Diff-ETS: Learning a Diffusion Probabilistic Model for Electromyography-to-Speech Conversion标题:Dist-ETS:学习肌电到语音转换的扩散概率模型作者:Zhao Ren,Kevin Scheck,Qinhan Hou,Stefano van Gogh,Michael Wand,Tanja Schultz摘要:肌电图到语音(ETS)的转换已被证明是无声的语音接口的潜力,从肌电图(EMG)信号在无声的发音过程中产生可听的语音。ETS模型通常由一个EMG编码器和一个声码器组成,EMG编码器将EMG信号转换为声学语音特征,声码器然后合成语音信号。由于可用数据量不足和噪声信号,合成语音通常表现出低水平的自然度。在这项工作中,我们提出了Diff-ETS,ETS模型,它使用基于分数的扩散概率模型,以提高合成语音的自然度。扩散模型被应用于改善由EMG编码器预测的声学特征的质量。在我们的实验中,我们评估了对预先训练的EMG编码器的预测进行微调的扩散模型,并以端到端的方式训练这两个模型。我们比较了Diff-ETS与基线ETS模型没有扩散使用客观指标和听力测试。实验结果表明,提出的Diff-ETS显著提高了语音的自然度。摘要:Electromyography-to-Speech (ETS) conversion has demonstrated its potential for silent speech interfaces by generating audible speech from Electromyography (EMG) signals during silent articulations. ETS models usually consist of an EMG encoder which converts EMG signals to acoustic speech features, and a vocoder which then synthesises the speech signals. Due to an inadequate amount of available data and noisy signals, the synthesised speech often exhibits a low level of naturalness. In this work, we propose Diff-ETS, an ETS model which uses a score-based diffusion probabilistic model to enhance the naturalness of synthesised speech. The diffusion model is applied to improve the quality of the acoustic features predicted by an EMG encoder. In our experiments, we evaluated fine-tuning the diffusion model on predictions of a pre-trained EMG encoder, and training both models in an end-to-end fashion. We compared Diff-ETS with a baseline ETS model without diffusion using objective metrics and a listening test. The results indicated the proposed Diff-ETS significantly improved speech naturalness over the baseline.