今日论文合集:cs.SD语音7篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for  Noise-Robust Speech Perception

标题:XLAVS—R:跨语言音视频语音表示学习的噪声鲁棒语音感知

链接:https://arxiv.org/abs/2403.14402

作者:HyoJung Han,Mohamed Anwar,Juan Pino,Wei-Ning Hsu,Marine Carpuat,Bowen Shi,Changhan Wang

摘要:语音识别和翻译系统在现实环境中经常出现的噪声输入上表现不佳。用视觉信号增强这些系统有可能提高对噪声的鲁棒性。但是,视听数据的数量有限,而且语种少于纯视听资源。为了解决这一差距,我们提出了XLAVS-R,一个跨语言的视听语音表示模型,用于100多种语言的噪声鲁棒语音识别和翻译。它旨在通过建立在仅音频的多语言预训练之上并简化现有的预训练方案,最大限度地发挥有限的多语言AV预训练数据的优势。对MuAViC基准测试的广泛评估显示了XLAVS-R在下游视听语音识别和翻译任务上的实力,在有噪声的AV输入下,它的WER和BLEU性能比之前的最先进水平高出18.5%,并通过仅音频微调实现了强大的zero-shot视听能力。

摘要:Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. However, audio-visual (AV) data is only available in limited amounts and for fewer languages than audio-only resources. To address this gap, we present XLAVS-R, a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. It is designed to maximize the benefits of limited multilingual AV pre-training data, by building on top of audio-only multilingual pre-training and simplifying existing pre-training schemes. Extensive evaluation on the MuAViC benchmark shows the strength of XLAVS-R on downstream audio-visual speech recognition and translation tasks, where it outperforms the previous state of the art by up to 18.5% WER and 4.7 BLEU given noisy AV inputs, and enables strong zero-shot audio-visual ability with audio-only fine-tuning.


【2】 Exploring Green AI for Audio Deepfake Detection
标题:探索绿色人工智能用于音频Deepfake检测
链接:https://arxiv.org/abs/2403.14290
作者:Subhajit Saha,Md Sahidullah,Swagatam Das
备注:This manuscript is under review in a conference
摘要:利用深度神经网络的最先进的音频deepfake检测器表现出令人印象深刻的识别性能。然而,这一优势伴随着显著的碳足迹。这主要是由于使用了具有加速器的高性能计算和高训练时间。研究表明,平均深度NLP模型产生约626 k lbs的CO\text下标{2},相当于美国汽车寿命期平均排放量的五倍。这对环境无疑是一个巨大的威胁。为了应对这一挑战,这项研究提出了一种新的音频deepfake检测框架,可以使用标准CPU资源进行无缝训练。我们提出的框架利用了现成的基于自监督学习(SSL)的模型,这些模型是预先训练好的,可以在公共存储库中使用。与现有的微调SSL模型并为下游任务采用额外的深度神经网络的方法相比,我们利用经典的机器学习算法,如逻辑回归和浅层神经网络,使用使用预训练模型提取的SSL嵌入。与常用的高碳足迹方法相比,我们的方法显示出具有竞争力的结果。在ASVspoof 2019 LA数据集的实验中,我们实现了0.90\%的等错误率(EER),可训练的模型参数少于1 k。为了鼓励在这一方向上的进一步研究并支持可重复的结果,Python代码将在接受后公开访问\footnote{\href{https://github.com/sahasubhajit/Speech-Spoofing-}{GitHub link}}。
摘要:The state-of-the-art audio deepfake detectors leveraging deep neural networks exhibit impressive recognition performance. Nonetheless, this advantage is accompanied by a significant carbon footprint. This is mainly due to the use of high-performance computing with accelerators and high training time. Studies show that average deep NLP model produces around 626k lbs of CO\textsubscript{2} which is equivalent to five times of average US car emission at its lifetime. This is certainly a massive threat to the environment. To tackle this challenge, this study presents a novel framework for audio deepfake detection that can be seamlessly trained using standard CPU resources. Our proposed framework utilizes off-the-shelve self-supervised learning (SSL) based models which are pre-trained and available in public repositories. In contrast to existing methods that fine-tune SSL models and employ additional deep neural networks for downstream tasks, we exploit classical machine learning algorithms such as logistic regression and shallow neural networks using the SSL embeddings extracted using the pre-trained model. Our approach shows competitive results compared to the commonly used high-carbon footprint approaches. In experiments with the ASVspoof 2019 LA dataset, we achieve a 0.90\% equal error rate (EER) with less than 1k trainable model parameters. To encourage further research in this direction and support reproducible results, the Python code will be made publicly accessible following acceptance\footnote{\href{https://github.com/sahasubhajit/Speech-Spoofing-}{GitHub link}}.


【3】 Assessing the Robustness of Spectral Clustering for Deep Speaker  Diarization
标题:谱聚类对深度说话人分类的鲁棒性评估
链接:https://arxiv.org/abs/2403.14286
作者:Nikhil Raghav,Md Sahidullah
备注:Manuscript Under Review
摘要:聚类说话人嵌入在说话人日志化中至关重要,但没有像其他组件那样受到关注。此外,当开发和评估数据来自不同领域时,说话人日志化在不同数据集上的鲁棒性还没有被探索。为了弥补这一差距,本研究彻底检查频谱聚类同域和跨域扬声器diarization。我们在两个广泛使用的语料库,AMI和DIHARD上进行了大量的实验,揭示了在领域失配的情况下说话人日记化的性能趋势。我们观察到,两个不同的域条件之间的性能差异可以归因于谱聚类的作用。特别是,保持其他模块不变,我们表明,最佳调谐参数以及扬声器数量估计的差异起源于由于不匹配。本研究为说话人日记化研究开辟了几个方向。
摘要:Clustering speaker embeddings is crucial in speaker diarization but hasn't received as much focus as other components. Moreover, the robustness of speaker diarization across various datasets hasn't been explored when the development and evaluation data are from different domains. To bridge this gap, this study thoroughly examines spectral clustering for both same-domain and cross-domain speaker diarization. Our extensive experiments on two widely used corpora, AMI and DIHARD, reveal the performance trend of speaker diarization in the presence of domain mismatch. We observe that the performance difference between two different domain conditions can be attributed to the role of spectral clustering. In particular, keeping other modules unchanged, we show that differences in optimal tuning parameters as well as speaker count estimation originates due to the mismatch. This study opens several future directions for speaker diarization research.

【4】 emoDARTS: Joint Optimisation of CNN & Sequential Neural Network  Architectures for Superior Speech Emotion Recognition
标题:emoDARTS:CNN和序列神经网络架构的联合优化,用于卓越语音情感识别
链接:https://arxiv.org/abs/2403.14083
作者:Thejan Rajapakshe,Rajib Rana,Sara Khalifa,Berrak Sisman,Bjorn W. Schuller,Carlos Busso
备注:Submitted to IEEE Transactions on Affective Computing on February 19, 2024. arXiv admin note: text overlap with arXiv:2305.14402
摘要:语音情感识别(SER)对于使计算机能够理解人类交流中传达的情感至关重要。随着深度学习(DL)的最新进展,SER模型的性能得到了显着提高。然而,设计最佳的DL架构需要专业知识和实验评估。幸运的是,神经架构搜索(NAS)为自动确定最佳DL模型提供了一种潜在的解决方案。微分结构搜索(DARTS)是一种发现最优模型的特别有效的方法。这项研究提出了emoDARTS,这是一种DARTS优化的联合CNN和顺序神经网络(SeqNN:LSTM,RNN)架构,可增强SER性能。文献支持选择CNN和LSTM耦合来提高性能。虽然DARTS以前曾用于独立选择CNN和LSTM操作,但我们的技术增加了一种新的机制,用于结合使用DARTS选择CNN和SeqNN操作。与早期的工作不同,我们不对CNN的层顺序施加限制。相反,我们让DARTS选择DARTS单元内的最佳层顺序。我们证明了emoDARTS优于传统设计的CNN-LSTM模型,并通过在IEMOCAP,MSP-IMPROV和MSP-Podcast数据集上评估我们的方法,超越了通过DARTS在CNN-LSTM上实现的最佳报告SER结果。
摘要:Speech Emotion Recognition (SER) is crucial for enabling computers to understand the emotions conveyed in human communication. With recent advancements in Deep Learning (DL), the performance of SER models has significantly improved. However, designing an optimal DL architecture requires specialised knowledge and experimental assessments. Fortunately, Neural Architecture Search (NAS) provides a potential solution for automatically determining the best DL model. The Differentiable Architecture Search (DARTS) is a particularly efficient method for discovering optimal models. This study presents emoDARTS, a DARTS-optimised joint CNN and Sequential Neural Network (SeqNN: LSTM, RNN) architecture that enhances SER performance. The literature supports the selection of CNN and LSTM coupling to improve performance.  While DARTS has previously been used to choose CNN and LSTM operations independently, our technique adds a novel mechanism for selecting CNN and SeqNN operations in conjunction using DARTS. Unlike earlier work, we do not impose limits on the layer order of the CNN. Instead, we let DARTS choose the best layer order inside the DARTS cell. We demonstrate that emoDARTS outperforms conventionally designed CNN-LSTM models and surpasses the best-reported SER results achieved through DARTS on CNN-LSTM by evaluating our approach on the IEMOCAP, MSP-IMPROV, and MSP-Podcast datasets.


【5】 The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio  Benchmarks and Novel Data
标题:NeurIPS 2023音频机器学习研讨会:情感音频基准和新数据
链接:https://arxiv.org/abs/2403.14048
作者:Alice Baird,Rachel Manzelli,Panagiotis Tzirakis,Chris Gagne,Haoqi Li,Sadie Allen,Sander Dieleman,Brian Kulis,Shrikanth S. Narayanan,Alan Cowen
摘要:NeurIPS 2023 Machine Learning for Audio研讨会汇集了来自各个音频领域的机器学习(ML)专家。有几个有价值的音频驱动的ML任务,从语音情感识别到音频事件检测,但与其他ML领域相比,社区是稀疏的,例如,计算机视觉或自然语言处理。音频的一个主要限制是可用的数据;由于音频是一种时间依赖的模式,高质量的数据收集既耗时又昂贵,这使得学术团体将其通常最先进的策略应用于更大,更普遍的数据集变得具有挑战性。在这篇简短的白皮书中,为了鼓励那些对大型数据集访问有限的研究人员,组织者首先概述了几个可供社区使用的开源数据集,并在研讨会期间提供了几个适当的数据集。也就是说,三个声音数据集,Hume-Prosody,Hume-VocalBurst,一个表演情感语音数据集Modulate-Sonata和一个游戏中的流数据集Modulate-Stream。我们概述了这些数据集的当前基线,但鼓励来自音频领域的研究人员在初始基线任务之外使用它们。
摘要:The NeurIPS 2023 Machine Learning for Audio Workshop brings together machine learning (ML) experts from various audio domains. There are several valuable audio-driven ML tasks, from speech emotion recognition to audio event detection, but the community is sparse compared to other ML areas, e.g., computer vision or natural language processing. A major limitation with audio is the available data; with audio being a time-dependent modality, high-quality data collection is time-consuming and costly, making it challenging for academic groups to apply their often state-of-the-art strategies to a larger, more generalizable dataset. In this short white paper, to encourage researchers with limited access to large-datasets, the organizers first outline several open-source datasets that are available to the community, and for the duration of the workshop are making several propriety datasets available. Namely, three vocal datasets, Hume-Prosody, Hume-VocalBurst, an acted emotional speech dataset Modulate-Sonata, and an in-game streamer dataset Modulate-Stream. We outline the current baselines on these datasets but encourage researchers from across audio to utilize them outside of the initial baseline tasks.


【6】 Speech-Aware Neural Diarization with Encoder-Decoder Attractor Guided by  Attention Constraints
标题:用注意约束引导的编解码器吸引子的语音感知神经二值化
链接:https://arxiv.org/abs/2403.14268
作者:PeiYing Lee,HauYun Guo,Berlin Chen备注:Accepted to The 28th International Conference on Technologies and Applications of Artificial Intelligence (TAAI), in Chinese language
摘要:EEND-EDA(End-to-End Neural Diarization with Encoder-Decoder based Attractor)是一种用于自动说话人分割和标记的端到端神经模型。通过估计吸引子的数目,实现了灵活处理说话人数目的能力。然而,EEND-EDA很难准确地捕捉本地扬声器动态。本文提出了一种辅助损失,旨在指导EEND-EDA模型的低层的Transformer编码器,以增强使用说话人活动信息的自我注意模块的效果。在公共数据集Mini LibriSpeech上的测试结果表明,该方法的有效性,将日志错误率从30.95%降低到28.17%。我们将在GitHub上发布源代码,以便进一步研究和再现。
摘要:End-to-End Neural Diarization with Encoder-Decoder based Attractor (EEND-EDA) is an end-to-end neural model for automatic speaker segmentation and labeling. It achieves the capability to handle flexible number of speakers by estimating the number of attractors. EEND-EDA, however, struggles to accurately capture local speaker dynamics. This work proposes an auxiliary loss that aims to guide the Transformer encoders at the lower layer of EEND-EDA model to enhance the effect of self-attention modules using speaker activity information. The results evaluated on public dataset Mini LibriSpeech, demonstrates the effectiveness of the work, reducing Diarization Error Rate from 30.95% to 28.17%. We will release the source code on GitHub to allow further research and reproducibility.

【7】 AdaProj: Adaptively Scaled Angular Margin Subspace Projections for  Anomalous Sound Detection with Auxiliary Classification Tasks
标题:AdaProj:自适应缩放角余量子空间投影,用于辅助分类任务的异常声音检测
链接:https://arxiv.org/abs/2403.14179
作者:Kevin Wilkinghoff
摘要:用于半监督异常声音检测的最新方法是首先通过使用基于Meta信息或自监督学习的辅助分类任务来学习嵌入空间,然后估计正常数据的分布。在这项工作中,AdaProj一个新的损失函数。与常用的角度余量损失(angular margin loss)相比,AdaProj学习将数据投影到特定于类的子空间上,角度余量损失将每个类的数据投影到尽可能接近其对应的类中心的位置。通过这样做,属于正常数据的嵌入的结果分布不需要像其他损失函数那样具有限制性,从而允许更详细地查看数据。在DCASE2022和DCASE2023数据集上进行的实验表明,使用AdaProj学习嵌入空间的性能明显优于其他常用的损失函数,并在DCASE2023数据集上获得了最先进的性能。
摘要:The state-of-the-art approach for semi-supervised anomalous sound detection is to first learn an embedding space by using auxiliary classification tasks based on meta information or self-supervised learning and then estimate the distribution of normal data. In this work, AdaProj a novel loss function is presented. In contrast to commonly used angular margin losses, which project data of each class as close as possible to their corresponding class centers, AdaProj learns to project data onto class-specific subspaces. By doing so, the resulting distributions of embeddings belonging to normal data are not required to be as restrictive as other loss functions allowing a more detailed view on the data. In experiments conducted on the DCASE2022 and DCASE2023 datasets, it is shown that using AdaProj to learn an embedding space significantly outperforms other commonly used loss functions and results in a state-of-the-art performance on the DCASE2023 dataset.


eess.AS音频处理
【1】 Speech-Aware Neural Diarization with Encoder-Decoder Attractor Guided by  Attention Constraints
标题:注意力约束引导下的语音感知神经元分类
链接:https://arxiv.org/abs/2403.14268
作者:PeiYing Lee,HauYun Guo,Berlin Chen
备注:Accepted to The 28th International Conference on Technologies and Applications of Artificial Intelligence (TAAI), in Chinese language
摘要:EEND-EDA(End-to-End Neural Diarization with Encoder-Decoder based Attractor)是一种用于自动说话人分割和标记的端到端神经模型。通过估计吸引子的数目,实现了灵活处理说话人数目的能力。然而,EEND-EDA很难准确地捕捉本地扬声器动态。本文提出了一种辅助损失,旨在指导EEND-EDA模型的低层的Transformer编码器,以增强使用说话人活动信息的自我注意模块的效果。在公共数据集Mini LibriSpeech上的测试结果表明,该方法的有效性,将日志错误率从30.95%降低到28.17%。我们将在GitHub上发布源代码,以便进一步研究和再现。
摘要:End-to-End Neural Diarization with Encoder-Decoder based Attractor (EEND-EDA) is an end-to-end neural model for automatic speaker segmentation and labeling. It achieves the capability to handle flexible number of speakers by estimating the number of attractors. EEND-EDA, however, struggles to accurately capture local speaker dynamics. This work proposes an auxiliary loss that aims to guide the Transformer encoders at the lower layer of EEND-EDA model to enhance the effect of self-attention modules using speaker activity information. The results evaluated on public dataset Mini LibriSpeech, demonstrates the effectiveness of the work, reducing Diarization Error Rate from 30.95% to 28.17%. We will release the source code on GitHub to allow further research and reproducibility.

【2】 CATSE: A Context-Aware Framework for Causal Target Sound Extraction
标题:CATSE:一种基于上下文感知的目标声提取框架
链接:https://arxiv.org/abs/2403.14246
作者:Shrishail Baligar,Mikolaj Kegler,Bryce Irvin,Marko Stamenovic,Shawn Newsam
备注:Submitted to EUSIPCO 2024
摘要:目标声音提取(TSE)的重点是从输入的混合声中分离出用户提示的感兴趣的声源。大多数现有的解决方案以离线方式操作,并且不适合由诸如增强听力的直播流内容中的应用所施加的低延迟因果处理约束。我们介绍了一个家庭的上下文感知的低延迟因果TSE模型适合于实时处理。首先,我们通过提供TSE模型关于什么样的声音类构成输入混合物的Oracle信息来探索上下文的效用,其中模型的目标是提取用户指示的一个或多个感兴趣的源。由于预言机模型的实际应用是有限的,由于他们的假设,我们引入了一个复合的多任务训练目标,涉及分离和分类损失。我们的评估涉及单源和多源提取显示了在模型中使用上下文信息的好处,无论是通过提供完整的上下文,还是通过建议的多任务训练损失,而不需要完整的上下文信息。具体来说,我们表明,我们提出的模型优于大小和延迟匹配的波形形成器,一个国家的最先进的实时TSE模型。
摘要:Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the low-latency causal processing constraints imposed by applications in live-streamed content such as augmented hearing. We introduce a family of context-aware low-latency causal TSE models suitable for real-time processing. First, we explore the utility of context by providing the TSE model with oracle information about what sound classes make up the input mixture, where the objective of the model is to extract one or more sources of interest indicated by the user. Since the practical applications of oracle models are limited due to their assumptions, we introduce a composite multi-task training objective involving separation and classification losses. Our evaluation involving single- and multi-source extraction shows the benefit of using context information in the model either by means of providing full context or via the proposed multi-task training loss without the need for full context information. Specifically, we show that our proposed model outperforms size- and latency-matched Waveformer, a state-of-the-art model for real-time TSE.

【3】 AdaProj: Adaptively Scaled Angular Margin Subspace Projections for  Anomalous Sound Detection with Auxiliary Classification Tasks
标题:AdaProj:自适应缩放角余量子空间投影,用于辅助分类任务的异常声音检测
链接:https://arxiv.org/abs/2403.14179
作者:Kevin Wilkinghoff
摘要:用于半监督异常声音检测的最新方法是首先通过使用基于Meta信息或自监督学习的辅助分类任务来学习嵌入空间,然后估计正常数据的分布。在这项工作中,AdaProj一个新的损失函数。与常用的角度余量损失(angular margin loss)相比,AdaProj学习将数据投影到特定于类的子空间上,角度余量损失将每个类的数据投影到尽可能接近其对应的类中心的位置。通过这样做,属于正常数据的嵌入的结果分布不需要像其他损失函数那样具有限制性,从而允许更详细地查看数据。在DCASE2022和DCASE2023数据集上进行的实验表明,使用AdaProj学习嵌入空间的性能明显优于其他常用的损失函数,并在DCASE2023数据集上获得了最先进的性能。
摘要:The state-of-the-art approach for semi-supervised anomalous sound detection is to first learn an embedding space by using auxiliary classification tasks based on meta information or self-supervised learning and then estimate the distribution of normal data. In this work, AdaProj a novel loss function is presented. In contrast to commonly used angular margin losses, which project data of each class as close as possible to their corresponding class centers, AdaProj learns to project data onto class-specific subspaces. By doing so, the resulting distributions of embeddings belonging to normal data are not required to be as restrictive as other loss functions allowing a more detailed view on the data. In experiments conducted on the DCASE2022 and DCASE2023 datasets, it is shown that using AdaProj to learn an embedding space significantly outperforms other commonly used loss functions and results in a state-of-the-art performance on the DCASE2023 dataset.


【4】 A Multimodal Approach to Device-Directed Speech Detection with Large  Language Models
标题:基于大语言模型的面向设备语音检测的多模态方法
链接:https://arxiv.org/abs/2403.14438
作者:Dominik Wager,Alexander Churchill,Siddharth Sigtia,Panayiotis Georgiou,Matt Mirsamadi,Aarshee Mishra,Erik Marchi
备注:arXiv admin note: text overlap with arXiv:2312.03632
摘要:与虚拟助理的交互通常以预定义的触发短语开始,然后是用户命令。为了使与助手的交互更加直观,我们探索是否可以放弃用户必须以触发短语开始每个命令的要求。我们探索这个任务在三个方面:首先,我们训练分类器只使用从音频波形获得的声学信息。其次,我们采取的解码器输出的自动语音识别(ASR)系统,如1-最好的假设,作为输入功能的大型语言模型(LLM)。最后,我们探讨了一个多模态系统,结合声学和词汇功能,以及ASR解码器信号在LLM。使用多模态信息产生相对相等的错误率提高了39%和61%的纯文本和纯音频模型。增加LLM的大小和使用低秩自适应的训练导致我们的数据集上的相对EER进一步降低高达18%。
摘要:Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users must begin each command with a trigger phrase. We explore this task in three ways: First, we train classifiers using only acoustic information obtained from the audio waveform. Second, we take the decoder outputs of an automatic speech recognition (ASR) system, such as 1-best hypotheses, as input features to a large language model (LLM). Finally, we explore a multimodal system that combines acoustic and lexical features, as well as ASR decoder signals in an LLM. Using multimodal information yields relative equal-error-rate improvements over text-only and audio-only models of up to 39% and 61%. Increasing the size of the LLM and training with low-rank adaption leads to further relative EER reductions of up to 18% on our dataset.


【5】 XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for  Noise-Robust Speech Perception
标题:XLAVS—R:跨语言音视频语音表示学习的噪声鲁棒语音感知
链接:https://arxiv.org/abs/2403.14402
作者:HyoJung Han,Mohamed Anwar,Juan Pino,Wei-Ning Hsu,Marine Carpuat,Bowen Shi,Changhan Wang
摘要:语音识别和翻译系统在现实环境中经常出现的噪声输入上表现不佳。用视觉信号增强这些系统有可能提高对噪声的鲁棒性。但是,视听数据的数量有限,而且语种少于纯视听资源。为了解决这一差距,我们提出了XLAVS-R,一个跨语言的视听语音表示模型,用于100多种语言的噪声鲁棒语音识别和翻译。它旨在通过建立在仅音频的多语言预训练之上并简化现有的预训练方案,最大限度地发挥有限的多语言AV预训练数据的优势。对MuAViC基准测试的广泛评估显示了XLAVS-R在下游视听语音识别和翻译任务上的实力,在有噪声的AV输入下,它的WER和BLEU性能比之前的最先进水平高出18.5%,并通过仅音频微调实现了强大的zero-shot视听能力。
摘要:Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. However, audio-visual (AV) data is only available in limited amounts and for fewer languages than audio-only resources. To address this gap, we present XLAVS-R, a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. It is designed to maximize the benefits of limited multilingual AV pre-training data, by building on top of audio-only multilingual pre-training and simplifying existing pre-training schemes. Extensive evaluation on the MuAViC benchmark shows the strength of XLAVS-R on downstream audio-visual speech recognition and translation tasks, where it outperforms the previous state of the art by up to 18.5% WER and 4.7 BLEU given noisy AV inputs, and enables strong zero-shot audio-visual ability with audio-only fine-tuning.

【6】 emoDARTS: Joint Optimisation of CNN & Sequential Neural Network  Architectures for Superior Speech Emotion Recognition
标题:emoDARTS:CNN和序列神经网络架构的联合优化,用于卓越语音情感识别
链接:https://arxiv.org/abs/2403.14083
作者:Thejan Rajapakshe,Rajib Rana,Sara Khalifa,Berrak Sisman,Bjorn W. Schuller,Carlos Busso
备注:Submitted to IEEE Transactions on Affective Computing on February 19, 2024. arXiv admin note: text overlap with arXiv:2305.14402
摘要:语音情感识别(SER)对于使计算机能够理解人类交流中传达的情感至关重要。随着深度学习(DL)的最新进展,SER模型的性能得到了显着提高。然而,设计最佳的DL架构需要专业知识和实验评估。幸运的是,神经架构搜索(NAS)为自动确定最佳DL模型提供了一种潜在的解决方案。微分结构搜索(DARTS)是一种发现最优模型的特别有效的方法。这项研究提出了emoDARTS,这是一种DARTS优化的联合CNN和顺序神经网络(SeqNN:LSTM,RNN)架构,可增强SER性能。文献支持选择CNN和LSTM耦合来提高性能。虽然DARTS以前曾用于独立选择CNN和LSTM操作,但我们的技术增加了一种新的机制,用于结合使用DARTS选择CNN和SeqNN操作。与早期的工作不同,我们不对CNN的层顺序施加限制。相反,我们让DARTS选择DARTS单元内的最佳层顺序。我们证明了emoDARTS优于传统设计的CNN-LSTM模型,并通过在IEMOCAP,MSP-IMPROV和MSP-Podcast数据集上评估我们的方法,超越了通过DARTS在CNN-LSTM上实现的最佳报告SER结果。
摘要:Speech Emotion Recognition (SER) is crucial for enabling computers to understand the emotions conveyed in human communication. With recent advancements in Deep Learning (DL), the performance of SER models has significantly improved. However, designing an optimal DL architecture requires specialised knowledge and experimental assessments. Fortunately, Neural Architecture Search (NAS) provides a potential solution for automatically determining the best DL model. The Differentiable Architecture Search (DARTS) is a particularly efficient method for discovering optimal models. This study presents emoDARTS, a DARTS-optimised joint CNN and Sequential Neural Network (SeqNN: LSTM, RNN) architecture that enhances SER performance. The literature supports the selection of CNN and LSTM coupling to improve performance.  While DARTS has previously been used to choose CNN and LSTM operations independently, our technique adds a novel mechanism for selecting CNN and SeqNN operations in conjunction using DARTS. Unlike earlier work, we do not impose limits on the layer order of the CNN. Instead, we let DARTS choose the best layer order inside the DARTS cell. We demonstrate that emoDARTS outperforms conventionally designed CNN-LSTM models and surpasses the best-reported SER results achieved through DARTS on CNN-LSTM by evaluating our approach on the IEMOCAP, MSP-IMPROV, and MSP-Podcast datasets.


【7】 The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio  Benchmarks and Novel Data
标题:面向音频研讨会的NeurIPS 2023机器学习:情感音频基准和新数据
链接:https://arxiv.org/abs/2403.14048
作者:Alice Baird,Rachel Manzelli,Panagiotis Tzirakis,Chris Gagne,Haoqi Li,Sadie Allen,Sander Dieleman,Brian Kulis,Shrikanth S. Narayanan,Alan Cowen
摘要:NeurIPS 2023 Machine Learning for Audio研讨会汇集了来自各个音频领域的机器学习(ML)专家。有几个有价值的音频驱动的ML任务,从语音情感识别到音频事件检测,但与其他ML领域相比,社区是稀疏的,例如,计算机视觉或自然语言处理。音频的一个主要限制是可用的数据;由于音频是一种时间依赖的模式,高质量的数据收集既耗时又昂贵,这使得学术团体将其通常最先进的策略应用于更大,更普遍的数据集变得具有挑战性。在这篇简短的白皮书中,为了鼓励那些对大型数据集访问有限的研究人员,组织者首先概述了几个可供社区使用的开源数据集,并在研讨会期间提供了几个适当的数据集。也就是说,三个声音数据集,Hume-Prosody,Hume-VocalBurst,一个表演情感语音数据集Modulate-Sonata和一个游戏中的流数据集Modulate-Stream。我们概述了这些数据集的当前基线,但鼓励来自音频领域的研究人员在初始基线任务之外使用它们。
摘要:The NeurIPS 2023 Machine Learning for Audio Workshop brings together machine learning (ML) experts from various audio domains. There are several valuable audio-driven ML tasks, from speech emotion recognition to audio event detection, but the community is sparse compared to other ML areas, e.g., computer vision or natural language processing. A major limitation with audio is the available data; with audio being a time-dependent modality, high-quality data collection is time-consuming and costly, making it challenging for academic groups to apply their often state-of-the-art strategies to a larger, more generalizable dataset. In this short white paper, to encourage researchers with limited access to large-datasets, the organizers first outline several open-source datasets that are available to the community, and for the duration of the workshop are making several propriety datasets available. Namely, three vocal datasets, Hume-Prosody, Hume-VocalBurst, an acted emotional speech dataset Modulate-Sonata, and an in-game streamer dataset Modulate-Stream. We outline the current baselines on these datasets but encourage researchers from across audio to utilize them outside of the initial baseline tasks.


机器翻译由腾讯交互翻译提供,仅供参考