今日论文合集:cs.SD语音7篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
标题: 音频蕴含:评估音频理解的演绎推理
作者:Soham Deshmukh,Shuo Han,Hazim Bukhari,Benjamin Elizalde,Hannes Gamper,Rita Singh,Bhiksha Raj
链接:点击下载PDF文件
摘要:最近的文献使用语言来建立音频的基础模型。这些音频语言模型(ALM)在大量的音频文本对上进行了训练,并在文本到音频检索,字幕和提问等任务中表现出出色的性能。然而,他们从事更复杂的开放式任务的能力,如交互式推理,需要精通逻辑推理-一项尚未基准的技能。我们引入了新的任务音频蕴涵来评估一个ALM的演绎推理能力。该任务评估音频内容的文本描述(假设)是否可以从音频记录(前提)中推导出来,根据证据的充分性,潜在的结论是蕴涵,中性或矛盾。我们为此任务创建了两个数据集,其中音频记录来自两个音频字幕数据集- AudioCaps和Clotho -以及使用大型语言模型(LLM)生成的假设。我们对最先进的ALM进行基准测试,发现zero-shot和线性探针评估在逻辑推理方面存在缺陷。最后,我们提出了“标题前的原因”,一个中间步骤的字幕,提高了zero-shot和线性探测性能的ALM的绝对6%和3%,分别。摘要:Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage in more complex open-ended tasks, like Interactive Question-Answering, requires proficiency in logical reasoning -- a skill not yet benchmarked. We introduce the novel task of Audio Entailment to evaluate an ALM's deductive reasoning ability. This task assesses whether a text description (hypothesis) of audio content can be deduced from an audio recording (premise), with potential conclusions being entailment, neutral, or contradiction, depending on the sufficiency of the evidence. We create two datasets for this task with audio recordings sourced from two audio captioning datasets -- AudioCaps and Clotho -- and hypotheses generated using Large Language Models (LLMs). We benchmark state-of-the-art ALMs and find deficiencies in logical reasoning with both zero-shot and linear probe evaluations. Finally, we propose "caption-before-reason", an intermediate step of captioning that improves the zero-shot and linear-probe performance of ALMs by an absolute 6% and 3%, respectively.

【2】 I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
标题: 我能听但不能读:仪器识别双塔多模式系统的评估
作者:Yannis Vasilakis,Rachel Bittner,Johan Pauwels
备注:Accepted to ISMIR 2024
链接:点击下载PDF文件
摘要:音乐双塔多模态系统将音频和文本模态集成到联合音频-文本空间中,从而实现歌曲及其相应标签之间的直接比较。这些系统使分类和检索的新方法,利用这两种形式。尽管他们在zero-shot分类和检索任务中显示了有希望的结果,但需要对嵌入进行更仔细的检查。本文评估了固有的zero-shot属性的联合音频文本空间的乐器识别的案例研究。我们提出了一个评估和分析的双塔系统的zero-shot仪器识别和详细分析的属性的预联合和联合嵌入空间。我们的研究结果表明,单独的音频编码器表现出良好的质量,而挑战仍然存在于文本编码器或联合空间投影。具体来说,双塔系统表现出对特定单词的敏感性,更喜欢通用提示而不是音乐提示。尽管文本编码器的规模很大,但它们还没有利用额外的文本上下文或从它们的描述中准确地推断乐器。最后,提出了一种利用工具本体来量化文本空间的语义意义的新方法。这种方法揭示了系统对乐器理解的缺陷,并提供了需要对音乐数据进行微调的文本编码器的证据。摘要:Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for classification and retrieval, leveraging both modalities. Despite the promising results they have shown for zero-shot classification and retrieval tasks, closer inspection of the embeddings is needed. This paper evaluates the inherent zero-shot properties of joint audio-text spaces for the case-study of instrument recognition. We present an evaluation and analysis of two-tower systems for zero-shot instrument recognition and a detailed analysis of the properties of the pre-joint and joint embeddings spaces. Our findings suggest that audio encoders alone demonstrate good quality, while challenges remain within the text encoder or joint space projection. Specifically, two-tower systems exhibit sensitivity towards specific words, favoring generic prompts over musically informed ones. Despite the large size of textual encoders, they do not yet leverage additional textual context or infer instruments accurately from their descriptions. Lastly, a novel approach for quantifying the semantic meaningfulness of the textual space leveraging an instrument ontology is proposed. This method reveals deficiencies in the systems' understanding of instruments and provides evidence of the need for fine-tuning text encoders on musical data.

【3】 On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
标题: 不同自动语音识别体系结构的动态合成训练数据的影响
作者:Nick Rossenbach,Benedikt Hilmes,Ralf Schlüter
备注:Accepted at the SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:在这项工作中,我们评估的实用程序的训练自动语音识别(ASR)的合成数据。我们使用ASR训练数据来训练类似于FastSpeech-2的文本到语音(TTS)系统。有了这个TTS,我们再现了原始的训练数据,仅在合成数据上训练ASR系统。对于ASR,我们使用了三种不同的架构,基于注意力的编码器-解码器,混合深度神经网络隐马尔可夫模型和高斯混合隐马尔可夫模型,显示了模型对合成数据生成的不同敏感性。为了扩展以前的工作,我们提出了一些消融研究的有效性合成与真实的训练数据的ASR。特别是,我们专注于如何通过改变说话人嵌入或缩放模型大小来改变合成数据和真实数据之间的训练差距。对于后者,我们表明TTS模型泛化良好,即使训练分数表明过拟合。摘要:In this work we evaluate the utility of synthetic data for training automatic speech recognition (ASR). We use the ASR training data to train a text-to-speech (TTS) system similar to FastSpeech-2. With this TTS we reproduce the original training data, training ASR systems solely on synthetic data. For ASR, we use three different architectures, attention-based encoder-decoder, hybrid deep neural network hidden Markov model and a Gaussian mixture hidden Markov model, showing the different sensitivity of the models to synthetic data generation. In order to extend previous work, we present a number of ablation studies on the effectiveness of synthetic vs. real training data for ASR. In particular we focus on how the gap between training on synthetic and real data changes by varying the speaker embedding or by scaling the model size. For the latter we show that the TTS models generalize well, even when training scores indicate overfitting.

【4】 Innovative Speech-Based Deep Learning Approaches for Parkinson's Disease Classification: A Systematic Review
标题: 基于语音的创新深度学习方法用于帕金森病分类:系统性评论
作者:Lisanne van Gelderen,Cristian Tejedor-García
备注:Submitted in Applied Sciences - peer reviewed Open Access journal. This research was funded by the NWO research programme AiNed Fellowship Grants under the project Responsible AI for Voice Diagnostics (RAIVD) - grant number NGF.1607.22.013
链接:点击下载PDF文件
摘要:帕金森病(PD)是世界范围内第二大流行的神经退行性疾病,经常表现为早期言语障碍。人工智能(AI)的最新进展,特别是深度学习(DL),通过分析语音数据显着增强了PD诊断。然而,研究的进展受到有限的公开访问的基于语音的PD数据集的限制,主要是由于隐私和伦理问题。本综述涵盖了最新的基于DL的人工智能方法,用于基于语音的PD分类,重点关注2020年至2024年3月期间发表的33篇科学著作的性能、可用资源和相关挑战。这些深度学习方法分为端到端(E2E)学习、迁移学习(TL)和深度声学特征(E2E)提取。在E2E方法中,卷积神经网络(CNN)是普遍的,尽管Transformers越来越受欢迎。E2E方法面临诸如有限的数据和计算资源的挑战,特别是对于Transformers。TL通过提供更强大的PD诊断和更好的跨语言通用性来解决这些问题。特征提取旨在通过检查深度特征对其他DL方法和更传统的机器学习(ML)方法的具体影响来提高结果的可解释性和可解释性。然而,与E2E和TL方法相比,它往往表现不佳。这篇综述还讨论了与偏见、可解释性和隐私有关的未解决问题,强调了未来研究的必要性。摘要:Parkinson's disease (PD), the second most prevalent neurodegenerative disorder worldwide, frequently presents with early-stage speech impairments. Recent advancements in Artificial Intelligence (AI), particularly deep learning (DL), have significantly enhanced PD diagnosis through the analysis of speech data. Nevertheless, the progress of research is restricted by the limited availability of publicly accessible speech-based PD datasets, primarily due to privacy and ethical concerns. This review covers the latest DL-based AI approaches for speech-based PD classification, focusing on performance, available resources and associated challenges of 33 scientific works published between 2020 and March 2024. These DL approaches are categorized into end-to-end (E2E) learning, transfer learning (TL) and deep acoustic features (DAF) extraction. Among E2E approaches, Convolutional Neural Networks (CNNs) are prevalent, though Transformers are increasingly popular. E2E approaches face challenges such as limited data and computational resources, especially with Transformers. TL addresses these issues by providing more robust PD diagnosis and better generalizability across languages. DAF extraction aims to improve the explainability and interpretability of results by examining the specific effects of deep features on both other DL approaches and more traditional machine learning (ML) methods. However, it often underperforms compared to E2E and TL approaches. This review also discusses unresolved issues related to bias, explainability and privacy, highlighting the need for future research.

【5】 Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
标题: 描述您在哪里:通过环境的文本描述提高语音情感识别的噪音稳健性
作者:Seong-Gyun Leem,Daniel Fulford,Jukka-Pekka Onnela,David Gard,Carlos Busso
链接:点击下载PDF文件
摘要:语音情感识别(SER)系统通常在现实环境中挣扎,其中环境噪声严重降低其性能。本文探讨了一种新的方法,利用先验知识的测试环境,以最大限度地提高SER性能在嘈杂的条件下。为了解决这个问题,我们提出了一个文本引导的,环境感知的训练,其中SER模型是用受污染的语音样本及其配对的噪声描述来训练的。我们使用一个预先训练的文本编码器来提取基于文本的环境嵌入,然后在训练和推理过程中将其融合到基于transformer的SER模型中。我们证明了我们的方法的有效性,通过我们的实验与MSP播客语料库和真实世界的加性噪声样本收集的Freesound存储库。我们的实验表明,基于文本的环境描述处理的大型语言模型(LLM)产生的表示,提高噪声鲁棒性的SER系统。此外,我们提出的LLM方法比我们的环境不可知基线产生更好的性能,特别是在低信噪比(SNR)条件下。当在-5dB SNR水平下测试时,我们提出的方法比我们最好的基线模型表现出更好的性能,分别为31.8%(唤醒),23.5%(优势)和9.5%(效价)。摘要:Speech emotion recognition (SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach that exploits prior knowledge of testing environments to maximize SER performance under noisy conditions. To address this task, we propose a text-guided, environment-aware training where an SER model is trained with contaminated speech samples and their paired noise description. We use a pre-trained text encoder to extract the text-based environment embedding and then fuse it to a transformer-based SER model during training and inference. We demonstrate the effectiveness of our approach through our experiment with the MSP-Podcast corpus and real-world additive noise samples collected from the Freesound repository. Our experiment indicates that the text-based environment descriptions processed by a large language model (LLM) produce representations that improve the noise-robustness of the SER system. In addition, our proposed approach with an LLM yields better performance than our environment-agnostic baselines, especially in low signal-to-noise ratio (SNR) conditions. When testing at -5dB SNR level, our proposed method shows better performance than our best baseline model by 31.8 % (arousal), 23.5% (dominance), and 9.5% (valence).

【6】 Coupling Speech Encoders with Downstream Text Models
标题: 将语音编码器与下游文本模型相结合
作者:Ciprian Chelba,Johan Schalkwyk
链接:点击下载PDF文件
摘要:我们提出了一种模块化方法来构建级联语音翻译(AST)模型,该模型可以保证生成的模型的性能不低于1-best级联基线,同时保留给定任务的最先进语音识别(ASR)和文本翻译(MT)性能。我们的新贡献是使用了一个在L2损失下训练的“导出器”层,以确保ASR嵌入和MT令牌嵌入之间的强匹配。“导出器”输出嵌入被直接馈送到MT模型,以代替1-best令牌嵌入,从而保证所得到的模型的性能不差于1-best级联基线,同时允许反向传播梯度从MT模型流入ASR组件。匹配嵌入级联架构在MT模型的增量训练不是一种选择的情况下提供了对其1-最佳对应物的显着改进,但我们试图通过利用AST任务提供的(语音,转录,翻译转录)数据来提高质量。当MT模型在AST任务可用的并行文本数据上进行增量训练时,增益消失。该方法有望用于寻求将ASR编码器和不可变文本模型耦合的其他场景,例如大型语言模型(LLM)。摘要:We present a modular approach to building cascade speech translation (AST) models that guarantees that the resulting model performs no worse than the 1-best cascade baseline while preserving state-of-the-art speech recognition (ASR) and text translation (MT) performance for a given task. Our novel contribution is the use of an exporter'' layer that is trained under L2-loss to ensure a strong match between ASR embeddings and the MT token embeddings for the 1-best sequence. The exporter'' output embeddings are fed directly to the MT model in lieu of 1-best token embeddings, thus guaranteeing that the resulting model performs no worse than the 1-best cascade baseline, while allowing back-propagation gradient to flow from the MT model into the ASR components. The matched-embeddings cascade architecture provide a significant improvement over its 1-best counterpart in scenarios where incremental training of the MT model is not an option and yet we seek to improve quality by leveraging (speech, transcription, translated transcription) data provided with the AST task. The gain disappears when the MT model is incrementally trained on the parallel text data available with the AST task. The approach holds promise for other scenarios that seek to couple ASR encoders and immutable text models, such at large language models (LLM).

【7】 Improved symbolic drum style classification with grammar-based hierarchical representations
标题: 使用基于语法的分层表示改进的符号鼓风格分类
作者:Léo Géré,Philippe Rigaux,Nicolas Audebert
备注:International Society for Music Information Retrieval Conference 2024, Nov 2024, San Francisco, United States
链接:点击下载PDF文件
摘要:深度学习模型已经成为音乐数据分析和分类的重要工具。这些模型要么对音频信号(例如波形或频谱图)进行操作,要么对符号表示(例如,音频)进行操作。在后者中,音乐信息通常被简化为基本特征,即持续时间,音高和速度。大多数现有的作品依赖于经典自然语言处理或矩阵表示(例如钢琴卷)的通用标记化策略。在这项工作中,我们评估了符号数据的丰富表示如何影响音乐风格分类的深度模型,即Transformers和RNN。特别是,我们研究表示,明确纳入音乐信息隐含在MIDI类编码,如节奏组织,并表明他们优于通用标记化策略。我们引入了一个新的基于树的表示的数据建立在一个上下文无关的音乐语法。我们表明,这种语法表示准确地编码高层次的节奏信息,并优于现有的编码的Groundwise数据集的鼓点风格分类,同时更紧凑和参数效率。摘要:Deep learning models have become a critical tool for analysis and classification of musical data. These models operate either on the audio signal, e.g. waveform or spectrogram, or on a symbolic representation, such as MIDI. In the latter, musical information is often reduced to basic features, i.e. durations, pitches and velocities. Most existing works then rely on generic tokenization strategies from classical natural language processing, or matrix representations, e.g. piano roll. In this work, we evaluate how enriched representations of symbolic data can impact deep models, i.e. Transformers and RNN, for music style classification. In particular, we examine representations that explicitly incorporate musical information implicitly present in MIDI-like encodings, such as rhythmic organization, and show that they outperform generic tokenization strategies. We introduce a new tree-based representation of MIDI data built upon a context-free musical grammar. We show that this grammar representation accurately encodes high-level rhythmic information and outperforms existing encodings on the GrooveMIDI Dataset for drumming style classification, while being more compact and parameter-efficient.


eess.AS音频处理
【1】 Reshape Dimensions Network for Speaker Recognition
标题: 重塑语音识别维度网络
作者:Ivan Yakovlev,Rostislav Makarov,Andrei Balykin,Pavel Malov,Anton Okhotnikov,Nikita Torgashov
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了重塑维度网络(ReDimNet),一种新的神经网络架构,用于提取话语级说话人表示。我们的方法利用2D特征映射到1D信号表示的维度重塑,反之亦然,从而实现1D和2D块的联合使用。我们提出了一个原始的网络拓扑结构,保留了1D和2D块的通道时间步长频率输出的体积,有利于有效的残差特征图聚合。此外,ReDimNet是有效的可扩展性,我们引入了一系列的模型大小,从1到15 M的参数和从0.5到20 GMAC。我们的实验结果表明,ReDimNet在说话人识别中实现了最先进的性能,同时降低了计算复杂度和模型参数的数量。摘要:In this paper, we present Reshape Dimensions Network (ReDimNet), a novel neural network architecture for extracting utterance-level speaker representations. Our approach leverages dimensionality reshaping of 2D feature maps to 1D signal representation and vice versa, enabling the joint usage of 1D and 2D blocks. We propose an original network topology that preserves the volume of channel-timestep-frequency outputs of 1D and 2D blocks, facilitating efficient residual feature maps aggregation. Moreover, ReDimNet is efficiently scalable, and we introduce a range of model sizes, varying from 1 to 15 M parameters and from 0.5 to 20 GMACs. Our experimental results demonstrate that ReDimNet achieves state-of-the-art performance in speaker recognition while reducing computational complexity and the number of model parameters.

【2】 Detection of manatee vocalisations using the Audio Spectrogram Transformer
标题: 使用音频频谱图Transformer检测海牛的叫声
作者:Stefano Schiappacasse,Taco de Wolff,Yann Henaut,Regina Cervera,Aviva Charles,Felipe Tobar
备注:Accepted at MLSP 2024
链接:点击下载PDF文件
摘要:安的列斯海牛(Trichechus manatus)是一种濒临灭绝的食草性水生哺乳动物,其作为生态平衡者和保护伞物种的作用强调了其保护的重要性。监测海牛种群的一种创新方法是被动声学监测(PAM),其中从潜艇音频中提取发声。我们提出了一种新的端到端的方法来检测海牛发声音频频谱图Transformer(AST)的建设。本着迁移学习的精神,我们通过重新设计其滤波器组和调整包含部分阳性标签的真实数据集来微调AST以检测海牛呼叫。我们的实验评估揭示了所提出的模型的两个关键特征:i)它的性能与最先进的技术水平相当,而不需要手动调整去噪或检测阶段,ii)它可以成功地识别训练数据集中错过的发声,从而减少专家生物声学标记的工作量。这项工作是一个初步的相关步骤,开发新的,用户友好的工具,保护不同种类的海牛。摘要:The Antillean manatee ( emph{Trichechus manatus}) is an endangered herbivorous aquatic mammal whose role as an ecological balancer and umbrella species underscores the importance of its conservation. An innovative approach to monitor manatee populations is passive acoustic monitoring (PAM), where vocalisations are extracted from submarine audio. We propose a novel end-to-end approach to detect manatee vocalisations building on the Audio Spectrogram Transformer (AST). In a transfer learning spirit, we fine-tune AST to detect manatee calls by redesigning its filterbanks and adapting a real-world dataset containing partial positive labels. Our experimental evaluation reveals the two key features of the proposed model: i) it performs on par with the state of the art without requiring hand-tuned denoising or detection stages, and ii) it can successfully identify missed vocalisations in the training dataset, thus reducing the workload of expert bioacoustic labellers. This work is a preliminary relevant step to develop novel, user-friendly tools for the conservation of the different species of manatees.

【3】 Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
标题: 带Keynote说话人二元化的多阶段脸音关联学习
作者:Ruijie Tao,Zhan Shi,Yidi Jiang,Duc-Tuan Truong,Eng-Siong Chng,Massimo Alioto,Haizhou Li
链接:点击下载PDF文件
摘要:人类大脑有能力通过利用他们的一般关系将未知的人的声音和面部联系起来,称为“跨模态说话者验证”。由于模式之间的复杂关系,这项任务构成了重大挑战。在本文中,我们提出了一个“多阶段的人脸语音联想学习与基调发言人拨号”~(MFV-KSD)框架。MFV-KSD包含一个基调发言人日记前端,以有效地解决嘈杂的语音输入问题。为了平衡和增强模态内特征学习和模态间相关性理解,MFV-KSD采用了一种新的三阶段训练策略。我们的实验结果证明了稳健的性能,在2024年多语言环境中的人脸语音关联(FAME)挑战中获得第一名,总体等错误率(EER)为19.9%。详细信息可在https: github.com TaoRuijie MFV-KSD上找到。摘要:The human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as cross-modal speaker verification''. This task poses significant challenges due to the complex relationship between the modalities. In this paper, we propose a Multi-stage Face-voice Association Learning with Keynote Speaker Diarization''~(MFV-KSD) framework. MFV-KSD contains a keynote speaker diarization front-end to effectively address the noisy speech inputs issue. To balance and enhance the intra-modal feature learning and inter-modal correlation understanding, MFV-KSD utilizes a novel three-stage training strategy. Our experimental results demonstrated robust performance, achieving the first rank in the 2024 Face-voice Association in Multilingual Environments (FAME) challenge with an overall Equal Error Rate (EER) of 19.9%. Details can be found in https: github.com TaoRuijie MFV-KSD.

【4】 Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
标题: 音频蕴含:评估音频理解的演绎推理
作者:Soham Deshmukh,Shuo Han,Hazim Bukhari,Benjamin Elizalde,Hannes Gamper,Rita Singh,Bhiksha Raj
链接:点击下载PDF文件
摘要:最近的文献使用语言来建立音频的基础模型。这些音频语言模型(ALM)在大量的音频文本对上进行了训练,并在文本到音频检索,字幕和提问等任务中表现出出色的性能。然而,他们从事更复杂的开放式任务的能力,如交互式推理,需要精通逻辑推理-一项尚未基准的技能。我们引入了新的任务音频蕴涵来评估一个ALM的演绎推理能力。该任务评估音频内容的文本描述(假设)是否可以从音频记录(前提)中推导出来,根据证据的充分性,潜在的结论是蕴涵,中性或矛盾。我们为此任务创建了两个数据集,其中音频记录来自两个音频字幕数据集- AudioCaps和Clotho -以及使用大型语言模型(LLM)生成的假设。我们对最先进的ALM进行基准测试,发现zero-shot和线性探针评估在逻辑推理方面存在缺陷。最后,我们提出了“标题前的原因”,一个中间步骤的字幕,提高了zero-shot和线性探测性能的ALM的绝对6%和3%,分别。摘要:Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage in more complex open-ended tasks, like Interactive Question-Answering, requires proficiency in logical reasoning -- a skill not yet benchmarked. We introduce the novel task of Audio Entailment to evaluate an ALM's deductive reasoning ability. This task assesses whether a text description (hypothesis) of audio content can be deduced from an audio recording (premise), with potential conclusions being entailment, neutral, or contradiction, depending on the sufficiency of the evidence. We create two datasets for this task with audio recordings sourced from two audio captioning datasets -- AudioCaps and Clotho -- and hypotheses generated using Large Language Models (LLMs). We benchmark state-of-the-art ALMs and find deficiencies in logical reasoning with both zero-shot and linear probe evaluations. Finally, we propose "caption-before-reason", an intermediate step of captioning that improves the zero-shot and linear-probe performance of ALMs by an absolute 6% and 3%, respectively.

【5】 I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
标题: 我能听但不能读:仪器识别双塔多模式系统的评估
作者:Yannis Vasilakis,Rachel Bittner,Johan Pauwels
备注:Accepted to ISMIR 2024
链接:点击下载PDF文件
摘要:音乐双塔多模态系统将音频和文本模态集成到联合音频-文本空间中,从而实现歌曲及其相应标签之间的直接比较。这些系统使分类和检索的新方法,利用这两种形式。尽管他们在zero-shot分类和检索任务中显示了有希望的结果,但需要对嵌入进行更仔细的检查。本文评估了固有的zero-shot属性的联合音频文本空间的乐器识别的案例研究。我们提出了一个评估和分析的双塔系统的zero-shot仪器识别和详细分析的属性的预联合和联合嵌入空间。我们的研究结果表明,单独的音频编码器表现出良好的质量,而挑战仍然存在于文本编码器或联合空间投影。具体来说,双塔系统表现出对特定单词的敏感性,更喜欢通用提示而不是音乐提示。尽管文本编码器的规模很大,但它们还没有利用额外的文本上下文或从它们的描述中准确地推断乐器。最后,提出了一种利用工具本体来量化文本空间的语义意义的新方法。这种方法揭示了系统对乐器理解的缺陷,并提供了需要对音乐数据进行微调的文本编码器的证据。摘要:Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for classification and retrieval, leveraging both modalities. Despite the promising results they have shown for zero-shot classification and retrieval tasks, closer inspection of the embeddings is needed. This paper evaluates the inherent zero-shot properties of joint audio-text spaces for the case-study of instrument recognition. We present an evaluation and analysis of two-tower systems for zero-shot instrument recognition and a detailed analysis of the properties of the pre-joint and joint embeddings spaces. Our findings suggest that audio encoders alone demonstrate good quality, while challenges remain within the text encoder or joint space projection. Specifically, two-tower systems exhibit sensitivity towards specific words, favoring generic prompts over musically informed ones. Despite the large size of textual encoders, they do not yet leverage additional textual context or infer instruments accurately from their descriptions. Lastly, a novel approach for quantifying the semantic meaningfulness of the textual space leveraging an instrument ontology is proposed. This method reveals deficiencies in the systems' understanding of instruments and provides evidence of the need for fine-tuning text encoders on musical data.

【6】 On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
标题: 不同自动语音识别体系结构的动态合成训练数据的影响
作者:Nick Rossenbach,Benedikt Hilmes,Ralf Schlüter
备注:Accepted at the SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:在这项工作中,我们评估的实用程序的训练自动语音识别(ASR)的合成数据。我们使用ASR训练数据来训练类似于FastSpeech-2的文本到语音(TTS)系统。有了这个TTS,我们再现了原始的训练数据,仅在合成数据上训练ASR系统。对于ASR,我们使用了三种不同的架构,基于注意力的编码器-解码器,混合深度神经网络隐马尔可夫模型和高斯混合隐马尔可夫模型,显示了模型对合成数据生成的不同敏感性。为了扩展以前的工作,我们提出了一些消融研究的有效性合成与真实的训练数据的ASR。特别是,我们专注于如何通过改变说话人嵌入或缩放模型大小来改变合成数据和真实数据之间的训练差距。对于后者,我们表明TTS模型泛化良好,即使训练分数表明过拟合。摘要:In this work we evaluate the utility of synthetic data for training automatic speech recognition (ASR). We use the ASR training data to train a text-to-speech (TTS) system similar to FastSpeech-2. With this TTS we reproduce the original training data, training ASR systems solely on synthetic data. For ASR, we use three different architectures, attention-based encoder-decoder, hybrid deep neural network hidden Markov model and a Gaussian mixture hidden Markov model, showing the different sensitivity of the models to synthetic data generation. In order to extend previous work, we present a number of ablation studies on the effectiveness of synthetic vs. real training data for ASR. In particular we focus on how the gap between training on synthetic and real data changes by varying the speaker embedding or by scaling the model size. For the latter we show that the TTS models generalize well, even when training scores indicate overfitting.

【7】 Innovative Speech-Based Deep Learning Approaches for Parkinson's Disease Classification: A Systematic Review
标题: 基于语音的创新深度学习方法用于帕金森病分类:系统性评论
作者:Lisanne van Gelderen,Cristian Tejedor-García
备注:Submitted in Applied Sciences - peer reviewed Open Access journal. This research was funded by the NWO research programme AiNed Fellowship Grants under the project Responsible AI for Voice Diagnostics (RAIVD) - grant number NGF.1607.22.013
链接:点击下载PDF文件
摘要:帕金森病(PD)是世界范围内第二大流行的神经退行性疾病,经常表现为早期言语障碍。人工智能(AI)的最新进展,特别是深度学习(DL),通过分析语音数据显着增强了PD诊断。然而,研究的进展受到有限的公开访问的基于语音的PD数据集的限制,主要是由于隐私和伦理问题。本综述涵盖了最新的基于DL的人工智能方法,用于基于语音的PD分类,重点关注2020年至2024年3月期间发表的33篇科学著作的性能、可用资源和相关挑战。这些深度学习方法分为端到端(E2E)学习、迁移学习(TL)和深度声学特征(E2E)提取。在E2E方法中,卷积神经网络(CNN)是普遍的,尽管Transformers越来越受欢迎。E2E方法面临诸如有限的数据和计算资源的挑战,特别是对于Transformers。TL通过提供更强大的PD诊断和更好的跨语言通用性来解决这些问题。特征提取旨在通过检查深度特征对其他DL方法和更传统的机器学习(ML)方法的具体影响来提高结果的可解释性和可解释性。然而,与E2E和TL方法相比,它往往表现不佳。这篇综述还讨论了与偏见、可解释性和隐私有关的未解决问题,强调了未来研究的必要性。摘要:Parkinson's disease (PD), the second most prevalent neurodegenerative disorder worldwide, frequently presents with early-stage speech impairments. Recent advancements in Artificial Intelligence (AI), particularly deep learning (DL), have significantly enhanced PD diagnosis through the analysis of speech data. Nevertheless, the progress of research is restricted by the limited availability of publicly accessible speech-based PD datasets, primarily due to privacy and ethical concerns. This review covers the latest DL-based AI approaches for speech-based PD classification, focusing on performance, available resources and associated challenges of 33 scientific works published between 2020 and March 2024. These DL approaches are categorized into end-to-end (E2E) learning, transfer learning (TL) and deep acoustic features (DAF) extraction. Among E2E approaches, Convolutional Neural Networks (CNNs) are prevalent, though Transformers are increasingly popular. E2E approaches face challenges such as limited data and computational resources, especially with Transformers. TL addresses these issues by providing more robust PD diagnosis and better generalizability across languages. DAF extraction aims to improve the explainability and interpretability of results by examining the specific effects of deep features on both other DL approaches and more traditional machine learning (ML) methods. However, it often underperforms compared to E2E and TL approaches. This review also discusses unresolved issues related to bias, explainability and privacy, highlighting the need for future research.

【8】 Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
标题: 描述您在哪里:通过环境的文本描述提高语音情感识别的噪音稳健性
作者:Seong-Gyun Leem,Daniel Fulford,Jukka-Pekka Onnela,David Gard,Carlos Busso
链接:点击下载PDF文件
摘要:语音情感识别(SER)系统通常在现实环境中挣扎,其中环境噪声严重降低其性能。本文探讨了一种新的方法,利用先验知识的测试环境,以最大限度地提高SER性能在嘈杂的条件下。为了解决这个问题,我们提出了一个文本引导的,环境感知的训练,其中SER模型是用受污染的语音样本及其配对的噪声描述来训练的。我们使用预训练的文本编码器来提取基于文本的环境嵌入,然后在训练和推理过程中将其融合到基于变换器的SER模型中。我们证明了我们的方法的有效性,通过我们的实验与MSP播客语料库和真实世界的加性噪声样本收集的Freesound存储库。我们的实验表明,基于文本的环境描述处理的大型语言模型(LLM)产生的表示,提高噪声鲁棒性的SER系统。此外,我们提出的LLM方法比我们的环境不可知基线产生更好的性能,特别是在低信噪比(SNR)条件下。当在-5dB SNR水平下测试时,我们提出的方法比我们最好的基线模型表现出更好的性能,分别为31.8%(唤醒),23.5%(优势)和9.5%(效价)。摘要:Speech emotion recognition (SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach that exploits prior knowledge of testing environments to maximize SER performance under noisy conditions. To address this task, we propose a text-guided, environment-aware training where an SER model is trained with contaminated speech samples and their paired noise description. We use a pre-trained text encoder to extract the text-based environment embedding and then fuse it to a transformer-based SER model during training and inference. We demonstrate the effectiveness of our approach through our experiment with the MSP-Podcast corpus and real-world additive noise samples collected from the Freesound repository. Our experiment indicates that the text-based environment descriptions processed by a large language model (LLM) produce representations that improve the noise-robustness of the SER system. In addition, our proposed approach with an LLM yields better performance than our environment-agnostic baselines, especially in low signal-to-noise ratio (SNR) conditions. When testing at -5dB SNR level, our proposed method shows better performance than our best baseline model by 31.8 % (arousal), 23.5% (dominance), and 9.5% (valence).

【9】 Coupling Speech Encoders with Downstream Text Models
标题: 将语音编码器与下游文本模型相结合
作者:Ciprian Chelba,Johan Schalkwyk
链接:点击下载PDF文件
摘要:我们提出了一种模块化方法来构建级联语音翻译(AST)模型,该模型可以保证生成的模型的性能不低于1-best级联基线,同时保留给定任务的最先进语音识别(ASR)和文本翻译(MT)性能。我们的新贡献是使用了一个在L2损失下训练的“导出器”层,以确保ASR嵌入和MT令牌嵌入之间的强匹配。“导出器”输出嵌入被直接馈送到MT模型,以代替1-best令牌嵌入,从而保证所得到的模型的性能不差于1-best级联基线,同时允许反向传播梯度从MT模型流入ASR组件。匹配嵌入级联架构在MT模型的增量训练不是一种选择的情况下提供了对其1-最佳对应物的显着改进,但我们试图通过利用AST任务提供的(语音,转录,翻译转录)数据来提高质量。当MT模型在AST任务可用的并行文本数据上进行增量训练时,增益就会消失。该方法有望用于寻求将ASR编码器和不可变文本模型耦合的其他场景,例如大型语言模型(LLM)。摘要:We present a modular approach to building cascade speech translation (AST) models that guarantees that the resulting model performs no worse than the 1-best cascade baseline while preserving state-of-the-art speech recognition (ASR) and text translation (MT) performance for a given task. Our novel contribution is the use of an exporter'' layer that is trained under L2-loss to ensure a strong match between ASR embeddings and the MT token embeddings for the 1-best sequence. The exporter'' output embeddings are fed directly to the MT model in lieu of 1-best token embeddings, thus guaranteeing that the resulting model performs no worse than the 1-best cascade baseline, while allowing back-propagation gradient to flow from the MT model into the ASR components. The matched-embeddings cascade architecture provide a significant improvement over its 1-best counterpart in scenarios where incremental training of the MT model is not an option and yet we seek to improve quality by leveraging (speech, transcription, translated transcription) data provided with the AST task. The gain disappears when the MT model is incrementally trained on the parallel text data available with the AST task. The approach holds promise for other scenarios that seek to couple ASR encoders and immutable text models, such at large language models (LLM).

【10】 Improved symbolic drum style classification with grammar-based hierarchical representations
标题: 使用基于语法的分层表示改进的符号鼓风格分类
作者:Léo Géré,Philippe Rigaux,Nicolas Audebert
备注:International Society for Music Information Retrieval Conference 2024, Nov 2024, San Francisco, United States
链接:点击下载PDF文件
摘要:深度学习模型已经成为音乐数据分析和分类的重要工具。这些模型要么对音频信号(例如波形或频谱图)进行操作,要么对符号表示(例如,音频)进行操作。在后者中,音乐信息通常被简化为基本特征,即持续时间,音高和速度。大多数现有的作品依赖于经典自然语言处理或矩阵表示(例如钢琴卷)的通用标记化策略。在这项工作中,我们评估了符号数据的丰富表示如何影响音乐风格分类的深度模型,即Transformers和RNN。特别是,我们研究表示,明确纳入音乐信息隐含在MIDI类编码,如节奏组织,并表明他们优于通用标记化策略。我们引入了一个新的基于树的表示的数据建立在上下文无关的音乐语法。我们表明,这种语法表示准确地编码高层次的节奏信息,并优于现有的编码的Groundwise数据集的鼓点风格分类,同时更紧凑和参数效率。摘要:Deep learning models have become a critical tool for analysis and classification of musical data. These models operate either on the audio signal, e.g. waveform or spectrogram, or on a symbolic representation, such as MIDI. In the latter, musical information is often reduced to basic features, i.e. durations, pitches and velocities. Most existing works then rely on generic tokenization strategies from classical natural language processing, or matrix representations, e.g. piano roll. In this work, we evaluate how enriched representations of symbolic data can impact deep models, i.e. Transformers and RNN, for music style classification. In particular, we examine representations that explicitly incorporate musical information implicitly present in MIDI-like encodings, such as rhythmic organization, and show that they outperform generic tokenization strategies. We introduce a new tree-based representation of MIDI data built upon a context-free musical grammar. We show that this grammar representation accurately encodes high-level rhythmic information and outperforms existing encodings on the GrooveMIDI Dataset for drumming style classification, while being more compact and parameter-efficient.


机器翻译,仅供参考