今日论文合集:cs.SD语音8篇,eess.AS音频处理9篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】Training-Free Deepfake Voice Recognition by Leveraging Large-Scale  Pre-Trained Models
标题:通过利用大规模预训练模型进行免训练Deepfake语音识别
链接:https://arxiv.org/abs/2405.02179
作者:Alessandro Pianese,Davide Cozzolino,Giovanni Poggi,Luisa Verdoliva
摘要:泛化是当前音频deepfake检测器的主要问题,这些检测器很难对分发数据提供可靠的结果。考虑到越来越多的精确合成方法的发展速度,设计出在未经训练的数据上也能很好地工作的技术是非常重要的。在本文中,我们研究了大规模预训练模型用于音频深度伪造检测的潜力,特别关注泛化能力。为此,在说话人验证框架中重新制定了检测问题,并且通过测试下的语音样本与所声称的身份的语音之间的不匹配来暴露假音频。使用这种范式,在训练中不需要假语音样本,从根本上切断了与生成方法的任何联系,并确保了完全的泛化能力。特征由通用的大型预训练模型提取,无需对特定的虚假检测或说话人验证数据集进行训练或微调。在检测时,只需要被测身份的有限语音片段集。在社区中广泛使用的几个数据集上的实验表明,基于预训练模型的检测器具有出色的性能,并显示出强大的泛化能力,可与分布数据上的监督方法相媲美,并在很大程度上克服了分布数据上的监督方法。
摘要:Generalization is a main issue for current audio deepfake detectors, which struggle to provide reliable results on out-of-distribution data. Given the speed at which more and more accurate synthesis methods are developed, it is very important to design techniques that work well also on data they were not trained for.In this paper we study the potential of large-scale pre-trained models for audio deepfake detection, with special focus on generalization ability. To this end, the detection problem is reformulated in a speaker verification framework and fake audios are exposed by the mismatch between the voice sample under test and the voice of the claimed identity. With this paradigm, no fake speech sample is necessary in training, cutting off any link with the generation method at the root, and ensuring full generalization ability. Features are extracted by general-purpose large pre-trained models, with no need for training or fine-tuning on specific fake detection or speaker verification datasets. At detection time only a limited set of voice fragments of the identity under test is required. Experiments on several datasets widespread in the community show that detectors based on pre-trained models achieve excellent performance and show strong generalization ability, rivaling supervised methods on in-distribution data and largely overcoming them on out-of-distribution data.

【2】 GMP-ATL: Gender-augmented Multi-scale Pseudo-label Enhanced Adaptive  Transfer Learning for Speech Emotion Recognition via HuBERT
标题:GMP-ATL:通过HuBERT进行语音情感识别的性别增强多尺度伪标签增强自适应转移学习
链接:https://arxiv.org/abs/2405.02151
作者:Yu Pan,Yuguang Yang,Heng Lu,Lei Ma,Jianjun Zhao
摘要:预先训练的语音模型的不断发展极大地促进了语音情感识别(SER)。然而,这些方法的性能仍有提高的潜力。GMP-ATL(Gender-augmented Multi-scale Pseudo-label Adaptive Transfer Learning)是一种新的基于HuBERT的自适应迁移学习框架,它首先使用预先训练好的HuBERT,通过多任务学习和多尺度k-means聚类来获取帧级的性别增强多尺度伪标签。然后,为了充分利用所获得的帧级和话语级情感标签,我们结合模型再训练和微调方法来进一步优化GMP-ATL。在IEMOCAP上的实验表明,GMP-ATL的WAR和UAR分别为80.0%和82.0%,超过了现有的单模态SER方法,同时也取得了与多模态SER方法相当的识别效果.
摘要:The continuous evolution of pre-trained speech models has greatly advanced Speech Emotion Recognition (SER). However, there is still potential for enhancement in the performance of these methods. In this paper, we present GMP-ATL (Gender-augmented Multi-scale Pseudo-label Adaptive Transfer Learning), a novel HuBERT-based adaptive transfer learning framework for SER. Specifically, GMP-ATL initially employs the pre-trained HuBERT, implementing multi-task learning and multi-scale k-means clustering to acquire frame-level gender-augmented multi-scale pseudo-labels. Then, to fully leverage both obtained frame-level and utterance-level emotion labels, we incorporate model retraining and fine-tuning methods to further optimize GMP-ATL. Experiments on IEMOCAP show that our GMP-ATL achieves superior recognition performance, with a WAR of 80.0\% and a UAR of 82.0\%, surpassing state-of-the-art unimodal SER methods, while also yielding comparable results with multimodal SER approaches.

【3】 Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
标题:在中国开源数据集中揭示基于LLM的ASB的潜力
链接:https://arxiv.org/abs/2405.02132
作者:Xuelong Geng,Tianyi Xu,Kun Wei,Bingsheng Mu,Hongfei Xue,He Wang,Yangze Li,Pengcheng Guo,Yuhang Dai,Longhao Li,Mingchen Shao,Lei Xie
摘要:大型语言模型在各种NLP任务中表现出无与伦比的有效性,将LLM与自动语音识别集成正在成为主流范例。在此基础上,我们的研究深入研究了一个大型开源中文数据集上的这种范式。具体来说,我们的研究旨在评估各种配置的语音编码器,LLM和投影仪模块的语音基础encoderLLM ASR范式的背景下的影响。此外,我们引入了一个三阶段的训练方法,明确开发,以提高模型的能力,使听觉和文本信息。这种方法的实施,以及ASR组件的战略集成,使我们能够在AISHELL1,TestNet和TestMeeting测试集上实现SOTA性能。我们的分析为基于LLM的ASR系统的未来研究提供了实证基础,并为使用中国数据集优化性能提供了见解。我们将公开发布用于数据准备、训练、推理和评分的所有脚本,以及预训练模型和训练日志,以促进可重复的研究。
摘要:Large Language Models have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition is becoming a mainstream paradigm. Building upon this momentum, our research delves into an indepth examination of this paradigm on a large opensource Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoderLLM ASR paradigm. Furthermore, we introduce a threestage training approach, expressly developed to enhance the model's ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL1, TestNet, and TestMeeting test sets. Our analysis presents an empirical foundation for future research in LLMbased ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pretrained models and training logs to promote reproducible research.

【4】 Can We Identify Unknown Audio Recording Environments in Forensic  Scenarios?
标题:我们可以在法医场景中识别未知的录音环境吗?
链接:https://arxiv.org/abs/2405.02119
作者:Denise Moussa,Germans Hirsch,Christian Riess
备注:This work has been submitted to the IEEE for possible publication
摘要:录音可为刑事调查提供重要证据。一种这样的情况是所记录的音频与记录位置的取证关联。例如,语音信息可能是缩小犯罪候选地点范围的唯一调查线索。到目前为止,一些工作提供了在相对干净的记录条件下进行闭集记录环境分类的工具。然而,在法证调查中,候选地点视具体案件而定。因此,闭集工具在没有对每个情况和相应候选集的足够量的训练样本进行再训练的情况下是不适用的。此外,取证工具必须处理来自不受控制的来源的具有可变属性和质量的音频材料。因此,在这项工作中,我们试图向实际的法医应用场景迈出重要一步。我们提出了一个称为EnvId的表示学习框架,简称环境识别。EnvId避免了针对具体情况的再培训。相反,它是第一个强大的Few-Shot分类看不见的环境位置的工具。我们证明了EnvId可以处理具有法医学挑战性的材料。即使在看不见的信号降级、环境特性或记录位置不匹配的情况下,它也能提供良好的质量预测。我们的代码和数据集将在接受后公开提供。
摘要:Audio recordings may provide important evidence in criminal investigations. One such case is the forensic association of the recorded audio to the recording location. For example, a voice message may be the only investigative cue to narrow down the candidate sites for a crime. Up to now, several works provide tools for closed-set recording environment classification under relatively clean recording conditions. However, in forensic investigations, the candidate locations are case-specific. Thus, closed-set tools are not applicable without retraining on a sufficient amount of training samples for each case and respective candidate set. In addition, a forensic tool has to deal with audio material from uncontrolled sources with variable properties and quality.  In this work, we therefore attempt a major step towards practical forensic application scenarios. We propose a representation learning framework called EnvId, short for environment identification. EnvId avoids case-specific retraining. Instead, it is the first tool for robust few-shot classification of unseen environment locations. We demonstrate that EnvId can handle forensically challenging material. It provides good quality predictions even under unseen signal degradations, environment characteristics or recording position mismatches.  Our code and datasets will be made publicly available upon acceptance.

【5】 Joint sentiment analysis of lyrics and audio in music
标题:音乐中歌词和音频的联合情感分析
链接:https://arxiv.org/abs/2405.01988
作者:Lea Schaab,Anna Kruspe
备注:published at DAGA 2024
摘要:情感或情绪可以在音乐的各个层面上表达出来。在自动分析中,通常分析实际的音频数据,但歌词也可以在情绪感知中发挥至关重要的作用。我们首先分别基于歌词和音频评估各种情感分析模型。相应的方法已经显示出令人满意的结果,但它们也表现出弱点,我们更详细地研究其原因。此外,不同的方法来结合音频和歌词的结果提出和评估。同时考虑这两种模式通常会提高性能。我们调查错误分类和(也是故意的)音频和歌词情感之间的矛盾更密切,并确定可能的原因。最后,我们解决了这个研究领域的基本问题,如高度主观性,缺乏数据,情绪分类不一致。摘要:Sentiment or mood can express themselves on various levels in music. In automatic analysis, the actual audio data is usually analyzed, but the lyrics can also play a crucial role in the perception of moods. We first evaluate various models for sentiment analysis based on lyrics and audio separately. The corresponding approaches already show satisfactory results, but they also exhibit weaknesses, the causes of which we examine in more detail. Furthermore, different approaches to combining the audio and lyrics results are proposed and evaluated. Considering both modalities generally leads to improved performance. We investigate misclassifications and (also intentional) contradictions between audio and lyrics sentiment more closely, and identify possible causes. Finally, we address fundamental problems in this research area, such as high subjectivity, lack of data, and inconsistency in emotion taxonomies.


【6】 Toward end-to-end interpretable convolutional neural networks for  waveform signals
标题:走向用于波信号的端到端可解释卷积神经网络
链接:https://arxiv.org/abs/2405.01815
作者:Linh Vu,Thu Tran,Wern-Han Lim,Raphael Phan
摘要:本文介绍了一种新的卷积神经网络(CNN)框架,该框架专为端到端音频深度学习模型定制,在效率和可解释性方面取得了进步。通过对三个标准语音情感识别数据集进行基准测试实验,我们的框架比Mel谱图特征高出7%。它可以潜在地取代梅尔频率倒谱系数(MFCC),同时保持轻量级。此外,我们使用PhysioNet心音数据库展示了前端层的效率和可解释性,说明了其处理和捕获复杂长波形模式的能力。我们的贡献提供了一个便携式的解决方案,为原始波形数据建立高效和可解释的模型。
摘要:This paper introduces a novel convolutional neural networks (CNN) framework tailored for end-to-end audio deep learning models, presenting advancements in efficiency and explainability. By benchmarking experiments on three standard speech emotion recognition datasets with five-fold cross-validation, our framework outperforms Mel spectrogram features by up to seven percent. It can potentially replace the Mel-Frequency Cepstral Coefficients (MFCC) while remaining lightweight. Furthermore, we demonstrate the efficiency and interpretability of the front-end layer using the PhysioNet Heart Sound Database, illustrating its ability to handle and capture intricate long waveform patterns. Our contributions offer a portable solution for building efficient and interpretable models for raw waveform data.

【7】 Real-time multichannel deep speech enhancement in hearing aids:  Comparing monaural and binaural processing in complex acoustic scenarios
标题:助听器中的实时多通道深度语音增强:比较复杂声学场景中的单耳和双耳处理
链接:https://arxiv.org/abs/2405.01967
作者:Nils L. Westhausen,Hendrik Kayser,Theresa Jansen,Bernd T. Meyer
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:深度学习有可能增强语音信号并提高助听器用户的可理解性。适合现实世界应用的深度模型应该具有低计算复杂度和仅几毫秒的低处理延迟。在本文中,我们将探讨符合这些要求的深度语音增强,并在两个复杂的声学场景中对比单声道和双耳处理算法。这两种算法进行评估与客观的指标,并在实验中与听力受损的听众进行语音噪声测试。结果进行了比较,两个传统的增强策略,即,自适应差分麦克风处理和双耳波束形成。虽然在扩散噪声中,所有算法的表现都相似,但双耳深度学习方法在存在空间干扰的情况下表现最好。通过后分析,这可以归因于低SNR下的改进和精确的空间滤波。
摘要:Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of only a few milliseconds. In this paper, we explore deep speech enhancement that matches these requirements and contrast monaural and binaural processing algorithms in two complex acoustic scenes. Both algorithms are evaluated with objective metrics and in experiments with hearing-impaired listeners performing a speech-in-noise test. Results are compared to two traditional enhancement strategies, i.e., adaptive differential microphone processing and binaural beamforming. While in diffuse noise, all algorithms perform similarly, the binaural deep learning approach performs best in the presence of spatial interferers. Through a post-analysis, this can be attributed to improvements at low SNRs and to precise spatial filtering.

【8】 Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a  Conditional Diffusion Model
标题:转换任何人的声音:使用条件扩散模型进行端到端表达性声音转换
链接:https://arxiv.org/abs/2405.01730
作者:Zongyang Du,Junchen Lu,Kun Zhou,Lakshmish Kaushik,Berrak Sisman
备注:Accepted by Speaker Odyssey 2024
摘要:表达性语音转换(VC)通过对说话人身份和情感风格的联合转换,为情感说话人进行说话人身份转换。表达性VC中任意说话者的情感风格建模尚未得到广泛研究。以前的方法依赖于声码器进行语音重建,这使得语音质量严重依赖于声码器的性能。表达性VC的一个主要挑战在于情感韵律建模。为了解决这些挑战,本文提出了一个完全端到端的表达VC框架的基础上的条件去噪扩散概率模型(DDPM)。我们利用来自自监督语音模型的语音单元作为内容调节,以及从语音情感识别和说话人验证系统中提取的深层特征来建模情感风格和说话人身份。客观和主观的评价表明我们的框架的有效性。代码和样本是公开的。
摘要:Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.

eess.AS音频处理
【1】 TIPAA-SSL: Text Independent Phone-to-Audio Alignment based on  Self-Supervised Learning and Knowledge Transfer
标题:TIPAA-SSL:基于自我监督学习和知识转移的文本独立电话到音频对齐
链接:https://arxiv.org/abs/2405.02124
作者:Noé Tits,Prernna Bhatnagar,Thierry Dutoit
摘要:在本文中,我们提出了一种新的方法,文本无关的音素到音频对齐的基础上音素识别,表示学习和知识转移。我们的方法利用了一个自监督模型(wav2vec2)微调音素识别使用连接时间分类(CTC)损失,降维模型和帧级音素分类器训练感谢强制对齐标签(使用蒙特利尔强制对齐器),以产生多语言的语音表示,因此需要最少的额外训练。我们使用TIMIT数据集和SCRIBE数据集分别针对美国英语和英国英语的合成本地数据来评估我们的模型。我们提出的模型优于国家的最先进的(charsiu)在统计指标,并在语言学习和语音处理系统中的应用。我们离开其他语言的实验,为未来的工作,但系统的设计,使它很容易适应其他语言。
摘要:In this paper, we present a novel approach for text independent phone-to-audio alignment based on phoneme recognition, representation learning and knowledge transfer. Our method leverages a self-supervised model (wav2vec2) fine-tuned for phoneme recognition using a Connectionist Temporal Classification (CTC) loss, a dimension reduction model and a frame-level phoneme classifier trained thanks to forced-alignment labels (using Montreal Forced Aligner) to produce multi-lingual phonetic representations, thus requiring minimal additional training. We evaluate our model using synthetic native data from the TIMIT dataset and the SCRIBE dataset for American and British English, respectively. Our proposed model outperforms the state-of-the-art (charsiu) in statistical metrics and has applications in language learning and speech processing systems. We leave experiments on other languages for future work but the design of the system makes it easily adaptable to other languages.

【2】 Real-time multichannel deep speech enhancement in hearing aids:  Comparing monaural and binaural processing in complex acoustic scenarios
标题:助听器中的实时多通道深度语音增强:比较复杂声学场景中的单耳和双耳处理
链接:https://arxiv.org/abs/2405.01967
作者:Nils L. Westhausen,Hendrik Kayser,Theresa Jansen,Bernd T. Meyer
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:深度学习有可能增强语音信号并提高助听器用户的可理解性。适合现实世界应用的深度模型应该具有低计算复杂度和仅几毫秒的低处理延迟。在本文中,我们将探讨符合这些要求的深度语音增强,并在两个复杂的声学场景中对比单声道和双耳处理算法。这两种算法进行评估与客观的指标,并在实验中与听力受损的听众进行语音噪声测试。结果进行了比较,两个传统的增强策略,即,自适应差分麦克风处理和双耳波束形成。虽然在扩散噪声中,所有算法的表现都相似,但双耳深度学习方法在存在空间干扰的情况下表现最好。通过后分析,这可以归因于低SNR下的改进和精确的空间滤波。
摘要:Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of only a few milliseconds. In this paper, we explore deep speech enhancement that matches these requirements and contrast monaural and binaural processing algorithms in two complex acoustic scenes. Both algorithms are evaluated with objective metrics and in experiments with hearing-impaired listeners performing a speech-in-noise test. Results are compared to two traditional enhancement strategies, i.e., adaptive differential microphone processing and binaural beamforming. While in diffuse noise, all algorithms perform similarly, the binaural deep learning approach performs best in the presence of spatial interferers. Through a post-analysis, this can be attributed to improvements at low SNRs and to precise spatial filtering.

【3】 Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a  Conditional Diffusion Model
标题:转换任何人的声音:使用条件扩散模型进行端到端表达性声音转换
链接:https://arxiv.org/abs/2405.01730
作者:Zongyang Du,Junchen Lu,Kun Zhou,Lakshmish Kaushik,Berrak Sisman备注:Accepted by Speaker Odyssey 2024
摘要:表达性语音转换(VC)通过对说话人身份和情感风格的联合转换,为情感说话人进行说话人身份转换。表达性VC中任意说话者的情感风格建模尚未得到广泛研究。以前的方法依赖于声码器进行语音重建,这使得语音质量严重依赖于声码器的性能。表达性VC的一个主要挑战在于情感韵律建模。为了解决这些挑战,本文提出了一个完全端到端的表达VC框架的基础上的条件去噪扩散概率模型(DDPM)。我们利用来自自监督语音模型的语音单元作为内容调节,以及从语音情感识别和说话人验证系统中提取的深层特征来建模情感风格和说话人身份。客观和主观的评价表明我们的框架的有效性。代码和样本是公开的。
摘要:Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.

【4】 Training-Free Deepfake Voice Recognition by Leveraging Large-Scale  Pre-Trained Models
标题:通过利用大规模预训练模型进行免训练Deepfake语音识别
链接:https://arxiv.org/abs/2405.02179
作者:Alessandro Pianese,Davide Cozzolino,Giovanni Poggi,Luisa Verdoliva
摘要:泛化是当前音频deepfake检测器的主要问题,这些检测器很难对分发数据提供可靠的结果。考虑到越来越多的精确合成方法的发展速度,设计出在未经训练的数据上也能很好地工作的技术是非常重要的。在本文中,我们研究了大规模预训练模型用于音频深度伪造检测的潜力,特别关注泛化能力。为此,在说话人验证框架中重新制定了检测问题,并且通过测试下的语音样本与所声称的身份的语音之间的不匹配来暴露假音频。使用这种范式,在训练中不需要假语音样本,从根本上切断了与生成方法的任何联系,并确保了完全的泛化能力。特征由通用的大型预训练模型提取,无需对特定的虚假检测或说话人验证数据集进行训练或微调。在检测时,只需要被测身份的有限语音片段集。在社区中广泛使用的几个数据集上的实验表明,基于预训练模型的检测器具有出色的性能,并显示出强大的泛化能力,可与分布数据上的监督方法相媲美,并在很大程度上克服了分布数据上的监督方法。
摘要:Generalization is a main issue for current audio deepfake detectors, which struggle to provide reliable results on out-of-distribution data. Given the speed at which more and more accurate synthesis methods are developed, it is very important to design techniques that work well also on data they were not trained for.In this paper we study the potential of large-scale pre-trained models for audio deepfake detection, with special focus on generalization ability. To this end, the detection problem is reformulated in a speaker verification framework and fake audios are exposed by the mismatch between the voice sample under test and the voice of the claimed identity. With this paradigm, no fake speech sample is necessary in training, cutting off any link with the generation method at the root, and ensuring full generalization ability. Features are extracted by general-purpose large pre-trained models, with no need for training or fine-tuning on specific fake detection or speaker verification datasets. At detection time only a limited set of voice fragments of the identity under test is required. Experiments on several datasets widespread in the community show that detectors based on pre-trained models achieve excellent performance and show strong generalization ability, rivaling supervised methods on in-distribution data and largely overcoming them on out-of-distribution data.

【5】 GMP-ATL: Gender-augmented Multi-scale Pseudo-label Enhanced Adaptive  Transfer Learning for Speech Emotion Recognition via HuBERT
标题:GMP-ATL:通过HuBERT进行语音情感识别的性别增强多尺度伪标签增强自适应转移学习
链接:https://arxiv.org/abs/2405.02151
作者:Yu Pan,Yuguang Yang,Heng Lu,Lei Ma,Jianjun Zhao
摘要:预先训练的语音模型的不断发展极大地促进了语音情感识别(SER)。然而,这些方法的性能仍有提高的潜力。GMP-ATL(Gender-augmented Multi-scale Pseudo-label Adaptive Transfer Learning)是一种新的基于HuBERT的自适应迁移学习框架,它首先使用预先训练好的HuBERT,通过多任务学习和多尺度k-means聚类来获取帧级的性别增强多尺度伪标签。然后,为了充分利用所获得的帧级和话语级情感标签,我们结合模型再训练和微调方法来进一步优化GMP-ATL。在IEMOCAP上的实验表明,GMP-ATL的WAR和UAR分别为80.0%和82.0%,超过了现有的单模态SER方法,同时也取得了与多模态SER方法相当的识别效果.
摘要:The continuous evolution of pre-trained speech models has greatly advanced Speech Emotion Recognition (SER). However, there is still potential for enhancement in the performance of these methods. In this paper, we present GMP-ATL (Gender-augmented Multi-scale Pseudo-label Adaptive Transfer Learning), a novel HuBERT-based adaptive transfer learning framework for SER. Specifically, GMP-ATL initially employs the pre-trained HuBERT, implementing multi-task learning and multi-scale k-means clustering to acquire frame-level gender-augmented multi-scale pseudo-labels. Then, to fully leverage both obtained frame-level and utterance-level emotion labels, we incorporate model retraining and fine-tuning methods to further optimize GMP-ATL. Experiments on IEMOCAP show that our GMP-ATL achieves superior recognition performance, with a WAR of 80.0\% and a UAR of 82.0\%, surpassing state-of-the-art unimodal SER methods, while also yielding comparable results with multimodal SER approaches.

【6】 Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
标题:在中国开源数据集中揭示基于LLM的ASB的潜力
链接:https://arxiv.org/abs/2405.02132
作者:Xuelong Geng,Tianyi Xu,Kun Wei,Bingsheng Mu,Hongfei Xue,He Wang,Yangze Li,Pengcheng Guo,Yuhang Dai,Longhao Li,Mingchen Shao,Lei Xie
摘要:大型语言模型在各种NLP任务中表现出无与伦比的有效性,将LLM与自动语音识别集成正在成为主流范例。在此基础上,我们的研究深入研究了一个大型开源中文数据集上的这种范式。具体来说,我们的研究旨在评估各种配置的语音编码器,LLM和投影仪模块的语音基础encoderLLM ASR范式的背景下的影响。此外,我们引入了一个三阶段的训练方法,明确开发,以提高模型的能力,使听觉和文本信息。这种方法的实施,以及ASR组件的战略集成,使我们能够在AISHELL1,TestNet和TestMeeting测试集上实现SOTA性能。我们的分析为基于LLM的ASR系统的未来研究提供了实证基础,并为使用中国数据集优化性能提供了见解。我们将公开发布用于数据准备、训练、推理和评分的所有脚本,以及预训练模型和训练日志,以促进可重复的研究。
摘要:Large Language Models have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition is becoming a mainstream paradigm. Building upon this momentum, our research delves into an indepth examination of this paradigm on a large opensource Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoderLLM ASR paradigm. Furthermore, we introduce a threestage training approach, expressly developed to enhance the model's ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL1, TestNet, and TestMeeting test sets. Our analysis presents an empirical foundation for future research in LLMbased ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pretrained models and training logs to promote reproducible research.

【7】 Can We Identify Unknown Audio Recording Environments in Forensic  Scenarios?
标题:我们可以在法医场景中识别未知的录音环境吗?
链接:https://arxiv.org/abs/2405.02119
作者:Denise Moussa,Germans Hirsch,Christian Riess
备注:This work has been submitted to the IEEE for possible publication
摘要:录音可为刑事调查提供重要证据。一种这样的情况是所记录的音频与记录位置的取证关联。例如,语音信息可能是缩小犯罪候选地点范围的唯一调查线索。到目前为止,一些工作提供了在相对干净的记录条件下进行闭集记录环境分类的工具。然而,在法证调查中,候选地点视具体案件而定。因此,闭集工具在没有对每个情况和相应候选集的足够量的训练样本进行再训练的情况下是不适用的。此外,取证工具必须处理来自不受控制的来源的具有可变属性和质量的音频材料。因此,在这项工作中,我们试图向实际的法医应用场景迈出重要一步。我们提出了一个称为EnvId的表示学习框架,简称环境识别。EnvId避免了针对具体情况的再培训。相反,它是第一个强大的Few-Shot分类看不见的环境位置的工具。我们证明了EnvId可以处理具有法医学挑战性的材料。即使在看不见的信号降级、环境特性或记录位置不匹配的情况下,它也能提供良好的质量预测。我们的代码和数据集将在接受后公开提供。
摘要:Audio recordings may provide important evidence in criminal investigations. One such case is the forensic association of the recorded audio to the recording location. For example, a voice message may be the only investigative cue to narrow down the candidate sites for a crime. Up to now, several works provide tools for closed-set recording environment classification under relatively clean recording conditions. However, in forensic investigations, the candidate locations are case-specific. Thus, closed-set tools are not applicable without retraining on a sufficient amount of training samples for each case and respective candidate set. In addition, a forensic tool has to deal with audio material from uncontrolled sources with variable properties and quality.  In this work, we therefore attempt a major step towards practical forensic application scenarios. We propose a representation learning framework called EnvId, short for environment identification. EnvId avoids case-specific retraining. Instead, it is the first tool for robust few-shot classification of unseen environment locations. We demonstrate that EnvId can handle forensically challenging material. It provides good quality predictions even under unseen signal degradations, environment characteristics or recording position mismatches.  Our code and datasets will be made publicly available upon acceptance.

【8】 Joint sentiment analysis of lyrics and audio in music
标题:音乐中歌词和音频的联合情感分析
链接:https://arxiv.org/abs/2405.01988
作者:Lea Schaab,Anna Kruspe
备注:published at DAGA 2024
摘要:情感或情绪可以在音乐的各个层面上表达出来。在自动分析中,通常分析实际的音频数据,但歌词也可以在情绪感知中发挥至关重要的作用。我们首先分别基于歌词和音频评估各种情感分析模型。相应的方法已经显示出令人满意的结果,但它们也表现出弱点,我们更详细地研究其原因。此外,不同的方法来结合音频和歌词的结果提出和评估。同时考虑这两种模式通常会提高性能。我们调查错误分类和(也是故意的)音频和歌词情感之间的矛盾更密切,并确定可能的原因。最后,我们解决了这个研究领域的基本问题,如高度主观性,缺乏数据,情绪分类不一致。
摘要:Sentiment or mood can express themselves on various levels in music. In automatic analysis, the actual audio data is usually analyzed, but the lyrics can also play a crucial role in the perception of moods. We first evaluate various models for sentiment analysis based on lyrics and audio separately. The corresponding approaches already show satisfactory results, but they also exhibit weaknesses, the causes of which we examine in more detail. Furthermore, different approaches to combining the audio and lyrics results are proposed and evaluated. Considering both modalities generally leads to improved performance. We investigate misclassifications and (also intentional) contradictions between audio and lyrics sentiment more closely, and identify possible causes. Finally, we address fundamental problems in this research area, such as high subjectivity, lack of data, and inconsistency in emotion taxonomies.

【9】 Toward end-to-end interpretable convolutional neural networks for  waveform signals
标题:走向用于波信号的端到端可解释卷积神经网络
链接:https://arxiv.org/abs/2405.01815
作者:Linh Vu,Thu Tran,Wern-Han Lim,Raphael Phan
摘要:本文介绍了一种新的卷积神经网络(CNN)框架,该框架专为端到端音频深度学习模型定制,在效率和可解释性方面取得了进步。通过对三个标准语音情感识别数据集进行基准测试实验,我们的框架比Mel谱图特征高出7%。它可以潜在地取代梅尔频率倒谱系数(MFCC),同时保持轻量级。此外,我们使用PhysioNet心音数据库展示了前端层的效率和可解释性,说明了其处理和捕获复杂长波形模式的能力。我们的贡献提供了一个便携式的解决方案,为原始波形数据建立高效和可解释的模型。
摘要:This paper introduces a novel convolutional neural networks (CNN) framework tailored for end-to-end audio deep learning models, presenting advancements in efficiency and explainability. By benchmarking experiments on three standard speech emotion recognition datasets with five-fold cross-validation, our framework outperforms Mel spectrogram features by up to seven percent. It can potentially replace the Mel-Frequency Cepstral Coefficients (MFCC) while remaining lightweight. Furthermore, we demonstrate the efficiency and interpretability of the front-end layer using the PhysioNet Heart Sound Database, illustrating its ability to handle and capture intricate long waveform patterns. Our contributions offer a portable solution for building efficient and interpretable models for raw waveform data.


机器翻译由腾讯交互翻译提供,仅供参考