今日论文合集:cs.SD语音2篇,eess.AS音频处理2篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】The RoyalFlush Automatic Speech Diarization and Recognition System for  In-Car Multi-Channel Automatic Speech Recognition Challenge

标题:用于车载多通道自动语音识别挑战赛的RoyalFlush自动语音扩展和识别系统

链接:https://arxiv.org/abs/2405.05498

作者:Jingguang Tian,Shuaishuai Ye,Shunfei Chen,Yang Xiang,Zhaohui Yin,Xinhui Hu,Xinkang Xu

摘要:本文介绍了我们的系统提交的车内多通道自动语音识别(ICMC-ASR)的挑战,重点是在复杂的多说话人场景中的说话人日记和语音识别。为了解决这些挑战,我们开发了端到端的说话人日记化模型,与开发集的官方基线相比,该模型将日记化错误率(DER)显著降低了49.58%。对于语音识别,我们利用自监督学习表示来训练端到端ASR模型。通过集成这些模型,我们在轨道1评估集上获得了16.93%的字符错误率(CER),在轨道2评估集上获得了25.88%的级联最小排列字符错误率(cpCER)。

摘要:This paper presents our system submission for the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) Challenge, which focuses on speaker diarization and speech recognition in complex multi-speaker scenarios. To address these challenges, we develop end-to-end speaker diarization models that notably decrease the diarization error rate (DER) by 49.58\% compared to the official baseline on the development set. For speech recognition, we utilize self-supervised learning representations to train end-to-end ASR models. By integrating these models, we achieve a character error rate (CER) of 16.93\% on the track 1 evaluation set, and a concatenated minimum permutation character error rate (cpCER) of 25.88\% on the track 2 evaluation set.


【2】 AFEN: Respiratory Disease Classification using Ensemble Learning
标题:AFEN:使用整体学习进行呼吸道疾病分类
链接:https://arxiv.org/abs/2405.05467
作者:Rahul Nadkarni,Emmanouil Nikolakakis,Razvan Marinescu
备注:Under Review Process for MLForHC 2024
摘要:我们提出了AFEN(音频特征包围学习),这是一种利用卷积神经网络(CNN)和XGBoost以集成学习方式对一系列呼吸系统疾病进行最先进的音频分类的模型。我们使用精心选择的音频特征组合,这些特征提供了数据的突出属性,并允许准确分类。然后将提取的特征用作两个单独的模型分类器的输入:1)多特征CNN分类器和2)XGBoost分类器。两个模型的输出,然后融合使用软表决。因此,通过利用集成学习,我们实现了更高的鲁棒性和准确性。我们评估了920个呼吸音的数据库上的模型的性能,该数据库经历了数据增强技术,以增加数据的多样性和模型的泛化能力。我们通过经验验证,AFEN使用精度和召回率作为指标,同时将训练时间减少了60%,从而达到了最新的水平。
摘要:We present AFEN (Audio Feature Ensemble Learning), a model that leverages Convolutional Neural Networks (CNN) and XGBoost in an ensemble learning fashion to perform state-of-the-art audio classification for a range of respiratory diseases. We use a meticulously selected mix of audio features which provide the salient attributes of the data and allow for accurate classification. The extracted features are then used as an input to two separate model classifiers 1) a multi-feature CNN classifier and 2) an XGBoost Classifier. The outputs of the two models are then fused with the use of soft voting. Thus, by exploiting ensemble learning, we achieve increased robustness and accuracy. We evaluate the performance of the model on a database of 920 respiratory sounds, which undergoes data augmentation techniques to increase the diversity of the data and generalizability of the model. We empirically verify that AFEN sets a new state-of-the-art using Precision and Recall as metrics, while decreasing training time by 60%.

eess.AS音频处理
【1】 The RoyalFlush Automatic Speech Diarization and Recognition System for  In-Car Multi-Channel Automatic Speech Recognition Challenge
标题:用于车载多通道自动语音识别挑战赛的RoyalFlush自动语音扩展和识别系统
链接:https://arxiv.org/abs/2405.05498
作者:Jingguang Tian,Shuaishuai Ye,Shunfei Chen,Yang Xiang,Zhaohui Yin,Xinhui Hu,Xinkang Xu
摘要:本文介绍了我们的系统提交的车内多通道自动语音识别(ICMC-ASR)的挑战,重点是在复杂的多说话人场景中的说话人日记和语音识别。为了解决这些挑战,我们开发了端到端的说话人日记化模型,与开发集的官方基线相比,该模型将日记化错误率(DER)显著降低了49.58%。对于语音识别,我们利用自监督学习表示来训练端到端ASR模型。通过集成这些模型,我们在轨道1评估集上获得了16.93%的字符错误率(CER),在轨道2评估集上获得了25.88%的级联最小排列字符错误率(cpCER)。
摘要:This paper presents our system submission for the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) Challenge, which focuses on speaker diarization and speech recognition in complex multi-speaker scenarios. To address these challenges, we develop end-to-end speaker diarization models that notably decrease the diarization error rate (DER) by 49.58\% compared to the official baseline on the development set. For speech recognition, we utilize self-supervised learning representations to train end-to-end ASR models. By integrating these models, we achieve a character error rate (CER) of 16.93\% on the track 1 evaluation set, and a concatenated minimum permutation character error rate (cpCER) of 25.88\% on the track 2 evaluation set.

【2】 AFEN: Respiratory Disease Classification using Ensemble Learning
标题:AFEN:使用整体学习进行呼吸道疾病分类
链接:https://arxiv.org/abs/2405.05467
作者:Rahul Nadkarni,Emmanouil Nikolakakis,Razvan Marinescu
备注:Under Review Process for MLForHC 2024
摘要:我们提出了AFEN(音频特征包围学习),这是一种利用卷积神经网络(CNN)和XGBoost以集成学习方式对一系列呼吸系统疾病进行最先进的音频分类的模型。我们使用精心选择的音频特征组合,这些特征提供了数据的突出属性,并允许准确分类。然后将提取的特征用作两个单独的模型分类器的输入:1)多特征CNN分类器和2)XGBoost分类器。两个模型的输出,然后融合使用软表决。因此,通过利用集成学习,我们实现了更高的鲁棒性和准确性。我们评估了920个呼吸音的数据库上的模型的性能,该数据库经历了数据增强技术,以增加数据的多样性和模型的泛化能力。我们通过经验验证,AFEN使用精度和召回率作为指标,同时将训练时间减少了60%,从而达到了最新的水平。
摘要:We present AFEN (Audio Feature Ensemble Learning), a model that leverages Convolutional Neural Networks (CNN) and XGBoost in an ensemble learning fashion to perform state-of-the-art audio classification for a range of respiratory diseases. We use a meticulously selected mix of audio features which provide the salient attributes of the data and allow for accurate classification. The extracted features are then used as an input to two separate model classifiers 1) a multi-feature CNN classifier and 2) an XGBoost Classifier. The outputs of the two models are then fused with the use of soft voting. Thus, by exploiting ensemble learning, we achieve increased robustness and accuracy. We evaluate the performance of the model on a database of 920 respiratory sounds, which undergoes data augmentation techniques to increase the diversity of the data and generalizability of the model. We empirically verify that AFEN sets a new state-of-the-art using Precision and Recall as metrics, while decreasing training time by 60%.

机器翻译由腾讯交互翻译提供,仅供参考