今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理7篇。

cs.SD语音

【1】 A Model You Can Hear: Audio Identification with Playable Prototypes

标题:您可以听到的模型:基于可播放原型的音频识别

链接:https://arxiv.org/abs/2208.03311

作者:Romain Loiseau,Baptiste Bouvier,Yann Teytaut,Elliot Vincent,Mathieu Aubry,Loic Landrieu
机构:LIGM, Ecole des Ponts, Univ Gustave Eiffel, CNRS, France,  LASTIG, Univ Gustave Eiffel, IGN, ENSG,  STMS Lab, UMR , (IRCAM, CNRS, Sorbonne University), Paris, France,  INRIA and DIENS (ENS-PSL, CNRS, INRIA)
摘要:机器学习技术已被证明对音频内容的分类和分析非常有用.然而,最近的方法通常依赖于难以解释的抽象和高维表示.受针对图像和3D数据开发的变换不变方法的启发,我们提出了一种基于可学习频谱原型的音频识别模型.配备专用变换网络,这些原型可用于对来自大量声音集合的输入音频样本进行聚类和分类。我们的模型可以在有监督或无监督的情况下进行训练,在保持易于解释的同时,达到最先进的说话者和乐器识别结果。代码可在以下网址获得:https://github.com/romainloiseau/a-model-you-can-hear摘要:Machine learning techniques have proved useful for classifying and analyzing audio content. However, recent methods typically rely on abstract and high-dimensional representations that are difficult to interpret. Inspired by transformation-invariant approaches developed for image and 3D data, we propose an audio identification model based on learnable spectral prototypes. Equipped with dedicated transformation networks, these prototypes can be used to cluster and classify input audio samples from large collections of sounds. Our model can be trained with or without supervision and reaches state-of-the-art results for speaker and instrument identification, while remaining easily interpretable. The code is available at: https://github.com/romainloiseau/a-model-you-can-hear


【2】 Robust Acoustic Domain Identification with its Application to Speaker  Diarization

标题:稳健的声域辨识及其在说话人识别中的应用

链接:https://arxiv.org/abs/2208.03162

作者:A Kishore Kumar,Shefali Waldekar,Md Sahidullah,Goutam Saha
机构:the date of receipt and acceptance should be inserted later
备注:International Journal of Speech Technology (2022)
摘要:随着近年来多媒体内容的增加,音频录制环境也出现了更多变化。如果音频处理系统的前端具有识别声学域的模块,则可能会对音频处理系统有所帮助。本文将演示\emph{声学域识别}的概念。(ADI),用于\emph{扬声器日志记录}。为此,我们首先对第三次DIHARD挑战的各个领域进行了详细的研究,突出了使它们相互区别的因素。我们的主要贡献是为ADI开发一种简单而有效的解决方案,在本工作中,我们将探讨扬声器嵌入的方法,我们将ADI模块与DIHARD III挑战的扬声器diarization框架集成在一起。当根据下式优化凝聚层次聚类的阈值时,性能与基线相比有了显著提高在DIHARD III评估集的Track 1上,我们在核心和完全条件下分别实现了超过$5 $和$8 $的DER相对改善。
摘要:With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this paper, we demonstrate the idea of \emph{acoustic domain identification} (ADI) for \emph{speaker diarization}. For this, we first present a detailed study of the various domains of the third DIHARD challenge highlighting the factors that differentiated them from each other. Our main contribution is to develop a simple and efficient solution for ADI. In the present work, we explore speaker embeddings for this task. Next, we integrate the ADI module with the speaker diarization framework of the DIHARD III challenge. The performance substantially improved over that of the baseline when the thresholds for agglomerative hierarchical clustering were optimized according to the respective domains. We achieved a relative improvement of more than $5\%$ and $8\%$ in DER for core and full conditions, respectively, on Track 1 of the DIHARD III evaluation set.


【3】 Deep Feature Learning for Medical Acoustics

标题:面向医学声学的深度特征学习

链接:https://arxiv.org/abs/2208.03084

作者:Alessandro Maria Poirè,Federico Simonetta,Stavros Ntalampiras
机构:Ntalampiras[,−,−,−,], LIM – Music Informatics Laboratory, Department of Computer Science, University of Milano
备注:Published at ICANN 2022
摘要:本文的目的是比较医学声学任务中不同的可学习前端。已经实现了一个框架来将人的呼吸声和心跳分为两类,即健康的和受病理影响的。在获得两个合适的数据集之后,我们继续使用两个最先进的可学习前端—-LEAF和nnAudio-—加上一个不可学习的基线前端来分类声音。然后将计算出的特征送入两个不同的CNN模型,即VGG16和EfficientNet。前端在参数数量、计算资源和有效性方面进行了仔细的基准测试。本研究展示了在神经音频分类系统中集成可学习前端如何提高性能,尤其是在医学声学领域。然而,使用此类框架会使所需的数据量更大。因此,如果可用于训练的数据量足够大,以辅助特征学习过程,则此类框架非常有用。
摘要:The purpose of this paper is to compare different learnable frontends in medical acoustics tasks. A framework has been implemented to classify human respiratory sounds and heartbeats in two categories, i.e. healthy or affected by pathologies. After obtaining two suitable datasets, we proceeded to classify the sounds using two learnable state-of-art frontends -- LEAF and nnAudio -- plus a non-learnable baseline frontend, i.e. Mel-filterbanks. The computed features are then fed into two different CNN models, namely VGG16 and EfficientNet. The frontends are carefully benchmarked in terms of the number of parameters, computational resources, and effectiveness.  This work demonstrates how the integration of learnable frontends in neural audio classification systems may improve performance, especially in the field of medical acoustics. However, the usage of such frameworks makes the needed amount of data even larger. Consequently, they are useful if the amount of data available for training is adequately large to assist the feature learning process.


【4】 Large vocabulary speech recognition for languages of Africa:  multilingual modeling and self-supervised learning

标题:非洲语言的大词汇量语音识别:多语言建模与自监督学习

链接:https://arxiv.org/abs/2208.03067

作者:Sandy Ritchie,You-Chi Cheng,Mingqing Chen,Rajiv Mathews,Daan van Esch,Bo Li,Khe Chai Sim
机构:Google Research
摘要:在非洲使用的2,000多种语言中,几乎没有一种语言拥有广泛可用的自动语音识别系统,而且所需的数据也只有少数几种语言可用。我们已经试验了两种技术,这两种技术可能为非洲语言的大词汇量语音识别提供途径:多语言建模和自监督学习。我们收集了可用的开源数据和15种语言的数据,并使用这些技术训练了实验模型。我们的结果表明,将多语言端到端模型中可用的少量数据合并,并对无监督数据进行预训练,有助于提高许多非洲语言的语音识别质量。
摘要:Almost none of the 2,000+ languages spoken in Africa have widely available automatic speech recognition systems, and the required data is also only available for a few languages. We have experimented with two techniques which may provide pathways to large vocabulary speech recognition for African languages: multilingual modeling and self-supervised learning. We gathered available open source data and collected data for 15 languages, and trained experimental models using these techniques. Our results show that pooling the small amounts of data available in multilingual end-to-end models, and pre-training on unsupervised data can help improve speech recognition quality for many African languages.


【5】 Hybrid Multimodal Feature Extraction, Mining and Fusion for Sentiment  Analysis

标题:面向情感分析的混合多模态特征提取、挖掘与融合

链接:https://arxiv.org/abs/2208.03051

作者:Jia Li,Ziyang Zhang,Junjie Lang,Yueqi Jiang,Liuwei An,Peng Zou,Yangyang Xu,Sheng Gao,Jie Lin,Chunxiao Fan,Xiao Sun,Meng Wang
机构:School of Computer Science and, Information Engineering, Hefei University of Technology, Hefei, China, ZhongJuYuan Intelligent Technology, Co., Ltd, Institute of Artificial Intelligence, Hefei Comprehensive National, Science Center
备注:8 pages, 2 figures, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
摘要:在本文中,我们提出了我们的解决方案,为多模态情感分析的挑战(MuSe)2022,包括MuSe-Humor、MuSe-Reaction和MuSe-Stress三个子挑战。MuSe 2022利用不同的模态和数据集,重点研究幽默检测、情绪反应和多模态情绪应激。在我们的工作中,提取了不同类型的多模态特征,包括听觉、视觉、提出了一种基于TEMMA和GRU的语音特征融合方法,2)通过对多模态情感特征的挖掘和融合,大大提高了多模态情感预测的准确性和可靠性; 3)提出了一种基于多模态情感特征的情感预测方法。在模型训练中采用了有效的数据扩充策略,有效地缓解了样本不均衡问题,避免了模型对主题特征的学习偏差,对于MuSe-Humor子挑战,模型的AUC值为0.8932,对于MuSe-Reaction子挑战,模型在测试集上的Pearson相关系数为0.3879,对于MuSe-Stress子挑战,我们的方法在测试数据集上的唤醒和效价方面都优于基线,达到0.5151的最终组合结果。
摘要:In this paper, we present our solutions for the Multimodal Sentiment Analysis Challenge (MuSe) 2022, which includes MuSe-Humor, MuSe-Reaction and MuSe-Stress Sub-challenges. The MuSe 2022 focuses on humor detection, emotional reactions and multimodal emotional stress utilising different modalities and data sets. In our work, different kinds of multimodal features are extracted, including acoustic, visual, text and biological features. These features are fused by TEMMA and GRU with self-attention mechanism frameworks. In this paper, 1) several new audio features, facial expression features and paragraph-level text embeddings are extracted for accuracy improvement. 2) we substantially improve the accuracy and reliability for multimodal sentiment prediction by mining and blending the multimodal features. 3) effective data augmentation strategies are applied in model training to alleviate the problem of sample imbalance and prevent the model form learning biased subject characters. For the MuSe-Humor sub-challenge, our model obtains the AUC score of 0.8932. For the MuSe-Reaction sub-challenge, the Pearson's Correlations Coefficient of our approach on the test set is 0.3879, which outperforms all other participants. For the MuSe-Stress sub-challenge, our approach outperforms the baseline in both arousal and valence on the test dataset, reaching a final combined result of 0.5151.


【6】 Time-Frequency Distributions of Heart Sound Signals: A Comparative Study  using Convolutional Neural Networks

标题:心音信号的时频分布:卷积神经网络的比较研究

链接:https://arxiv.org/abs/2208.03128

作者:Xinqi Bao,Yujia Xu,Hak-Keung Lam,Mohamed Trabelsi,Ines Chihi,Lilia Sidhom,Ernest N. Kamavuako
机构:Department of Engineering, King’s College London, Strand, London, WC,R ,LS, United Kingdom, Department of Electronic and Communications Engineering, Kuwait College of Science and Technology, Kuwait
摘要:时频分布(TFD)支持早期心脏筛查中的心音表征和分类。然而,尽管TFD在信号分析中经常使用,但没有研究全面比较它们在深度学习自动诊断方面的性能。此外,作为卷积神经网络输入的信号处理方法的组合(CNNs)已被证明是一种提高信号分类性能的实用方法,因此,本研究旨在研究TFD/组合TFD作为CNNs输入的最佳使用,结果表明:1)将心音信号变换到TF域比使用原始信号具有更高的分类性能,在所有CNN模型的TF域中,性能差异很小平均准确率在1.3%以内,连续小波变换连续小波变换与线性调频小波变换2)适当增加CNN的容量和结构优化可以提高CNN的性能,而网络结构不应过于复杂.根据ResNet或SEResNet族的结果,3)将TFDs作为CNN的输入并没有显著地改善分类结果。研究结果为选择TFDs作为CNN输入和设计用于心音分类的CNN结构提供了知识。
摘要:Time-Frequency Distributions (TFDs) support the heart sound characterisation and classification in early cardiac screening. However, despite the frequent use of TFDs in signal analysis, no study comprehensively compared their performances on deep learning for automatic diagnosis. Furthermore, the combination of signal processing methods as inputs for Convolutional Neural Networks (CNNs) has been proved as a practical approach to increasing signal classification performance. Therefore, this study aimed to investigate the optimal use of TFD/ combined TFDs as input for CNNs. The presented results revealed that: 1) The transformation of the heart sound signal into the TF domain achieves higher classification performance than using of raw signals. Among the TFDs, the difference in the performance was slight for all the CNN models (within $1.3\%$ in average accuracy). However, Continuous wavelet transform (CWT) and Chirplet transform (CT) outperformed the rest. 2) The appropriate increase of the CNN capacity and architecture optimisation can improve the performance, while the network architecture should not be overly complicated. Based on the ResNet or SEResNet family results, the increase in the number of parameters and the depth of the structure do not improve the performance apparently. 3) Combining TFDs as CNN inputs did not significantly improve the classification results. The findings of this study provided the knowledge for selecting TFDs as CNN input and designing CNN architecture for heart sound classification.


【7】 AID: Open-source Anechoic Interferer Dataset

标题:AID:开源消声干扰器数据集

链接:https://arxiv.org/abs/2208.03023

作者:Philipp Götz,Cagdas Tuna,Andreas Walther,Emanuël A. P. Habets
机构:International Audio Laboratories Erlangen†, Germany., Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany.
备注:Accepted for publication at IWAENC 2022
摘要:介绍了一个家庭环境中各种声源的消声记录数据集,该数据集旨在成为非平稳环境噪声信号的资源,当与声学脉冲响应进行卷积时,可用于模拟复杂的声学场景。此外,还提供了一个Python库,用于生成数据集中记录的随机混合。其可用作非平稳干扰信号。
摘要:A dataset of anechoic recordings of various sound sources encountered in domestic environments is presented. The dataset is intended to be a resource of non-stationary, environmental noise signals that, when convolved with acoustic impulse responses, can be used to simulate complex acoustic scenes. Additionally, a Python library is provided to generate random mixtures of the recordings in the dataset, which can be used as non-stationary interference signals.


eess.AS音频处理

【1】 Time-Frequency Distributions of Heart Sound Signals: A Comparative Study  using Convolutional Neural Networks

标题:心音信号的时频分布:卷积神经网络的比较研究

链接:https://arxiv.org/abs/2208.03128

* 与cs.SD语音【6】为同一篇

作者:Xinqi Bao,Yujia Xu,Hak-Keung Lam,Mohamed Trabelsi,Ines Chihi,Lilia Sidhom,Ernest N. Kamavuako
机构:Department of Engineering, King’s College London, Strand, London, WC,R ,LS, United Kingdom, Department of Electronic and Communications Engineering, Kuwait College of Science and Technology, Kuwait
摘要:时频分布(TFD)支持早期心脏筛查中的心音表征和分类。然而,尽管TFD在信号分析中经常使用,但没有研究全面比较它们在深度学习自动诊断方面的性能。此外,作为卷积神经网络输入的信号处理方法的组合(CNNs)已被证明是一种提高信号分类性能的实用方法,因此,本研究旨在研究TFD/组合TFD作为CNNs输入的最佳使用,结果表明:1)将心音信号变换到TF域比使用原始信号具有更高的分类性能,在所有CNN模型的TF域中,性能差异很小平均准确率在1.3%以内,连续小波变换连续小波变换与线性调频小波变换2)适当增加CNN的容量和结构优化可以提高CNN的性能,而网络结构不应过于复杂.根据ResNet或SEResNet族的结果,3)将TFDs作为CNN的输入并没有显著地改善分类结果。研究结果为选择TFDs作为CNN输入和设计用于心音分类的CNN结构提供了知识。
摘要:Time-Frequency Distributions (TFDs) support the heart sound characterisation and classification in early cardiac screening. However, despite the frequent use of TFDs in signal analysis, no study comprehensively compared their performances on deep learning for automatic diagnosis. Furthermore, the combination of signal processing methods as inputs for Convolutional Neural Networks (CNNs) has been proved as a practical approach to increasing signal classification performance. Therefore, this study aimed to investigate the optimal use of TFD/ combined TFDs as input for CNNs. The presented results revealed that: 1) The transformation of the heart sound signal into the TF domain achieves higher classification performance than using of raw signals. Among the TFDs, the difference in the performance was slight for all the CNN models (within $1.3\%$ in average accuracy). However, Continuous wavelet transform (CWT) and Chirplet transform (CT) outperformed the rest. 2) The appropriate increase of the CNN capacity and architecture optimisation can improve the performance, while the network architecture should not be overly complicated. Based on the ResNet or SEResNet family results, the increase in the number of parameters and the depth of the structure do not improve the performance apparently. 3) Combining TFDs as CNN inputs did not significantly improve the classification results. The findings of this study provided the knowledge for selecting TFDs as CNN input and designing CNN architecture for heart sound classification.


【2】 AID: Open-source Anechoic Interferer Dataset

标题:AID:开源消声干扰器数据集

链接:https://arxiv.org/abs/2208.03023

* 与cs.SD语音【7】为同一篇

作者:Philipp Götz,Cagdas Tuna,Andreas Walther,Emanuël A. P. Habets
机构:International Audio Laboratories Erlangen†, Germany., Fraunhofer Institute for Integrated Circuits IIS, Erlangen, Germany.
备注:Accepted for publication at IWAENC 2022
摘要:介绍了一个家庭环境中各种声源的消声记录数据集,该数据集旨在成为非平稳环境噪声信号的资源,当与声学脉冲响应进行卷积时,可用于模拟复杂的声学场景。此外,还提供了一个Python库,用于生成数据集中记录的随机混合。其可用作非平稳干扰信号。
摘要:A dataset of anechoic recordings of various sound sources encountered in domestic environments is presented. The dataset is intended to be a resource of non-stationary, environmental noise signals that, when convolved with acoustic impulse responses, can be used to simulate complex acoustic scenes. Additionally, a Python library is provided to generate random mixtures of the recordings in the dataset, which can be used as non-stationary interference signals.


【3】 A Model You Can Hear: Audio Identification with Playable Prototypes

标题:您可以听到的模型:基于可播放原型的音频识别

链接:https://arxiv.org/abs/2208.03311

* 与cs.SD语音【1】为同一篇

作者:Romain Loiseau,Baptiste Bouvier,Yann Teytaut,Elliot Vincent,Mathieu Aubry,Loic Landrieu
机构:LIGM, Ecole des Ponts, Univ Gustave Eiffel, CNRS, France,  LASTIG, Univ Gustave Eiffel, IGN, ENSG,  STMS Lab, UMR , (IRCAM, CNRS, Sorbonne University), Paris, France,  INRIA and DIENS (ENS-PSL, CNRS, INRIA)
摘要:机器学习技术已被证明对音频内容的分类和分析非常有用.然而,最近的方法通常依赖于难以解释的抽象和高维表示.受针对图像和3D数据开发的变换不变方法的启发,我们提出了一种基于可学习频谱原型的音频识别模型.配备专用变换网络,这些原型可用于对来自大量声音集合的输入音频样本进行聚类和分类。我们的模型可以在有监督或无监督的情况下进行训练,在保持易于解释的同时,达到最先进的说话者和乐器识别结果。代码可在以下网址获得:https://github.com/romainloiseau/a-model-you-can-hear
摘要:Machine learning techniques have proved useful for classifying and analyzing audio content. However, recent methods typically rely on abstract and high-dimensional representations that are difficult to interpret. Inspired by transformation-invariant approaches developed for image and 3D data, we propose an audio identification model based on learnable spectral prototypes. Equipped with dedicated transformation networks, these prototypes can be used to cluster and classify input audio samples from large collections of sounds. Our model can be trained with or without supervision and reaches state-of-the-art results for speaker and instrument identification, while remaining easily interpretable. The code is available at: https://github.com/romainloiseau/a-model-you-can-hear


【4】 Robust Acoustic Domain Identification with its Application to Speaker  Diarization

标题:稳健的声域辨识及其在说话人识别中的应用

链接:https://arxiv.org/abs/2208.03162

* 与cs.SD语音【2】为同一篇

作者:A Kishore Kumar,Shefali Waldekar,Md Sahidullah,Goutam Saha
机构:the date of receipt and acceptance should be inserted later
备注:International Journal of Speech Technology (2022)
摘要:随着近年来多媒体内容的增加,音频录制环境也出现了更多变化。如果音频处理系统的前端具有识别声学域的模块,则可能会对音频处理系统有所帮助。本文将演示\emph{声学域识别}的概念。(ADI),用于\emph{扬声器日志记录}。为此,我们首先对第三次DIHARD挑战的各个领域进行了详细的研究,突出了使它们相互区别的因素。我们的主要贡献是为ADI开发一种简单而有效的解决方案,在本工作中,我们将探讨扬声器嵌入的方法,我们将ADI模块与DIHARD III挑战的扬声器diarization框架集成在一起。当根据下式优化凝聚层次聚类的阈值时,性能与基线相比有了显著提高在DIHARD III评估集的Track 1上,我们在核心和完全条件下分别实现了超过$5 $和$8 $的DER相对改善。
摘要:With the rise in multimedia content over the years, more variety is observed in the recording environments of audio. An audio processing system might benefit when it has a module to identify the acoustic domain at its front-end. In this paper, we demonstrate the idea of \emph{acoustic domain identification} (ADI) for \emph{speaker diarization}. For this, we first present a detailed study of the various domains of the third DIHARD challenge highlighting the factors that differentiated them from each other. Our main contribution is to develop a simple and efficient solution for ADI. In the present work, we explore speaker embeddings for this task. Next, we integrate the ADI module with the speaker diarization framework of the DIHARD III challenge. The performance substantially improved over that of the baseline when the thresholds for agglomerative hierarchical clustering were optimized according to the respective domains. We achieved a relative improvement of more than $5\%$ and $8\%$ in DER for core and full conditions, respectively, on Track 1 of the DIHARD III evaluation set.


【5】 Deep Feature Learning for Medical Acoustics

标题:面向医学声学的深度特征学习

链接:https://arxiv.org/abs/2208.03084

* 与cs.SD语音【3】为同一篇

作者:Alessandro Maria Poirè,Federico Simonetta,Stavros Ntalampiras
机构:Ntalampiras[,−,−,−,], LIM – Music Informatics Laboratory, Department of Computer Science, University of Milano
备注:Published at ICANN 2022
摘要:本文的目的是比较医学声学任务中不同的可学习前端。已经实现了一个框架来将人的呼吸声和心跳分为两类,即健康的和受病理影响的。在获得两个合适的数据集之后,我们继续使用两个最先进的可学习前端—-LEAF和nnAudio-—加上一个不可学习的基线前端来分类声音。然后将计算出的特征送入两个不同的CNN模型,即VGG16和EfficientNet。前端在参数数量、计算资源和有效性方面进行了仔细的基准测试。本研究展示了在神经音频分类系统中集成可学习前端如何提高性能,尤其是在医学声学领域。然而,使用此类框架会使所需的数据量更大。因此,如果可用于训练的数据量足够大,以辅助特征学习过程,则此类框架非常有用。
摘要:The purpose of this paper is to compare different learnable frontends in medical acoustics tasks. A framework has been implemented to classify human respiratory sounds and heartbeats in two categories, i.e. healthy or affected by pathologies. After obtaining two suitable datasets, we proceeded to classify the sounds using two learnable state-of-art frontends -- LEAF and nnAudio -- plus a non-learnable baseline frontend, i.e. Mel-filterbanks. The computed features are then fed into two different CNN models, namely VGG16 and EfficientNet. The frontends are carefully benchmarked in terms of the number of parameters, computational resources, and effectiveness.  This work demonstrates how the integration of learnable frontends in neural audio classification systems may improve performance, especially in the field of medical acoustics. However, the usage of such frameworks makes the needed amount of data even larger. Consequently, they are useful if the amount of data available for training is adequately large to assist the feature learning process.


【6】 Large vocabulary speech recognition for languages of Africa:  multilingual modeling and self-supervised learning

标题:非洲语言的大词汇量语音识别:多语言建模与自监督学习

链接:https://arxiv.org/abs/2208.03067

* 与cs.SD语音【4】为同一篇

作者:Sandy Ritchie,You-Chi Cheng,Mingqing Chen,Rajiv Mathews,Daan van Esch,Bo Li,Khe Chai Sim
机构:Google Research
摘要:在非洲使用的2,000多种语言中,几乎没有一种语言拥有广泛可用的自动语音识别系统,而且所需的数据也只有少数几种语言可用。我们已经试验了两种技术,这两种技术可能为非洲语言的大词汇量语音识别提供途径:多语言建模和自监督学习。我们收集了可用的开源数据和15种语言的数据,并使用这些技术训练了实验模型。我们的结果表明,将多语言端到端模型中可用的少量数据合并,并对无监督数据进行预训练,有助于提高许多非洲语言的语音识别质量。
摘要:Almost none of the 2,000+ languages spoken in Africa have widely available automatic speech recognition systems, and the required data is also only available for a few languages. We have experimented with two techniques which may provide pathways to large vocabulary speech recognition for African languages: multilingual modeling and self-supervised learning. We gathered available open source data and collected data for 15 languages, and trained experimental models using these techniques. Our results show that pooling the small amounts of data available in multilingual end-to-end models, and pre-training on unsupervised data can help improve speech recognition quality for many African languages.


【7】 Hybrid Multimodal Feature Extraction, Mining and Fusion for Sentiment  Analysis

标题:面向情感分析的混合多模态特征提取、挖掘与融合

链接:https://arxiv.org/abs/2208.03051

* 与cs.SD语音【5】为同一篇

作者:Jia Li,Ziyang Zhang,Junjie Lang,Yueqi Jiang,Liuwei An,Peng Zou,Yangyang Xu,Sheng Gao,Jie Lin,Chunxiao Fan,Xiao Sun,Meng Wang
机构:School of Computer Science and, Information Engineering, Hefei University of Technology, Hefei, China, ZhongJuYuan Intelligent Technology, Co., Ltd, Institute of Artificial Intelligence, Hefei Comprehensive National, Science Center
备注:8 pages, 2 figures, to appear in MuSe 2022 (ACM MM2022 co-located workshop)
摘要:在本文中,我们提出了我们的解决方案,为多模态情感分析的挑战(MuSe)2022,包括MuSe-Humor、MuSe-Reaction和MuSe-Stress三个子挑战。MuSe 2022利用不同的模态和数据集,重点研究幽默检测、情绪反应和多模态情绪应激。在我们的工作中,提取了不同类型的多模态特征,包括听觉、视觉、提出了一种基于TEMMA和GRU的语音特征融合方法,2)通过对多模态情感特征的挖掘和融合,大大提高了多模态情感预测的准确性和可靠性; 3)提出了一种基于多模态情感特征的情感预测方法。在模型训练中采用了有效的数据扩充策略,有效地缓解了样本不均衡问题,避免了模型对主题特征的学习偏差,对于MuSe-Humor子挑战,模型的AUC值为0.8932,对于MuSe-Reaction子挑战,模型在测试集上的Pearson相关系数为0.3879,对于MuSe-Stress子挑战,我们的方法在测试数据集上的唤醒和效价方面都优于基线,达到0.5151的最终组合结果。
摘要:In this paper, we present our solutions for the Multimodal Sentiment Analysis Challenge (MuSe) 2022, which includes MuSe-Humor, MuSe-Reaction and MuSe-Stress Sub-challenges. The MuSe 2022 focuses on humor detection, emotional reactions and multimodal emotional stress utilising different modalities and data sets. In our work, different kinds of multimodal features are extracted, including acoustic, visual, text and biological features. These features are fused by TEMMA and GRU with self-attention mechanism frameworks. In this paper, 1) several new audio features, facial expression features and paragraph-level text embeddings are extracted for accuracy improvement. 2) we substantially improve the accuracy and reliability for multimodal sentiment prediction by mining and blending the multimodal features. 3) effective data augmentation strategies are applied in model training to alleviate the problem of sample imbalance and prevent the model form learning biased subject characters. For the MuSe-Humor sub-challenge, our model obtains the AUC score of 0.8932. For the MuSe-Reaction sub-challenge, the Pearson's Correlations Coefficient of our approach on the test set is 0.3879, which outperforms all other participants. For the MuSe-Stress sub-challenge, our approach outperforms the baseline in both arousal and valence on the test dataset, reaching a final combined result of 0.5151.


机器翻译,仅供参考