今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理8篇。

cs.SD语音

【1】 A Multi-tasking Model of Speaker-Keyword Classification for Keeping  Human in the Loop of Drone-assisted Inspection

标题:无人机辅助检测中人参与的说话人-关键词分类多任务模型

链接:https://arxiv.org/abs/2207.04027

作者:Yu Li,Anisha Parsan,Bill Wang,Penghao Dong,Shanshan Yao,Ruwen Qin
备注:Submitted to Engineering Applications of Artificial Intelligence journal in the end of June 2022. Currently it's under review
摘要:音频命令是一种首选通信媒介,用于让检查员在半自主无人机执行的民用基础设施检查循环中。为了理解来自一组异构和动态检查器的特定于工作的命令,需要为该组开发一个经济高效的模型,并在组发生变化时易于调整。本文旨在构建一个具有共享-拆分-协作架构的多任务深度学习模型。该体系结构允许两个分类任务共享特征提取器,然后通过特征投影和协作训练分割交织在提取特征中的主题特定和关键字特定特征。在本研究收集的检验关键词数据集上,对一组五名授权受试者的基础模型进行了训练和测试。该模型在对任何授权检查员的关键词进行分类时,平均准确率达到95.3%或更高。其说话人分类的平均准确率为99.2%。由于模型从汇集的训练数据中学习到更丰富的关键字表示,使基础模型适应新的检查器只需要该检查器的少量训练数据,例如每个关键字五个句子。使用说话人分类分数进行检查员验证,在验证授权检查员和检测未授权检查员时的成功率分别至少为93.9%和76.1%。此外,本文还证明了该模型适用于公共数据集上较大规模的群体。本文提供了一种解决方案,以解决人工智能辅助人机交互面临的挑战,包括工人异质性、工人动态和工作异质性。
摘要:Audio commands are a preferred communication medium to keep inspectors in the loop of civil infrastructure inspection performed by a semi-autonomous drone. To understand job-specific commands from a group of heterogeneous and dynamic inspectors, a model needs to be developed cost-effectively for the group and easily adapted when the group changes. This paper is motivated to build a multi-tasking deep learning model that possesses a Share-Split-Collaborate architecture. This architecture allows the two classification tasks to share the feature extractor and then split subject-specific and keyword-specific features intertwined in the extracted features through feature projection and collaborative training. A base model for a group of five authorized subjects is trained and tested on the inspection keyword dataset collected by this study. The model achieved a 95.3% or higher mean accuracy in classifying the keywords of any authorized inspectors. Its mean accuracy in speaker classification is 99.2%. Due to the richer keyword representations that the model learns from the pooled training data, adapting the base model to a new inspector requires only a little training data from that inspector, like five utterances per keyword. Using the speaker classification scores for inspector verification can achieve a success rate of at least 93.9% in verifying authorized inspectors and 76.1\% in detecting unauthorized ones. Further, the paper demonstrates the applicability of the proposed model to larger-size groups on a public dataset. This paper provides a solution to addressing challenges facing AI-assisted human-robot interaction, including worker heterogeneity, worker dynamics, and job heterogeneity.


【2】 BAST: Binaural Audio Spectrogram Transformer for Binaural Sound  Localization

标题:BAST:用于双耳声音定位的双耳音谱图转换器

链接:https://arxiv.org/abs/2207.03927

作者:Sheng Kuang,Kiki van der Heijden,Siamak Mehrkanoon
备注:7
摘要:混响环境中准确的声音定位对于人类听觉感知至关重要。最近,卷积神经网络(CNN)已被用于模拟双耳人类听觉通路。然而,CNN在捕捉全球声学特征方面存在障碍。为了解决这个问题,我们提出了一种新的端到端双耳音频频谱变换器(BAST)模型来预测消声和混响环境中的声音方位角。探讨了两种实现模式,即分别对应于具有共享和非共享参数的BAST模型的BAST-SP和BAST-NSP。我们的模型具有减法双耳积分和混合损耗,实现了1.29度的角距离和所有方位角的均方误差1e-3,显著优于基于CNN的模型。对BAST在左右半场、消声和混响环境中的性能进行了探索性分析,表明了其泛化能力以及双耳变换器在声音定位中的可行性。此外,还提供了注意力图的分析,以进一步了解在自然混响环境中对定位过程的解释。
摘要:Accurate sound localization in a reverberation environment is essential for human auditory perception. Recently, Convolutional Neural Networks (CNNs) have been utilized to model the binaural human auditory pathway. However, CNN shows barriers in capturing the global acoustic features. To address this issue, we propose a novel end-to-end Binaural Audio Spectrogram Transformer (BAST) model to predict the sound azimuth in both anechoic and reverberation environments. Two modes of implementation, i.e. BAST-SP and BAST-NSP corresponding to BAST model with shared and non-shared parameters respectively, are explored. Our model with subtraction interaural integration and hybrid loss achieves an angular distance of 1.29 degrees and a Mean Square Error of 1e-3 at all azimuths, significantly surpassing CNN based model. The exploratory analysis of the BAST's performance on the left-right hemifields and anechoic and reverberation environments shows its generalization ability as well as the feasibility of binaural Transformers in sound localization. Furthermore, the analysis of the attention maps is provided to give additional insights on the interpretation of the localization process in a natural reverberant environment.


【3】 FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech  Synthesis

标题:FastLTS:非自回归端到端无约束唇语合成

链接:https://arxiv.org/abs/2207.03800

作者:Yongqi Wang,Zhou Zhao
备注:10 pages, 5 figures, accepted by ACMMM 2022
摘要:无约束唇语音合成的目的是从有声人脸的无声视频中生成相应的语音,而不受头部姿势或词汇量的限制。目前的工作主要使用序列到序列模型来解决这个问题,无论是在自回归架构中还是在基于流的非自回归架构中。然而,这些模型有几个缺点:1)它们没有直接生成音频,而是使用两级管道,首先生成mel频谱,然后从频谱图重建音频。这会由于错误传播而导致部署繁琐和语音质量下降;2) 这些模型使用的音频重建算法限制了推理速度和音频质量,而神经声码器不适用于这些模型,因为它们的输出频谱不够精确;3) 自回归模型具有较高的推理延迟,而基于流的模型具有较高的内存占用率:两者在时间和内存使用方面都不够有效。为了解决这些问题,我们提出了FastLTS,这是一种非自回归端到端模型,它可以从无约束的语音视频中直接合成高质量的语音音频,延迟较低,并且模型尺寸相对较小。此外,与广泛使用的3D-CNN视觉前端进行唇部运动编码不同,我们首次提出了一种基于变换器的视觉前端。实验表明,与当前的自回归模型相比,我们的模型在3秒的输入序列上实现了19.76美元的音频波形生成速度,并获得了优异的音频质量。
摘要:Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to solve this problem, either in an autoregressive architecture or a flow-based non-autoregressive architecture. However, these models suffer from several drawbacks: 1) Instead of directly generating audios, they use a two-stage pipeline that first generates mel-spectrograms and then reconstructs audios from the spectrograms. This causes cumbersome deployment and degradation of speech quality due to error propagation; 2) The audio reconstruction algorithm used by these models limits the inference speed and audio quality, while neural vocoders are not available for these models since their output spectrograms are not accurate enough; 3) The autoregressive model suffers from high inference latency, while the flow-based model has high memory occupancy: neither of them is efficient enough in both time and memory usage. To tackle these problems, we propose FastLTS, a non-autoregressive end-to-end model which can directly synthesize high-quality speech audios from unconstrained talking videos with low latency, and has a relatively small model size. Besides, different from the widely used 3D-CNN visual frontend for lip movement encoding, we for the first time propose a transformer-based visual frontend for this task. Experiments show that our model achieves $19.76\times$ speedup for audio waveform generation compared with the current autoregressive model on input sequences of 3 seconds, and obtains superior audio quality.


【4】 End-to-End Binaural Speech Synthesis

标题:端到端双耳语音合成

链接:https://arxiv.org/abs/2207.03697

作者:Wen Chin Huang,Dejan Markovic,Alexander Richard,Israel Dejene Gebru,Anjali Menon
备注:Accepted to INTERSPEECH 2022. Demo link: this https URL
摘要:在这项工作中,我们提出了一种端到端的双耳语音合成系统,该系统将低比特率音频编解码器与强大的双耳解码器相结合,该解码器能够准确地进行语音双耳化,同时忠实地重建环境因素,如环境噪声或混响。该网络是一种改进的矢量量化变分自动编码器,通过几个精心设计的目标进行训练,包括对抗性损耗。我们在一个具有客观指标和感知研究的内部双耳数据集上评估了所提出的系统。结果表明,该方法比以前的方法更接近地面真实数据。特别是,我们展示了对抗性损失在捕捉创建真实听觉场景所需的环境效果方面的能力。
摘要:In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing environmental factors like ambient noise or reverb. The network is a modified vector-quantized variational autoencoder, trained with several carefully designed objectives, including an adversarial loss. We evaluate the proposed system on an internal binaural dataset with objective metrics and a perceptual study. Results show that the proposed approach matches the ground truth data more closely than previous methods. In particular, we demonstrate the capability of the adversarial loss in capturing environment effects needed to create an authentic auditory scene.


【5】 Tandem Multitask Training of Speaker Diarisation and Speech Recognition  for Meeting Transcription

标题:用于会议记录的说话人发音和语音识别的串联多任务训练

链接:https://arxiv.org/abs/2207.03852

作者:Xianrui Zheng,Chao Zhang,Philip C. Woodland
备注:To appear in Interspeech 2022
摘要:基于自监督学习的语音数据预训练模型,如Wav2Vec 2.0(W2V2),已成为许多语音任务的主干。本文提出了一种串联多任务训练(TMT)方法来微调W2V2,以使用单个模型实现说话人日记和语音识别。对于说话人日记,需要语音活动检测(VAD)和说话人分类(SC)任务,ASR使用连接时序分类(CTC)。多任务框架使用W2V2的早期层、中间层和后期层实现VAD、SC和ASR,这与使用VAD分割音频、基于说话人嵌入聚类片段以及使用ASR转录每个片段的顺序一致。在增强多方(AMI)数据集上的实验结果表明,将不同的W2V2层用于VAD、SC和ASR,从早期层到后期层用于TMT,不仅节省了计算成本,而且还降低了日记错误率(DER)。与单独微调模型的基线相比,通过手动/自动分割,VAD、SC和ASR的联合微调分别使DER相对减少了16%/17%,并一致降低了说话人归因的文字错误率。
摘要:Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a tandem multitask training (TMT) method is proposed to fine-tune W2V2. For speaker diarisation, the tasks of voice activity detection (VAD) and speaker classification (SC) are required, and connectionist temporal classification (CTC) is used for ASR. The multitask framework implements VAD, SC, and ASR using an early layer, middle layer, and late layer of W2V2, which coincides with the order of segmenting the audio with VAD, clustering the segments based on speaker embeddings, and transcribing each segment with ASR. Experimental results on the augmented multi-party (AMI) dataset showed that using different W2V2 layers for VAD, SC, and ASR from the earlier to later layers for TMT not only saves computational cost, but also reduces diarisation error rates (DERs). Joint fine-tuning of VAD, SC, and ASR yielded 16%/17% relative reductions of DER with manual/automatic segmentation respectively, and consistent reductions in speaker attributed word error rate, compared to the baseline with separately fine-tuned models.


【6】 Rhythm and form in music: a complex systems approach

标题:音乐中的节奏和形式:一种复杂的系统方法

链接:https://arxiv.org/abs/2207.03602

作者:Blas Kolic,Mateo Tonatiuh Rodriguez-Cervantes,Pablo Padilla-Longoria,Francis Knights
备注:10 pages, 8 figures
摘要:关于音乐中的形式概念一直有着永恒的讨论。这项工作的动机是这样的辩论,通过使用一个复杂的系统框架,我们研究的形式作为一个紧急属性的节奏。这样一个框架与传统的音乐形式概念相对应,并允许我们将这个概念推广到音乐中更一般的形状和结构。我们开发了以下三个衡量乐曲及其部分节奏复杂性的指标:1)基于排列熵的节奏异质性,其中高值表示各种各样的节奏模式;2) 切分音,基于节拍上音集的分布,其中高值表示非节拍音符的比例高;3)成分提取器,基于节奏图形随时间变化的可见性图的社区,我们在感知水平上识别构成作品的结构成分(待解释)。在参数相同的情况下,我们的指标在一个片段内或片段之间具有可比性。
摘要:There has been an everlasting discussion around the concept of form in music. This work is motivated by such debate by using a complex systems framework in which we study the form as an emergent property of rhythm. Such a framework corresponds with the traditional notion of musical form and allows us to generalize this concept to more general shapes and structures in music. We develop the three following metrics of the rhythmic complexity of a musical piece and its parts: 1) the rhythmic heterogeneity, based on the permutation entropy, where high values indicate a wide variety of rhythmic patterns; 2) the syncopation, based on the distribution of on-beat onsets, where high values indicate a high proportion of off-the-beat notes; and 3) the component extractor, based on the communities of a visibility graph of the rhythmic figures over time, where we identify structural components that constitute the piece at a (to be explained) perceptual level. With the same parameters, our metrics are comparable within a piece or between pieces.


【7】 The ACII 2022 Affective Vocal Bursts Workshop & Competition:  Understanding a critically understudied modality of emotional expression

标题:ACII 2022情绪爆发研讨会与竞赛:理解一种未被充分研究的情绪表达方式

链接:https://arxiv.org/abs/2207.03572

作者:Alice Baird,Panagiotis Tzirakis,Jeffrey A. Brooks,Christopher B. Gregory,Björn Schuller,Anton Batliner,Dacher Keltner,Alan Cowen
摘要:ACII情感性发声爆发研讨会和比赛的重点是了解发声爆发的多个情感维度:笑、喘息、哭、尖叫和许多其他非语言发声,这些发声对情感表达和人类交流更为普遍。今年的比赛包括四首曲目,使用了来自1702名演讲者的59299首歌曲的大规模野外数据集。第一个是A-VB高任务,要求比赛参与者在一个新的情感模型上执行多标签回归,利用十类有丰富注释的情感表达强度,包括:;敬畏、恐惧和惊讶。第二个任务是A-VB-2任务,它利用了更为传统的情感、唤醒和配价的二维模型。第三,A-VB-Culture任务,要求参与者探索数据集的文化方面,训练依赖于本国的模型。最后,对于第四个任务A-VB-Type,参与者应将发声爆发的类型(例如笑、哭、咕噜)识别为8级分类。本文描述了使用最先进的机器学习方法的四个轨迹和基线系统。通过利用端到端深度学习模型获得每个轨迹的基线性能,如下所示:对于A-VB-High,获得了0.5687 CCC的平均(10维以上)一致性相关系数(CCC);对于A-VB-Two,平均(在二维上)CCC为0.5084;对于A-VB-训练,四种训练物的平均CCC为0.4401;对于A-VB类型,8个类别的基线未加权平均召回率(UAR)为0.4172 UAR。
摘要:The ACII Affective Vocal Bursts Workshop & Competition is focused on understanding multiple affective dimensions of vocal bursts: laughs, gasps, cries, screams, and many other non-linguistic vocalizations central to the expression of emotion and to human communication more generally. This year's competition comprises four tracks using a large-scale and in-the-wild dataset of 59,299 vocalizations from 1,702 speakers. The first, the A-VB-High task, requires competition participants to perform a multi-label regression on a novel model for emotion, utilizing ten classes of richly annotated emotional expression intensities, including; Awe, Fear, and Surprise. The second, the A-VB-Two task, utilizes the more conventional 2-dimensional model for emotion, arousal, and valence. The third, the A-VB-Culture task, requires participants to explore the cultural aspects of the dataset, training native-country dependent models. Finally, for the fourth task, A-VB-Type, participants should recognize the type of vocal burst (e.g., laughter, cry, grunt) as an 8-class classification. This paper describes the four tracks and baseline systems, which use state-of-the-art machine learning methods. The baseline performance for each track is obtained by utilizing an end-to-end deep learning model and is as follows: for A-VB-High, a mean (over the 10-dimensions) Concordance Correlation Coefficient (CCC) of 0.5687 CCC is obtained; for A-VB-Two, a mean (over the 2-dimensions) CCC of 0.5084 is obtained; for A-VB-Culture, a mean CCC from the four cultures of 0.4401 is obtained; and for A-VB-Type, the baseline Unweighted Average Recall (UAR) from the 8-classes is 0.4172 UAR.


【8】 BibleTTS: a large, high-fidelity, multilingual, and uniquely African  speech corpus

标题:BibleTTS:一个大型、高保真、多语言、独一无二的非洲语音语料库

链接:https://arxiv.org/abs/2207.03546

作者:Josh Meyer,David Ifeoluwa Adelani,Edresson Casanova,Alp Öktem,Daniel Whitenack Julian Weber,Salomon Kabongo,Elizabeth Salesky,Iroro Orife,Colin Leong,Perez Ogayo,Chris Emezue,Jonathan Mukiibi,Salomey Osei,Apelete Agbolo,Victor Akinode,Bernard Opoku,Samuel Olanrewaju,Jesujoba Alabi,Shamsuddeen Muhammad
备注:Accepted to INTERSPEECH 2022
摘要:BibleTTS是撒哈拉以南非洲十种语言的大型、高质量、开放式语音数据集。语料库包含每种语言多达86小时的对齐录音棚质量48kHz单声道录音,能够开发高质量的文本到语音模型。所代表的十种语言是:阿夸佩姆语、阿桑特语、奇切瓦语、埃维语、豪萨语、基库尤语、林加拉语、卢甘达语、卢奥语和约鲁巴语。该语料库是由公开发行的圣经录音的衍生作品。Biblica的圣经项目。我们对原始录音进行了对齐、清理和过滤,此外还手动检查了每种语言的对齐子集。我们给出了使用Coqui TTS的文本到语音模型的结果。该数据是根据商业友好的CC-BY-SA许可证发布的。
摘要:BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models. The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba. This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica. We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language. We present results for text-to-speech models with Coqui TTS. The data is released under a commercial-friendly CC-BY-SA license.


eess.AS音频处理

【1】 Tandem Multitask Training of Speaker Diarisation and Speech Recognition  for Meeting Transcription

标题:用于会议记录的说话人发音和语音识别的串联多任务训练

链接:https://arxiv.org/abs/2207.03852

* 与cs.SD语音【5】为同一篇

作者:Xianrui Zheng,Chao Zhang,Philip C. Woodland
备注:To appear in Interspeech 2022
摘要:基于自监督学习的语音数据预训练模型,如Wav2Vec 2.0(W2V2),已成为许多语音任务的主干。本文提出了一种串联多任务训练(TMT)方法来微调W2V2,以使用单个模型实现说话人日记和语音识别。对于说话人日记,需要语音活动检测(VAD)和说话人分类(SC)任务,ASR使用连接时序分类(CTC)。多任务框架使用W2V2的早期层、中间层和后期层实现VAD、SC和ASR,这与使用VAD分割音频、基于说话人嵌入聚类片段以及使用ASR转录每个片段的顺序一致。在增强多方(AMI)数据集上的实验结果表明,将不同的W2V2层用于VAD、SC和ASR,从早期层到后期层用于TMT,不仅节省了计算成本,而且还降低了日记错误率(DER)。与单独微调模型的基线相比,通过手动/自动分割,VAD、SC和ASR的联合微调分别使DER相对减少了16%/17%,并一致降低了说话人归因的文字错误率。
摘要:Self-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a tandem multitask training (TMT) method is proposed to fine-tune W2V2. For speaker diarisation, the tasks of voice activity detection (VAD) and speaker classification (SC) are required, and connectionist temporal classification (CTC) is used for ASR. The multitask framework implements VAD, SC, and ASR using an early layer, middle layer, and late layer of W2V2, which coincides with the order of segmenting the audio with VAD, clustering the segments based on speaker embeddings, and transcribing each segment with ASR. Experimental results on the augmented multi-party (AMI) dataset showed that using different W2V2 layers for VAD, SC, and ASR from the earlier to later layers for TMT not only saves computational cost, but also reduces diarisation error rates (DERs). Joint fine-tuning of VAD, SC, and ASR yielded 16%/17% relative reductions of DER with manual/automatic segmentation respectively, and consistent reductions in speaker attributed word error rate, compared to the baseline with separately fine-tuned models.


【2】 Rhythm and form in music: a complex systems approach

标题:音乐中的节奏和形式:一种复杂的系统方法

链接:https://arxiv.org/abs/2207.03602

* 与cs.SD语音【6】为同一篇

作者:Blas Kolic,Mateo Tonatiuh Rodriguez-Cervantes,Pablo Padilla-Longoria,Francis Knights
备注:10 pages, 8 figures
摘要:关于音乐中的形式概念一直有着永恒的讨论。这项工作的动机是这样的辩论,通过使用一个复杂的系统框架,我们研究的形式作为一个紧急属性的节奏。这样一个框架与传统的音乐形式概念相对应,并允许我们将这个概念推广到音乐中更一般的形状和结构。我们开发了以下三个衡量乐曲及其部分节奏复杂性的指标:1)基于排列熵的节奏异质性,其中高值表示各种各样的节奏模式;2) 切分音,基于节拍上音集的分布,其中高值表示非节拍音符的比例高;3)成分提取器,基于节奏图形随时间变化的可见性图的社区,我们在感知水平上识别构成作品的结构成分(待解释)。在参数相同的情况下,我们的指标在一个片段内或片段之间具有可比性。
摘要:There has been an everlasting discussion around the concept of form in music. This work is motivated by such debate by using a complex systems framework in which we study the form as an emergent property of rhythm. Such a framework corresponds with the traditional notion of musical form and allows us to generalize this concept to more general shapes and structures in music. We develop the three following metrics of the rhythmic complexity of a musical piece and its parts: 1) the rhythmic heterogeneity, based on the permutation entropy, where high values indicate a wide variety of rhythmic patterns; 2) the syncopation, based on the distribution of on-beat onsets, where high values indicate a high proportion of off-the-beat notes; and 3) the component extractor, based on the communities of a visibility graph of the rhythmic figures over time, where we identify structural components that constitute the piece at a (to be explained) perceptual level. With the same parameters, our metrics are comparable within a piece or between pieces.


【3】 The ACII 2022 Affective Vocal Bursts Workshop & Competition:  Understanding a critically understudied modality of emotional expression

标题:ACII 2022情绪爆发研讨会与竞赛:理解一种未被充分研究的情绪表达方式

链接:https://arxiv.org/abs/2207.03572

* 与cs.SD语音【7】为同一篇

作者:Alice Baird,Panagiotis Tzirakis,Jeffrey A. Brooks,Christopher B. Gregory,Björn Schuller,Anton Batliner,Dacher Keltner,Alan Cowen
摘要:ACII情感性发声爆发研讨会和比赛的重点是了解发声爆发的多个情感维度:笑、喘息、哭、尖叫和许多其他非语言发声,这些发声对情感表达和人类交流更为普遍。今年的比赛包括四首曲目,使用了来自1702名演讲者的59299首歌曲的大规模野外数据集。第一个是A-VB高任务,要求比赛参与者在一个新的情感模型上执行多标签回归,利用十类有丰富注释的情感表达强度,包括:;敬畏、恐惧和惊讶。第二个任务是A-VB-2任务,它利用了更为传统的情感、唤醒和配价的二维模型。第三,A-VB-Culture任务,要求参与者探索数据集的文化方面,训练依赖于本国的模型。最后,对于第四个任务A-VB-Type,参与者应将发声爆发的类型(例如笑、哭、咕噜)识别为8级分类。本文描述了使用最先进的机器学习方法的四个轨迹和基线系统。通过利用端到端深度学习模型获得每个轨迹的基线性能,如下所示:对于A-VB-High,获得了0.5687 CCC的平均(10维以上)一致性相关系数(CCC);对于A-VB-Two,平均(在二维上)CCC为0.5084;对于A-VB-训练,四种训练物的平均CCC为0.4401;对于A-VB类型,8个类别的基线未加权平均召回率(UAR)为0.4172 UAR。
摘要:The ACII Affective Vocal Bursts Workshop & Competition is focused on understanding multiple affective dimensions of vocal bursts: laughs, gasps, cries, screams, and many other non-linguistic vocalizations central to the expression of emotion and to human communication more generally. This year's competition comprises four tracks using a large-scale and in-the-wild dataset of 59,299 vocalizations from 1,702 speakers. The first, the A-VB-High task, requires competition participants to perform a multi-label regression on a novel model for emotion, utilizing ten classes of richly annotated emotional expression intensities, including; Awe, Fear, and Surprise. The second, the A-VB-Two task, utilizes the more conventional 2-dimensional model for emotion, arousal, and valence. The third, the A-VB-Culture task, requires participants to explore the cultural aspects of the dataset, training native-country dependent models. Finally, for the fourth task, A-VB-Type, participants should recognize the type of vocal burst (e.g., laughter, cry, grunt) as an 8-class classification. This paper describes the four tracks and baseline systems, which use state-of-the-art machine learning methods. The baseline performance for each track is obtained by utilizing an end-to-end deep learning model and is as follows: for A-VB-High, a mean (over the 10-dimensions) Concordance Correlation Coefficient (CCC) of 0.5687 CCC is obtained; for A-VB-Two, a mean (over the 2-dimensions) CCC of 0.5084 is obtained; for A-VB-Culture, a mean CCC from the four cultures of 0.4401 is obtained; and for A-VB-Type, the baseline Unweighted Average Recall (UAR) from the 8-classes is 0.4172 UAR.


【4】 BibleTTS: a large, high-fidelity, multilingual, and uniquely African  speech corpus

标题:BibleTTS:一个大型、高保真、多语言、独一无二的非洲语音语料库

链接:https://arxiv.org/abs/2207.03546

* 与cs.SD语音【8】为同一篇

作者:Josh Meyer,David Ifeoluwa Adelani,Edresson Casanova,Alp Öktem,Daniel Whitenack Julian Weber,Salomon Kabongo,Elizabeth Salesky,Iroro Orife,Colin Leong,Perez Ogayo,Chris Emezue,Jonathan Mukiibi,Salomey Osei,Apelete Agbolo,Victor Akinode,Bernard Opoku,Samuel Olanrewaju,Jesujoba Alabi,Shamsuddeen Muhammad
备注:Accepted to INTERSPEECH 2022
摘要:BibleTTS是撒哈拉以南非洲十种语言的大型、高质量、开放式语音数据集。语料库包含每种语言多达86小时的对齐录音棚质量48kHz单声道录音,能够开发高质量的文本到语音模型。所代表的十种语言是:阿夸佩姆语、阿桑特语、奇切瓦语、埃维语、豪萨语、基库尤语、林加拉语、卢甘达语、卢奥语和约鲁巴语。该语料库是由公开发行的圣经录音的衍生作品。Biblica的圣经项目。我们对原始录音进行了对齐、清理和过滤,此外还手动检查了每种语言的对齐子集。我们给出了使用Coqui TTS的文本到语音模型的结果。该数据是根据商业友好的CC-BY-SA许可证发布的。
摘要:BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models. The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba. This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica. We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language. We present results for text-to-speech models with Coqui TTS. The data is released under a commercial-friendly CC-BY-SA license.


【5】 A Multi-tasking Model of Speaker-Keyword Classification for Keeping  Human in the Loop of Drone-assisted Inspection

标题:无人机辅助检测中人参与的说话人-关键词分类多任务模型

链接:https://arxiv.org/abs/2207.04027

* 与cs.SD语音【1】为同一篇

作者:Yu Li,Anisha Parsan,Bill Wang,Penghao Dong,Shanshan Yao,Ruwen Qin
备注:Submitted to Engineering Applications of Artificial Intelligence journal in the end of June 2022. Currently it's under review
摘要:音频命令是一种首选通信媒介,用于让检查员在半自主无人机执行的民用基础设施检查循环中。为了理解来自一组异构和动态检查器的特定于工作的命令,需要为该组开发一个经济高效的模型,并在组发生变化时易于调整。本文旨在构建一个具有共享-拆分-协作架构的多任务深度学习模型。该体系结构允许两个分类任务共享特征提取器,然后通过特征投影和协作训练分割交织在提取特征中的主题特定和关键字特定特征。在本研究收集的检验关键词数据集上,对一组五名授权受试者的基础模型进行了训练和测试。该模型在对任何授权检查员的关键词进行分类时,平均准确率达到95.3%或更高。其说话人分类的平均准确率为99.2%。由于模型从汇集的训练数据中学习到更丰富的关键字表示,使基础模型适应新的检查器只需要该检查器的少量训练数据,例如每个关键字五个句子。使用说话人分类分数进行检查员验证,在验证授权检查员和检测未授权检查员时的成功率分别至少为93.9%和76.1%。此外,本文还证明了该模型适用于公共数据集上较大规模的群体。本文提供了一种解决方案,以解决人工智能辅助人机交互面临的挑战,包括工人异质性、工人动态和工作异质性。
摘要:Audio commands are a preferred communication medium to keep inspectors in the loop of civil infrastructure inspection performed by a semi-autonomous drone. To understand job-specific commands from a group of heterogeneous and dynamic inspectors, a model needs to be developed cost-effectively for the group and easily adapted when the group changes. This paper is motivated to build a multi-tasking deep learning model that possesses a Share-Split-Collaborate architecture. This architecture allows the two classification tasks to share the feature extractor and then split subject-specific and keyword-specific features intertwined in the extracted features through feature projection and collaborative training. A base model for a group of five authorized subjects is trained and tested on the inspection keyword dataset collected by this study. The model achieved a 95.3% or higher mean accuracy in classifying the keywords of any authorized inspectors. Its mean accuracy in speaker classification is 99.2%. Due to the richer keyword representations that the model learns from the pooled training data, adapting the base model to a new inspector requires only a little training data from that inspector, like five utterances per keyword. Using the speaker classification scores for inspector verification can achieve a success rate of at least 93.9% in verifying authorized inspectors and 76.1\% in detecting unauthorized ones. Further, the paper demonstrates the applicability of the proposed model to larger-size groups on a public dataset. This paper provides a solution to addressing challenges facing AI-assisted human-robot interaction, including worker heterogeneity, worker dynamics, and job heterogeneity.


【6】 BAST: Binaural Audio Spectrogram Transformer for Binaural Sound  Localization

标题:BAST:用于双耳声音定位的双耳音谱图转换器

链接:https://arxiv.org/abs/2207.03927

* 与cs.SD语音【2】为同一篇

作者:Sheng Kuang,Kiki van der Heijden,Siamak Mehrkanoon
备注:7
摘要:混响环境中准确的声音定位对于人类听觉感知至关重要。最近,卷积神经网络(CNN)已被用于模拟双耳人类听觉通路。然而,CNN在捕捉全球声学特征方面存在障碍。为了解决这个问题,我们提出了一种新的端到端双耳音频频谱变换器(BAST)模型来预测消声和混响环境中的声音方位角。探讨了两种实现模式,即分别对应于具有共享和非共享参数的BAST模型的BAST-SP和BAST-NSP。我们的模型具有减法双耳积分和混合损耗,实现了1.29度的角距离和所有方位角的均方误差1e-3,显著优于基于CNN的模型。对BAST在左右半场、消声和混响环境中的性能进行了探索性分析,表明了其泛化能力以及双耳变换器在声音定位中的可行性。此外,还提供了注意力图的分析,以进一步了解在自然混响环境中对定位过程的解释。
摘要:Accurate sound localization in a reverberation environment is essential for human auditory perception. Recently, Convolutional Neural Networks (CNNs) have been utilized to model the binaural human auditory pathway. However, CNN shows barriers in capturing the global acoustic features. To address this issue, we propose a novel end-to-end Binaural Audio Spectrogram Transformer (BAST) model to predict the sound azimuth in both anechoic and reverberation environments. Two modes of implementation, i.e. BAST-SP and BAST-NSP corresponding to BAST model with shared and non-shared parameters respectively, are explored. Our model with subtraction interaural integration and hybrid loss achieves an angular distance of 1.29 degrees and a Mean Square Error of 1e-3 at all azimuths, significantly surpassing CNN based model. The exploratory analysis of the BAST's performance on the left-right hemifields and anechoic and reverberation environments shows its generalization ability as well as the feasibility of binaural Transformers in sound localization. Furthermore, the analysis of the attention maps is provided to give additional insights on the interpretation of the localization process in a natural reverberant environment.


【7】 FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech  Synthesis

标题:FastLTS:非自回归端到端无约束唇语合成

链接:https://arxiv.org/abs/2207.03800

* 与cs.SD语音【3】为同一篇

作者:Yongqi Wang,Zhou Zhao
备注:10 pages, 5 figures, accepted by ACMMM 2022
摘要:无约束唇语音合成的目的是从有声人脸的无声视频中生成相应的语音,而不受头部姿势或词汇量的限制。目前的工作主要使用序列到序列模型来解决这个问题,无论是在自回归架构中还是在基于流的非自回归架构中。然而,这些模型有几个缺点:1)它们没有直接生成音频,而是使用两级管道,首先生成mel频谱,然后从频谱图重建音频。这会由于错误传播而导致部署繁琐和语音质量下降;2) 这些模型使用的音频重建算法限制了推理速度和音频质量,而神经声码器不适用于这些模型,因为它们的输出频谱不够精确;3) 自回归模型具有较高的推理延迟,而基于流的模型具有较高的内存占用率:两者在时间和内存使用方面都不够有效。为了解决这些问题,我们提出了FastLTS,这是一种非自回归端到端模型,它可以从无约束的语音视频中直接合成高质量的语音音频,延迟较低,并且模型尺寸相对较小。此外,与广泛使用的3D-CNN视觉前端进行唇部运动编码不同,我们首次提出了一种基于变换器的视觉前端。实验表明,与当前的自回归模型相比,我们的模型在3秒的输入序列上实现了19.76美元的音频波形生成速度,并获得了优异的音频质量。
摘要:Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to solve this problem, either in an autoregressive architecture or a flow-based non-autoregressive architecture. However, these models suffer from several drawbacks: 1) Instead of directly generating audios, they use a two-stage pipeline that first generates mel-spectrograms and then reconstructs audios from the spectrograms. This causes cumbersome deployment and degradation of speech quality due to error propagation; 2) The audio reconstruction algorithm used by these models limits the inference speed and audio quality, while neural vocoders are not available for these models since their output spectrograms are not accurate enough; 3) The autoregressive model suffers from high inference latency, while the flow-based model has high memory occupancy: neither of them is efficient enough in both time and memory usage. To tackle these problems, we propose FastLTS, a non-autoregressive end-to-end model which can directly synthesize high-quality speech audios from unconstrained talking videos with low latency, and has a relatively small model size. Besides, different from the widely used 3D-CNN visual frontend for lip movement encoding, we for the first time propose a transformer-based visual frontend for this task. Experiments show that our model achieves $19.76\times$ speedup for audio waveform generation compared with the current autoregressive model on input sequences of 3 seconds, and obtains superior audio quality.


【8】 End-to-End Binaural Speech Synthesis

标题:端到端双耳语音合成

链接:https://arxiv.org/abs/2207.03697

* 与cs.SD语音【4】为同一篇

作者:Wen Chin Huang,Dejan Markovic,Alexander Richard,Israel Dejene Gebru,Anjali Menon
备注:Accepted to INTERSPEECH 2022. Demo link: this https URL
摘要:在这项工作中,我们提出了一种端到端的双耳语音合成系统,该系统将低比特率音频编解码器与强大的双耳解码器相结合,该解码器能够准确地进行语音双耳化,同时忠实地重建环境因素,如环境噪声或混响。该网络是一种改进的矢量量化变分自动编码器,通过几个精心设计的目标进行训练,包括对抗性损耗。我们在一个具有客观指标和感知研究的内部双耳数据集上评估了所提出的系统。结果表明,该方法比以前的方法更接近地面真实数据。特别是,我们展示了对抗性损失在捕捉创建真实听觉场景所需的环境效果方面的能力。
摘要:In this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing environmental factors like ambient noise or reverb. The network is a modified vector-quantized variational autoencoder, trained with several carefully designed objectives, including an adversarial loss. We evaluate the proposed system on an internal binaural dataset with objective metrics and a perceptual study. Results show that the proposed approach matches the ground truth data more closely than previous methods. In particular, we demonstrate the capability of the adversarial loss in capturing environment effects needed to create an authentic auditory scene.


机器翻译,仅供参考