今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理9篇。

cs.SD语音

【1】 Radio2Speech: High Quality Speech Recovery from Radio Frequency Signals

标题:Radio2Speech:从射频信号中恢复高质量语音

链接:https://arxiv.org/abs/2206.11066

作者:Running Zhao,Jiangtao Yu,Tingle Li,Hang Zhao,Edith C. H. Ngai
机构:The University of Hong Kong, Hong Kong SAR, China, IIIS, Tsinghua University, Beijing, China, Shanghai Qi Zhi Institute, Shanghai, China
备注:Accepted to INTERSPEECH 2022
摘要:考虑到麦克风很容易受到噪声和隔音材料的影响,射频(RF)信号是恢复音频的一个很有希望的候选者,因为它不受噪声的影响,并且可以穿过许多隔音对象。本文介绍了Radio2Speech,一种使用射频信号从扬声器中恢复高质量语音的系统。Radio2Speech可以恢复与麦克风质量相当的语音,而现有方法只能恢复单音音乐或无法理解的语音。我们使用无线UNet从有限频带的射频信号中准确地恢复时频域中的语音。此外,我们还结合了神经声码器,在不使用污染相位的情况下,根据估计的时频表示合成语音波形。定量和定性评估表明,在安静、嘈杂和隔音的情况下,Radio2Speech实现了最先进的性能,与在安静情况下工作的麦克风不相上下。
摘要:Considering the microphone is easily affected by noise and soundproof materials, the radio frequency (RF) signal is a promising candidate to recover audio as it is immune to noise and can traverse many soundproof objects. In this paper, we introduce Radio2Speech, a system that uses RF signals to recover high quality speech from the loudspeaker. Radio2Speech can recover speech comparable to the quality of the microphone, advancing from recovering only single tone music or incomprehensible speech in existing approaches. We use Radio UNet to accurately recover speech in time-frequency domain from RF signals with limited frequency band. Also, we incorporate the neural vocoder to synthesize the speech waveform from the estimated time-frequency representation without using the contaminated phase. Quantitative and qualitative evaluations show that in quiet, noisy and soundproof scenarios, Radio2Speech achieves state-of-the-art performance and is on par with the microphone that works in quiet scenarios.


【2】 Dynamic Restrained Uncertainty Weighting Loss for Multitask Learning of  Vocal Expression

标题:动态约束不确定性加权损失多任务语音表达学习

链接:https://arxiv.org/abs/2206.11049

作者:Meishu Song,Zijiang Yang,Andreas Triantafyllopoulos,Xin Jing,Vincent Karas,Xie Jiangjian,Zixing Zhang,Yamamoto Yoshiharu,Bjoern W. Schuller

机构:University of Augsburg,  Universityof Tokyo,  Japan  3School of Technology

备注:5 pages

摘要:我们提出了一种新的动态约束不确定性加权损失,以实验方式处理在ICML ExVo 2022挑战中平衡多个任务的贡献的问题。多任务旨在从声音爆发中共同识别所表达的情绪和人口统计学特征。我们的策略结合了不确定性权重和动态权重平均的优点,通过使用约束项扩展权重,使学习过程更具解释性。我们使用一个轻量级的多出口CNN架构来实现我们提出的丢失方法。实验H平均得分(0.394)显示,与基线H平均得分(0.335)相比,有显著改善。

摘要:We propose a novel Dynamic Restrained Uncertainty Weighting Loss to experimentally handle the problem of balancing the contributions of multiple tasks on the ICML ExVo 2022 Challenge. The multitask aims to recognize expressed emotions and demographic traits from vocal bursts jointly. Our strategy combines the advantages of Uncertainty Weight and Dynamic Weight Average, by extending weights with a restraint term to make the learning process more explainable. We use a lightweight multi-exit CNN architecture to implement our proposed loss approach. The experimental H-Mean score (0.394) shows a substantial improvement over the baseline H-Mean score (0.335).


【3】 UniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at  ActivityNet Challenge 2022

标题:UNICON+:ICTCAS-UCAS在2022年ActivityNet挑战赛上提交给AVA-ActiveSpeaker任务

链接:https://arxiv.org/abs/2206.10861

作者:Yuanhang Zhang,Susan Liang,Shuang Yang,Shiguang Shan

机构:Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing , China, School of Computer Science and Technology, University of Chinese Academy of Sciences

备注:5 pages, 3 figures; technical report for AVA Challenge (see this https URL) at the International Challenge on Activity Recognition (ActivityNet), CVPR 2022

摘要:本报告简要介绍了我们在2022年ActivityNet Challenge 2022上成功完成AVA主动说话人检测(ASD)任务的解决方案。我们的基础模型UniCon+继续构建在我们之前的工作基础上,即统一上下文网络(UniCon)和扩展UniCon,它们是为健壮的场景级ASD而设计的。我们使用一个简单的基于GRU的模块来扩充该体系结构,该模块允许重复身份的信息通过读取和更新操作在场景中流动。我们报告的AVA ActiveSpeaker测试集的最佳成绩为94.47%,继续在今年的挑战排行榜上排名第一,显著推动了最先进的水平。

摘要:This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Unified Context Network (UniCon) and Extended UniCon which are designed for robust scene-level ASD. We augment the architecture with a simple GRU-based module that allows information of recurring identities to flow across scenes through read and update operations. We report a best result of 94.47% mAP on the AVA-ActiveSpeaker test set, which continues to rank first on this year's challenge leaderboard and significantly pushes the state-of-the-art.


【4】 Jointist: Joint Learning for Multi-instrument Transcription and Its  Applications

标题:Jointist:多乐器转录的联合学习及其应用

链接:https://arxiv.org/abs/2206.10805

作者:Kin Wai Cheuk,Keunwoo Choi,Qiuqiang Kong,Bochen Li,Minz Won,Amy Hung,Ju-Chiang Wang,Dorien Herremans
机构:Singapore University of Technology and Design,  Gaudio Lab,  Agency for Science, Technology and Researc, Singapore,  ByteDance
备注:Submitted to ISMIR
摘要:在本文中,我们介绍Jointist,这是一个乐器感知的多乐器框架,能够转录、识别和从音频片段中分离多个乐器。Jointist包括调节其他模块的乐器识别模块:输出乐器特定钢琴卷的转录模块,以及利用乐器信息和转录结果的音源分离模块。仪器调节设计用于明确的多仪器功能,而转录和源分离模块之间的连接用于更好的转录性能。鉴于现代流行音乐通常由多种乐器组成,我们提出的具有挑战性的问题公式使得该模型在现实世界中非常有用。然而,它的新颖性要求对如何评估这种模型有一个新的视角。在实验过程中,我们从各个方面对模型进行了评估,为多仪器转录提供了一个新的评估视角。我们还认为,转录模型可以用作其他音乐分析任务的预处理模块。在几个下游任务的实验中,我们的转录模型提供的符号表示结果有助于谱图解决下拍检测、和弦识别和键估计问题。
摘要:In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of the instrument recognition module that conditions the other modules: the transcription module that outputs instrument-specific piano rolls, and the source separation module that utilizes instrument information and transcription results.  The instrument conditioning is designed for an explicit multi-instrument functionality while the connection between the transcription and source separation modules is for better transcription performance. Our challenging problem formulation makes the model highly useful in the real world given that modern popular music typically consists of multiple instruments. However, its novelty necessitates a new perspective on how to evaluate such a model. During the experiment, we assess the model from various aspects, providing a new evaluation perspective for multi-instrument transcription. We also argue that transcription models can be utilized as a preprocessing module for other music analysis tasks. In the experiment on several downstream tasks, the symbolic representation provided by our transcription model turned out to be helpful to spectrograms in solving downbeat detection, chord recognition, and key estimation.


【5】 Exploring the Effectiveness of Self-supervised Learning and Classifier  Chains in Emotion Recognition of Nonverbal Vocalizations

标题:自监督学习和分类器链在非言语发声情绪识别中的有效性研究

链接:https://arxiv.org/abs/2206.10695

作者:Detai Xin,Shinnosuke Takamichi,Hiroshi Saruwatari

机构:ExVo 1The University of Tokyo

备注:Accepted by the ICML Expressive Vocalizations Workshop and Competition 2022

摘要:我们提出了一个非言语发声的情感识别系统(NVs),该系统提交给2022年ICML表达性发声比赛的ExVo少数镜头曲目。该方法使用自监督学习(SSL)模型从NVs中提取特征,并使用分类器链建模情感之间的标签依赖关系。实验结果表明,与几种基线方法相比,该方法可以显著提高该任务的性能。我们提出的方法在验证集中获得的平均一致性相关系数(CCC)为0.725美元,在测试集中为0.739美元,而最佳基线方法在验证集中仅获得0.554美元。我们在https://github.com/Aria-K-Alethia/ExVo帮助他人复制我们的实验结果。

摘要:We present an emotion recognition system for nonverbal vocalizations (NVs) submitted to the ExVo Few-Shot track of the ICML Expressive Vocalizations Competition 2022. The proposed method uses self-supervised learning (SSL) models to extract features from NVs and uses a classifier chain to model the label dependency between emotions. Experimental results demonstrate that the proposed method can significantly improve the performance of this task compared to several baseline methods. Our proposed method obtained a mean concordance correlation coefficient (CCC) of $0.725$ in the validation set and $0.739$ in the test set, while the best baseline method only obtained $0.554$ in the validation set. We publicate our code at https://github.com/Aria-K-Alethia/ExVo to help others to reproduce our experimental results.


【6】 On the Role of Spatial, Spectral, and Temporal Processing for DNN-based  Non-linear Multi-channel Speech Enhancement

标题:空间、频谱和时间处理在基于DNN的非线性多通道语音增强中的作用

链接:https://arxiv.org/abs/2206.11181

作者:Kristina Tesch,Nils-Hendrik Mohrmann,Timo Gerkmann

机构:Signal Processing, Universit¨at Hamburg, Germany

备注:Accepted at Interspeech 2022

摘要:与将线性空间滤波器与独立的速度谱后滤波器相结合的传统方法相比,使用深度神经网络(DNN)直接学习滤波器进行多通道语音增强具有两个潜在的关键优势:1)非线性空间滤波允许克服源于线性处理模型的潜在限制;2)空间滤波的联合处理而节奏谱信息允许利用不同信息源之间的相互依赖性。最近,人们提出了各种基于DNN的非线性滤波器,并报告了其良好的增强性能。然而,人们对将网络架构设计变成一场机会游戏的内部机制知之甚少。因此,在本文中,我们通过实验来更好地理解基于DNN的非线性滤波器对空间、光谱和时间信息的内部处理。一方面,我们在一个困难的语音提取场景中的实验证实了非线性空间滤波的重要性,它比oracle线性空间滤波器的POLQA分数高出0.24。另一方面,我们证明,除了空间信息外,联合处理还可以在利用光谱与时间信息的网络架构之间产生0.4 POLQA分数的巨大性能差距。

摘要:Employing deep neural networks (DNNs) to directly learn filters for multi-channel speech enhancement has potentially two key advantages over a traditional approach combining a linear spatial filter with an independent tempo-spectral post-filter: 1) non-linear spatial filtering allows to overcome potential restrictions originating from a linear processing model and 2) joint processing of spatial and tempo-spectral information allows to exploit interdependencies between different sources of information. A variety of DNN-based non-linear filters have been proposed recently, for which good enhancement performance is reported. However, little is known about the internal mechanisms which turns network architecture design into a game of chance. Therefore, in this paper, we perform experiments to better understand the internal processing of spatial, spectral and temporal information by DNN-based non-linear filters. On the one hand, our experiments in a difficult speech extraction scenario confirm the importance of non-linear spatial filtering, which outperforms an oracle linear spatial filter by 0.24 POLQA score. On the other hand, we demonstrate that joint processing results in a large performance gap of 0.4 POLQA score between network architectures exploiting spectral versus temporal information besides spatial information.


【7】 COVYT: Introducing the Coronavirus YouTube and TikTok speech dataset  featuring the same speakers with and without infection

标题:COVYT:介绍冠状病毒YouTube和TikTok语音数据集,其中包含感染和未感染的相同演讲者

链接:https://arxiv.org/abs/2206.11045

作者:Andreas Triantafyllopoulos,Anastasia Semertzidou,Meishu Song,Florian B. Pokorny,Björn W. Schuller

机构:Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Augsburg, Germany,  Division of Phoniatrics and the Division of Physiology, Medical University of Graz, Graz, Germany

摘要:在疫情爆发两年多后,2019冠状病毒疾病大流行继续困扰着世界各地的医疗系统,给稀缺的资源带来了压力,并夺走了人类的生命。从一开始,人们就在寻求各种基于人工智能的2019冠状病毒疾病检测和监测工具,试图通过及时诊断遏制感染潮。特别是,计算机听觉被认为是一种非侵入性、经济高效、环保的方法,可以通过人声检测2019冠状病毒疾病感染。然而,与所有人工智能方法一样,计算机试听也严重依赖于可用数据的数量和质量,由于此类数据的敏感性,很难获得大规模的2019冠状病毒疾病声音数据集。为此,我们介绍了COVYT数据集,这是一个新的2019冠状病毒疾病数据集,从公共来源收集,包含65位演讲者8个多小时的演讲。与其他现有的2019冠状病毒疾病声音数据集相比,COVYT数据集的独特之处在于,它包含来自所有65名发言者的2019冠状病毒疾病阳性和阴性样本。我们使用可解释的音频描述符,基于这些完美的说话人特征平衡的“野外”数据,分析了2019冠状病毒疾病的声学表现,并研究了几种分类场景,这些场景为基于公平言语的2019冠状病毒疾病检测提供了正确的划分策略。

摘要:More than two years after its outbreak, the COVID-19 pandemic continues to plague medical systems around the world, putting a strain on scarce resources, and claiming human lives. From the very beginning, various AI-based COVID-19 detection and monitoring tools have been pursued in an attempt to stem the tide of infections through timely diagnosis. In particular, computer audition has been suggested as a non-invasive, cost-efficient, and eco-friendly alternative for detecting COVID-19 infections through vocal sounds. However, like all AI methods, also computer audition is heavily dependent on the quantity and quality of available data, and large-scale COVID-19 sound datasets are difficult to acquire -- amongst other reasons -- due to the sensitive nature of such data. To that end, we introduce the COVYT dataset -- a novel COVID-19 dataset collected from public sources containing more than 8 hours of speech from 65 speakers. As compared to other existing COVID-19 sound datasets, the unique feature of the COVYT dataset is that it comprises both COVID-19 positive and negative samples from all 65 speakers. We analyse the acoustic manifestation of COVID-19 on the basis of these perfectly speaker characteristic balanced `in-the-wild' data using interpretable audio descriptors, and investigate several classification scenarios that shed light into proper partitioning strategies for a fair speech-based COVID-19 detection.


【8】 A Systematic Comparison of Phonetic Aware Techniques for Speech  Enhancement

标题:语音增强中语音感知技术的系统比较

链接:https://arxiv.org/abs/2206.11000

作者:Or Tal,Moshe Mandel,Felix Kreuk,Yossi Adi

机构:Hebrew University of Jerusalem, Meta AI Research

备注:Published @ Interspeech 2022

摘要:近年来,使用端到端神经网络进行语音增强取得了很大的进步。然而,大多数模型对口语语音内容是不可知的。最近,一些研究建议语音感知语音增强,主要使用感知监督。然而,在模型优化期间注入语音特征可以采取其他形式(例如,模型调节)。在本文中,我们对在语音增强模型中加入语音信息的不同方法进行了系统的比较。通过一系列的对照实验,我们观察了不同语音内容模型以及各种特征注入技术对增强性能的影响,包括因果模型和非因果模型。具体来说,我们评估了三种注入语音信息的设置,即:i)特征条件;ii)感性监督;和iii)正规化。语音特征是通过有监督的预训练自动语音识别(ASR)模型的中间层或通过预训练的自监督学习(SSL)模型获得的。考虑到手动和学习配置,我们进一步观察了选择不同嵌入层对性能的影响。结果表明,在大多数情况下,使用SSL模型作为语音特征优于ASR模型。有趣的是,在经过评估的配置中,调节设置的性能最好。

摘要:Speech enhancement has seen great improvement in recent years using end-to-end neural networks. However, most models are agnostic to the spoken phonetic content. Recently, several studies suggested phonetic-aware speech enhancement, mostly using perceptual supervision. Yet, injecting phonetic features during model optimization can take additional forms (e.g., model conditioning). In this paper, we conduct a systematic comparison between different methods of incorporating phonetic information in a speech enhancement model. By conducting a series of controlled experiments, we observe the influence of different phonetic content models as well as various feature-injection techniques on enhancement performance, considering both causal and non-causal models. Specifically, we evaluate three settings for injecting phonetic information, namely: i) feature conditioning; ii) perceptual supervision; and iii) regularization. Phonetic features are obtained using an intermediate layer of either a supervised pre-trained Automatic Speech Recognition (ASR) model or by using a pre-trained Self-Supervised Learning (SSL) model. We further observe the effect of choosing different embedding layers on performance, considering both manual and learned configurations. Results suggest that using a SSL model as phonetic features outperforms the ASR one in most cases. Interestingly, the conditioning setting performs best among the evaluated configurations.

eess.AS音频处理

【1】 On the Role of Spatial, Spectral, and Temporal Processing for DNN-based  Non-linear Multi-channel Speech Enhancement

标题:空间、频谱和时间处理在基于DNN的非线性多通道语音增强中的作用

链接:https://arxiv.org/abs/2206.11181

* 与cs.SD语音【6】为同一篇

作者:Kristina Tesch,Nils-Hendrik Mohrmann,Timo Gerkmann
机构:Signal Processing, Universit¨at Hamburg, Germany
备注:Accepted at Interspeech 2022
摘要:与将线性空间滤波器与独立的速度谱后滤波器相结合的传统方法相比,使用深度神经网络(DNN)直接学习滤波器进行多通道语音增强具有两个潜在的关键优势:1)非线性空间滤波允许克服源于线性处理模型的潜在限制;2)空间滤波的联合处理而节奏谱信息允许利用不同信息源之间的相互依赖性。最近,人们提出了各种基于DNN的非线性滤波器,并报告了其良好的增强性能。然而,人们对将网络架构设计变成一场机会游戏的内部机制知之甚少。因此,在本文中,我们通过实验来更好地理解基于DNN的非线性滤波器对空间、光谱和时间信息的内部处理。一方面,我们在一个困难的语音提取场景中的实验证实了非线性空间滤波的重要性,它比oracle线性空间滤波器的POLQA分数高出0.24。另一方面,我们证明,除了空间信息外,联合处理还可以在利用光谱与时间信息的网络架构之间产生0.4 POLQA分数的巨大性能差距。
摘要:Employing deep neural networks (DNNs) to directly learn filters for multi-channel speech enhancement has potentially two key advantages over a traditional approach combining a linear spatial filter with an independent tempo-spectral post-filter: 1) non-linear spatial filtering allows to overcome potential restrictions originating from a linear processing model and 2) joint processing of spatial and tempo-spectral information allows to exploit interdependencies between different sources of information. A variety of DNN-based non-linear filters have been proposed recently, for which good enhancement performance is reported. However, little is known about the internal mechanisms which turns network architecture design into a game of chance. Therefore, in this paper, we perform experiments to better understand the internal processing of spatial, spectral and temporal information by DNN-based non-linear filters. On the one hand, our experiments in a difficult speech extraction scenario confirm the importance of non-linear spatial filtering, which outperforms an oracle linear spatial filter by 0.24 POLQA score. On the other hand, we demonstrate that joint processing results in a large performance gap of 0.4 POLQA score between network architectures exploiting spectral versus temporal information besides spatial information.


【2】 Conformer with dual-mode chunked attention for joint online and offline  ASR

标题:一种线上线下联合ASR的双模块关注整形器

链接:https://arxiv.org/abs/2206.11157

作者:Felix Weninger,Marco Gaudesi,Md Akmal Haidar,Nicola Ferri,Jesús Andrés-Ferrer,Puming Zhan机构:Nuance Communications, Inc.

备注:To appear in INTERSPEECH 2022

摘要:在本文中,我们利用构象传感器对双模(即在线和离线联合)ASR的在线注意机制和蒸馏技术进行了深入研究。在双模Conformer传感器模型中,各层可以在线或离线模式工作,同时共享参数,并在训练中应用从离线到在线模式的就地知识提取,以提高在线精度。在我们的研究中,我们首先证明,与有无前瞻的自回归注意相比,在一致性编码器中使用分块注意可以提高准确性。此外,我们还探讨了知识提取中在线和离线输出之间具有不同位移的有效KLD和1-best KLD损失。最后,我们证明了一个仅具有模式特异性自我注意的简化双模构象与同样具有模式特异性卷积和归一化的双模构象的性能相同。我们的实验基于两个非常不同的数据集:Librispeech任务和内部医学对话语料库。结果表明,与使用平均前瞻性相似的自回归注意的双模系统相比,使用分块注意的双模系统在Librispeech和医疗任务上的相对WER分别提高了5%和4%。

摘要:In this paper, we present an in-depth study on online attention mechanisms and distillation techniques for dual-mode (i.e., joint online and offline) ASR using the Conformer Transducer. In the dual-mode Conformer Transducer model, layers can function in online or offline mode while sharing parameters, and in-place knowledge distillation from offline to online mode is applied in training to improve online accuracy. In our study, we first demonstrate accuracy improvements from using chunked attention in the Conformer encoder compared to autoregressive attention with and without lookahead. Furthermore, we explore the efficient KLD and 1-best KLD losses with different shifts between online and offline outputs in the knowledge distillation. Finally, we show that a simplified dual-mode Conformer that only has mode-specific self-attention performs equally well as the one also having mode-specific convolutions and normalization. Our experiments are based on two very different datasets: the Librispeech task and an internal corpus of medical conversations. Results show that the proposed dual-mode system using chunked attention yields 5% and 4% relative WER improvement on the Librispeech and medical tasks, compared to the dual-mode system using autoregressive attention with similar average lookahead.


【3】 COVYT: Introducing the Coronavirus YouTube and TikTok speech dataset  featuring the same speakers with and without infection

标题:COVYT:介绍冠状病毒YouTube和TikTok语音数据集,其中包含感染和未感染的相同演讲者

链接:https://arxiv.org/abs/2206.11045

* 与cs.SD语音【7】为同一篇

作者:Andreas Triantafyllopoulos,Anastasia Semertzidou,Meishu Song,Florian B. Pokorny,Björn W. Schuller

机构:Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Augsburg, Germany,  Division of Phoniatrics and the Division of Physiology, Medical University of Graz, Graz, Germany

摘要:在疫情爆发两年多后,2019冠状病毒疾病大流行继续困扰着世界各地的医疗系统,给稀缺的资源带来了压力,并夺走了人类的生命。从一开始,人们就在寻求各种基于人工智能的2019冠状病毒疾病检测和监测工具,试图通过及时诊断遏制感染潮。特别是,计算机听觉被认为是一种非侵入性、经济高效、环保的方法,可以通过人声检测2019冠状病毒疾病感染。然而,与所有人工智能方法一样,计算机试听也严重依赖于可用数据的数量和质量,由于此类数据的敏感性,很难获得大规模的2019冠状病毒疾病声音数据集。为此,我们介绍了COVYT数据集,这是一个新的2019冠状病毒疾病数据集,从公共来源收集,包含65位演讲者8个多小时的演讲。与其他现有的2019冠状病毒疾病声音数据集相比,COVYT数据集的独特之处在于,它包含来自所有65名发言者的2019冠状病毒疾病阳性和阴性样本。我们使用可解释的音频描述符,基于这些完美的说话人特征平衡的“野外”数据,分析了2019冠状病毒疾病的声学表现,并研究了几种分类场景,这些场景为基于公平言语的2019冠状病毒疾病检测提供了正确的划分策略。

摘要:More than two years after its outbreak, the COVID-19 pandemic continues to plague medical systems around the world, putting a strain on scarce resources, and claiming human lives. From the very beginning, various AI-based COVID-19 detection and monitoring tools have been pursued in an attempt to stem the tide of infections through timely diagnosis. In particular, computer audition has been suggested as a non-invasive, cost-efficient, and eco-friendly alternative for detecting COVID-19 infections through vocal sounds. However, like all AI methods, also computer audition is heavily dependent on the quantity and quality of available data, and large-scale COVID-19 sound datasets are difficult to acquire -- amongst other reasons -- due to the sensitive nature of such data. To that end, we introduce the COVYT dataset -- a novel COVID-19 dataset collected from public sources containing more than 8 hours of speech from 65 speakers. As compared to other existing COVID-19 sound datasets, the unique feature of the COVYT dataset is that it comprises both COVID-19 positive and negative samples from all 65 speakers. We analyse the acoustic manifestation of COVID-19 on the basis of these perfectly speaker characteristic balanced `in-the-wild' data using interpretable audio descriptors, and investigate several classification scenarios that shed light into proper partitioning strategies for a fair speech-based COVID-19 detection.


【4】 A Systematic Comparison of Phonetic Aware Techniques for Speech  Enhancement

标题:语音增强中语音感知技术的系统比较

链接:https://arxiv.org/abs/2206.11000

* 与cs.SD语音【8】为同一篇

作者:Or Tal,Moshe Mandel,Felix Kreuk,Yossi Adi

机构:Hebrew University of Jerusalem, Meta AI Research

备注:Published @ Interspeech 2022

摘要:近年来,使用端到端神经网络进行语音增强取得了很大的进步。然而,大多数模型对口语语音内容是不可知的。最近,一些研究建议语音感知语音增强,主要使用感知监督。然而,在模型优化期间注入语音特征可以采取其他形式(例如,模型调节)。在本文中,我们对在语音增强模型中加入语音信息的不同方法进行了系统的比较。通过一系列的对照实验,我们观察了不同语音内容模型以及各种特征注入技术对增强性能的影响,包括因果模型和非因果模型。具体来说,我们评估了三种注入语音信息的设置,即:i)特征条件;ii)感性监督;和iii)正规化。语音特征是通过有监督的预训练自动语音识别(ASR)模型的中间层或通过预训练的自监督学习(SSL)模型获得的。考虑到手动和学习配置,我们进一步观察了选择不同嵌入层对性能的影响。结果表明,在大多数情况下,使用SSL模型作为语音特征优于ASR模型。有趣的是,在经过评估的配置中,调节设置的性能最好。

摘要:Speech enhancement has seen great improvement in recent years using end-to-end neural networks. However, most models are agnostic to the spoken phonetic content. Recently, several studies suggested phonetic-aware speech enhancement, mostly using perceptual supervision. Yet, injecting phonetic features during model optimization can take additional forms (e.g., model conditioning). In this paper, we conduct a systematic comparison between different methods of incorporating phonetic information in a speech enhancement model. By conducting a series of controlled experiments, we observe the influence of different phonetic content models as well as various feature-injection techniques on enhancement performance, considering both causal and non-causal models. Specifically, we evaluate three settings for injecting phonetic information, namely: i) feature conditioning; ii) perceptual supervision; and iii) regularization. Phonetic features are obtained using an intermediate layer of either a supervised pre-trained Automatic Speech Recognition (ASR) model or by using a pre-trained Self-Supervised Learning (SSL) model. We further observe the effect of choosing different embedding layers on performance, considering both manual and learned configurations. Results suggest that using a SSL model as phonetic features outperforms the ASR one in most cases. Interestingly, the conditioning setting performs best among the evaluated configurations.


【5】 Radio2Speech: High Quality Speech Recovery from Radio Frequency Signals

标题:Radio2Speech:从射频信号中恢复高质量语音

链接:https://arxiv.org/abs/2206.11066

* 与cs.SD语音【1】为同一篇

作者:Running Zhao,Jiangtao Yu,Tingle Li,Hang Zhao,Edith C. H. Ngai
机构:The University of Hong Kong, Hong Kong SAR, China, IIIS, Tsinghua University, Beijing, China, Shanghai Qi Zhi Institute, Shanghai, China
备注:Accepted to INTERSPEECH 2022
摘要:考虑到麦克风很容易受到噪声和隔音材料的影响,射频(RF)信号是恢复音频的一个很有希望的候选者,因为它不受噪声的影响,并且可以穿过许多隔音对象。本文介绍了Radio2Speech,一种使用射频信号从扬声器中恢复高质量语音的系统。Radio2Speech可以恢复与麦克风质量相当的语音,而现有方法只能恢复单音音乐或无法理解的语音。我们使用无线UNet从有限频带的射频信号中准确地恢复时频域中的语音。此外,我们还结合了神经声码器,在不使用污染相位的情况下,根据估计的时频表示合成语音波形。定量和定性评估表明,在安静、嘈杂和隔音的情况下,Radio2Speech实现了最先进的性能,与在安静情况下工作的麦克风不相上下。
摘要:Considering the microphone is easily affected by noise and soundproof materials, the radio frequency (RF) signal is a promising candidate to recover audio as it is immune to noise and can traverse many soundproof objects. In this paper, we introduce Radio2Speech, a system that uses RF signals to recover high quality speech from the loudspeaker. Radio2Speech can recover speech comparable to the quality of the microphone, advancing from recovering only single tone music or incomprehensible speech in existing approaches. We use Radio UNet to accurately recover speech in time-frequency domain from RF signals with limited frequency band. Also, we incorporate the neural vocoder to synthesize the speech waveform from the estimated time-frequency representation without using the contaminated phase. Quantitative and qualitative evaluations show that in quiet, noisy and soundproof scenarios, Radio2Speech achieves state-of-the-art performance and is on par with the microphone that works in quiet scenarios.


【6】 Dynamic Restrained Uncertainty Weighting Loss for Multitask Learning of  Vocal Expression标题:动态约束不确定性加权损失多任务语音表达学习

链接:https://arxiv.org/abs/2206.11049

* 与cs.SD语音【2】为同一篇

作者:Meishu Song,Zijiang Yang,Andreas Triantafyllopoulos,Xin Jing,Vincent Karas,Xie Jiangjian,Zixing Zhang,Yamamoto Yoshiharu,Bjoern W. Schuller

机构:University of Augsburg,  Universityof Tokyo,  Japan  3School of Technology

备注:5 pages

摘要:我们提出了一种新的动态约束不确定性加权损失,以实验方式处理在ICML ExVo 2022挑战中平衡多个任务的贡献的问题。多任务旨在从声音爆发中共同识别所表达的情绪和人口统计学特征。我们的策略结合了不确定性权重和动态权重平均的优点,通过使用约束项扩展权重,使学习过程更具解释性。我们使用一个轻量级的多出口CNN架构来实现我们提出的丢失方法。实验H平均得分(0.394)显示,与基线H平均得分(0.335)相比,有显著改善。

摘要:We propose a novel Dynamic Restrained Uncertainty Weighting Loss to experimentally handle the problem of balancing the contributions of multiple tasks on the ICML ExVo 2022 Challenge. The multitask aims to recognize expressed emotions and demographic traits from vocal bursts jointly. Our strategy combines the advantages of Uncertainty Weight and Dynamic Weight Average, by extending weights with a restraint term to make the learning process more explainable. We use a lightweight multi-exit CNN architecture to implement our proposed loss approach. The experimental H-Mean score (0.394) shows a substantial improvement over the baseline H-Mean score (0.335).


【7】 UniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at  ActivityNet Challenge 2022

标题:UNICON+:ICTCAS-UCAS在2022年ActivityNet挑战赛上提交给AVA-ActiveSpeaker任务

链接:https://arxiv.org/abs/2206.10861

* 与cs.SD语音【3】为同一篇

作者:Yuanhang Zhang,Susan Liang,Shuang Yang,Shiguang Shan

机构:Key Laboratory of Intelligent Information Processing of Chinese Academy of Sciences (CAS), Institute of Computing Technology, CAS, Beijing , China, School of Computer Science and Technology, University of Chinese Academy of Sciences

备注:5 pages, 3 figures; technical report for AVA Challenge (see this https URL) at the International Challenge on Activity Recognition (ActivityNet), CVPR 2022

摘要:本报告简要介绍了我们在2022年ActivityNet Challenge 2022上成功完成AVA主动说话人检测(ASD)任务的解决方案。我们的基础模型UniCon+继续构建在我们之前的工作基础上,即统一上下文网络(UniCon)和扩展UniCon,它们是为健壮的场景级ASD而设计的。我们使用一个简单的基于GRU的模块来扩充该体系结构,该模块允许重复身份的信息通过读取和更新操作在场景中流动。我们报告的AVA ActiveSpeaker测试集的最佳成绩为94.47%,继续在今年的挑战排行榜上排名第一,显著推动了最先进的水平。

摘要:This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Unified Context Network (UniCon) and Extended UniCon which are designed for robust scene-level ASD. We augment the architecture with a simple GRU-based module that allows information of recurring identities to flow across scenes through read and update operations. We report a best result of 94.47% mAP on the AVA-ActiveSpeaker test set, which continues to rank first on this year's challenge leaderboard and significantly pushes the state-of-the-art.


【8】 Jointist: Joint Learning for Multi-instrument Transcription and Its  Applications

标题:Jointist:多乐器转录的联合学习及其应用

链接:https://arxiv.org/abs/2206.10805

* 与cs.SD语音【4】为同一篇

作者:Kin Wai Cheuk,Keunwoo Choi,Qiuqiang Kong,Bochen Li,Minz Won,Amy Hung,Ju-Chiang Wang,Dorien Herremans

机构:Singapore University of Technology and Design,  Gaudio Lab,  Agency for Science, Technology and Researc, Singapore,  ByteDance

备注:Submitted to ISMIR

摘要:在本文中,我们介绍Jointist,这是一个乐器感知的多乐器框架,能够转录、识别和从音频片段中分离多个乐器。Jointist包括调节其他模块的乐器识别模块:输出乐器特定钢琴卷的转录模块,以及利用乐器信息和转录结果的音源分离模块。仪器调节设计用于明确的多仪器功能,而转录和源分离模块之间的连接用于更好的转录性能。鉴于现代流行音乐通常由多种乐器组成,我们提出的具有挑战性的问题公式使得该模型在现实世界中非常有用。然而,它的新颖性要求对如何评估这种模型有一个新的视角。在实验过程中,我们从各个方面对模型进行了评估,为多仪器转录提供了一个新的评估视角。我们还认为,转录模型可以用作其他音乐分析任务的预处理模块。在几个下游任务的实验中,我们的转录模型提供的符号表示结果有助于谱图解决下拍检测、和弦识别和键估计问题。

摘要:In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of the instrument recognition module that conditions the other modules: the transcription module that outputs instrument-specific piano rolls, and the source separation module that utilizes instrument information and transcription results.  The instrument conditioning is designed for an explicit multi-instrument functionality while the connection between the transcription and source separation modules is for better transcription performance. Our challenging problem formulation makes the model highly useful in the real world given that modern popular music typically consists of multiple instruments. However, its novelty necessitates a new perspective on how to evaluate such a model. During the experiment, we assess the model from various aspects, providing a new evaluation perspective for multi-instrument transcription. We also argue that transcription models can be utilized as a preprocessing module for other music analysis tasks. In the experiment on several downstream tasks, the symbolic representation provided by our transcription model turned out to be helpful to spectrograms in solving downbeat detection, chord recognition, and key estimation.

【9】 Exploring the Effectiveness of Self-supervised Learning and Classifier  Chains in Emotion Recognition of Nonverbal Vocalizations

标题:自监督学习和分类器链在非言语发声情绪识别中的有效性研究

链接:https://arxiv.org/abs/2206.10695

* 与cs.SD语音【5】为同一篇

作者:Detai Xin,Shinnosuke Takamichi,Hiroshi Saruwatari
机构:ExVo 1The University of Tokyo
备注:Accepted by the ICML Expressive Vocalizations Workshop and Competition 2022
摘要:我们提出了一个非言语发声的情感识别系统(NVs),该系统提交给2022年ICML表达性发声比赛的ExVo少数镜头曲目。该方法使用自监督学习(SSL)模型从NVs中提取特征,并使用分类器链建模情感之间的标签依赖关系。实验结果表明,与几种基线方法相比,该方法可以显著提高该任务的性能。我们提出的方法在验证集中获得的平均一致性相关系数(CCC)为0.725美元,在测试集中为0.739美元,而最佳基线方法在验证集中仅获得0.554美元。我们在https://github.com/Aria-K-Alethia/ExVo帮助他人复制我们的实验结果。
摘要:We present an emotion recognition system for nonverbal vocalizations (NVs) submitted to the ExVo Few-Shot track of the ICML Expressive Vocalizations Competition 2022. The proposed method uses self-supervised learning (SSL) models to extract features from NVs and uses a classifier chain to model the label dependency between emotions. Experimental results demonstrate that the proposed method can significantly improve the performance of this task compared to several baseline methods. Our proposed method obtained a mean concordance correlation coefficient (CCC) of $0.725$ in the validation set and $0.739$ in the test set, while the best baseline method only obtained $0.554$ in the validation set. We publicate our code at https://github.com/Aria-K-Alethia/ExVo to help others to reproduce our experimental results.