今天跟大家分享一篇语音相关的论文合集:cs.SD语音9篇,eess.AS音频处理9篇。本文经arXiv每日学术速递授权转载
【1】 Coswara: A website application enabling COVID-19 screening by analysing respiratory sound samples and health symptoms标题:Coswara:一个网站应用程序,可以通过分析呼吸音样本和健康症状进行新冠肺炎筛查
链接:https://arxiv.org/abs/2206.05053
作者:Debarpan Bhattacharya,Debottam Dutta,Neeraj Kumar Sharma,Srikanth Raj Chetupalli,Pravin Mote,Sriram Ganapathy,Chandrakiran C,Sahiti Nori,Suhail K K,Sadhana Gonuguntla,Murali Alagesan机构:LEAP lab, Indian Institute of Science, Bangalore, India,Erlangen, Germany,Ramaiah Medical, College Hospital, Bangalore, India,General Hospital, Hoskote, Bangalore, India,PSG Institute of, Medical Sciences and Research, India摘要:2019冠状病毒疾病大流行加速了对2019冠状病毒疾病替代、快速和有效诊断方法设计的研究。在本文中,我们描述了Coswara工具,这是一个网站应用程序,旨在通过分析呼吸音样本和健康症状来检测2019冠状病毒疾病。使用此服务的用户可以使用连接到internet的任何设备登录网站,提供当前的健康症状信息,并记录与呼吸、咳嗽和说话相对应的少量声音采样。在云服务器上分析此信息后的一分钟内,网站工具将向用户输出2019冠状病毒疾病概率得分。由于2019冠状病毒疾病大流行继续需要大规模和可扩展的人群水平检测,我们假设拟议的工具提供了一个潜在的解决方案。摘要:The COVID-19 pandemic has accelerated research on design of alternative, quick and effective COVID-19 diagnosis approaches. In this paper, we describe the Coswara tool, a website application designed to enable COVID-19 detection by analysing respiratory sound samples and health symptoms. A user using this service can log into a website using any device connected to the internet, provide there current health symptom information and record few sound sampled corresponding to breathing, cough, and speech. Within a minute of analysis of this information on a cloud server the website tool will output a COVID-19 probability score to the user. As the COVID-19 pandemic continues to demand massive and scalable population level testing, we hypothesize that the proposed tool provides a potential solution towards this.
【2】 Going Beyond the Cookie Theft Picture Test: Detecting Cognitive Impairments using Acoustic Features
标题:超越曲奇偷窃图片测试:利用声学特征检测认知障碍
链接:https://arxiv.org/abs/2206.05018
作者:Franziska Braun,Andreas Erzigkeit,Hartmut Lehfeld,Thomas Hillemacher,Korbinian Riedhammer,Sebastian P. Bayerl机构:Technische Hochschule N¨urnberg Georg Simon Ohm, Germany, Psychiatrische Klinik und Psychotherapie, Universit¨atsklinikum Erlangen, Germany, Klinik f¨ur Psychiatrie und Psychotherapie, Universit¨atsklinik der Paracelsus备注:Accepted at the 25th International Conference on Text, Speech and Dialogue (TSD 2022)摘要:标准化测试在认知障碍的检测中起着至关重要的作用。先前的研究表明,使用标准化图片描述任务中的音频数据可以自动检测认知障碍。本研究超越了这一点,评估了我们的方法,这些数据来自两个标准化的神经心理学测试,即德国SKT和德国版的CERAD-NB,以及患者和心理学家之间的半结构化临床访谈。在测试中,我们重点关注三个子测试的语音记录:阅读数字(SKT 3)、干扰(SKT 7)和语言流利性(CERAD-NB 1)。我们表明,标准化测试的声学特征可以可靠地区分认知障碍个体和非认知障碍个体。此外,我们提供的证据表明,即使是从随机的访谈语音样本中提取的特征,也可以作为认知障碍的鉴别指标。在我们的基线实验中,我们使用OpenSMILE特征和支持向量机分类器。在一个改进的设置中,我们展示了使用wav2vec 2.0功能,我们可以实现高达85%的精度。摘要:Standardized tests play a crucial role in the detection of cognitive impairment. Previous work demonstrated that automatic detection of cognitive impairment is possible using audio data from a standardized picture description task. The presented study goes beyond that, evaluating our methods on data taken from two standardized neuropsychological tests, namely the German SKT and a German version of the CERAD-NB, and a semi-structured clinical interview between a patient and a psychologist. For the tests, we focus on speech recordings of three sub-tests: reading numbers (SKT 3), interference (SKT 7), and verbal fluency (CERAD-NB 1). We show that acoustic features from standardized tests can be used to reliably discriminate cognitively impaired individuals from non-impaired ones. Furthermore, we provide evidence that even features extracted from random speech samples of the interview can be a discriminator of cognitive impairment. In our baseline experiments, we use OpenSMILE features and Support Vector Machine classifiers. In an improved setup, we show that using wav2vec 2.0 features instead, we can achieve an accuracy of up to 85%.
【3】 Zero-Shot Audio Classification using Image Embeddings
标题:基于图像嵌入的Zero-Shot音频分类
链接:https://arxiv.org/abs/2206.04984
作者:Duygu Dogan,Huang Xie,Toni Heittola,Tuomas Virtanen机构:∗Unit of Computing Sciences, Tampere University, Finland备注:Accepted to the European Signal Processing Conference (EUSIPCO) 2022摘要:监督学习方法可以在存在大量标记数据的情况下解决给定的问题。然而,获取覆盖所有目标类的数据集通常需要手动标记,这既昂贵又耗时。零炮学习模型能够利用不可见概念的语义信息对其进行分类。本研究利用非线性声学语义投影将图像嵌入作为零炮音频分类的辅助信息。我们从开放的图像数据集中提取语义图像表示,并使用不同领域的语义信息评估模型在音频集的音频子集上的性能;图像、音频和文本。我们证明了图像嵌入可以作为语义信息来执行Zero-Shot音频分类。实验结果表明,图像嵌入和文本嵌入无论是单独嵌入还是共同嵌入都表现出相似的性能。我们还从测试样本中计算语义声学嵌入,以提供性能上限。结果表明,分类性能对测试类和训练类之间的语义关系非常敏感,当看到的类和看不到的类语义相似时,文本和图像嵌入可以达到语义声学嵌入。摘要:Supervised learning methods can solve the given problem in the presence of a large set of labeled data. However, the acquisition of a dataset covering all the target classes typically requires manual labeling which is expensive and time-consuming. Zero-shot learning models are capable of classifying the unseen concepts by utilizing their semantic information. The present study introduces image embeddings as side information on zero-shot audio classification by using a nonlinear acoustic-semantic projection. We extract the semantic image representations from the Open Images dataset and evaluate the performance of the models on an audio subset of AudioSet using semantic information in different domains; image, audio, and textual. We demonstrate that the image embeddings can be used as semantic information to perform zero-shot audio classification. The experimental results show that the image and textual embeddings display similar performance both individually and together. We additionally calculate the semantic acoustic embeddings from the test samples to provide an upper limit to the performance. The results show that the classification performance is highly sensitive to the semantic relation between test and training classes and textual and image embeddings can reach up to the semantic acoustic embeddings when the seen and unseen classes are semantically similar.
【4】 Feature Learning and Ensemble Pre-Tasks Based Self-Supervised Speech Denoising and Dereverberation
标题:基于特征学习和集成预任务的自监督语音去噪和混响
链接:https://arxiv.org/abs/2206.04962
作者:Yi Li,ShuangLin Li,Yang Sun,Syed Mohsen Naqvi备注:arXiv admin note: text overlap with arXiv:2112.11142摘要:自监督学习(SSL)在单耳语音增强方面取得了巨大的成功,而目标语音估计的准确性,尤其是对于看不见的说话人,在现有的预任务中仍然不够。由于语音信号包含说话人身份、副语言学和语音内容等多方面的信息,语音增强的潜在表征成为一项艰巨的任务。在本文中,我们研究了语音增强中常用的每个特征的有效性,并在SSL情况下利用了这些特征的组合。此外,我们还提出了一种集成训练策略。学习干净语音信号的潜在表示,同时利用去冗余掩模和估计比率掩模对混合信号进行去噪和去冗余。在训练阶段,将潜在表征学习和掩码估计作为两个预处理任务。此外,为了研究预任务之间的有效性,我们比较了不同的训练例程来训练模型并进一步完善性能。使用NOISEX和DAPS语料库对所提方法的有效性进行了评估,其效果也优于最新的方法。摘要:Self-supervised learning (SSL) achieves great success in monaural speech enhancement, while the accuracy of the target speech estimation, particularly for unseen speakers, remains inadequate with existing pre-tasks. As speech signal contains multi-faceted information including speaker identity, paralinguistics, and spoken content, the latent representation for speech enhancement becomes a tough task. In this paper, we study the effectiveness of each feature which is commonly used in speech enhancement and exploit the feature combination in the SSL case. Besides, we propose an ensemble training strategy. The latent representation of the clean speech signal is learned, meanwhile, the dereverberated mask and the estimated ratio mask are exploited to denoise and dereverberate the mixture. The latent representation learning and the masks estimation are considered as two pre-tasks in the training stage. In addition, to study the effectiveness between the pre-tasks, we compare different training routines to train the model and further refine the performance. The NOISEX and DAPS corpora are used to evaluate the efficacy of the proposed method, which also outperforms the state-of-the-art methods.
【5】 A Novel Chinese Dialect TTS Frontend with Non-Autoregressive Neural Machine Translation
标题:一种新的非自回归神经机器翻译汉语方言TTS前端
链接:https://arxiv.org/abs/2206.04922
作者:Wudi Bao,Junhui Zhang,Junjie Pan,Xiang Yin注:Submitted to INTERSPEECH 2022, 5 pages,5 figures摘要:汉语方言文语转换(TTS)系统通常只能由母语语言学家使用,因为汉语方言的书面形式与普通话有着不同的字符、习语、语法和用法,甚至当地人也无法输入正确的句子。对于普通话文本输入,汉语方言TTS只能生成部分有意义的语音,韵律和自然度相对较差。为了降低使用门槛,使其在商业上更加实用,我们提出了一种带有翻译模块的汉语方言TTS前端。它有助于将汉语文本转换成具有正确拼写和语法的惯用表达,从而提高合成语音的清晰度和自然度。针对翻译任务,提出了一种具有扫描采样策略的非自回归神经机器翻译模型。这是第一个将翻译与TTS前端相结合的已知作品。我们在广东话上的实验表明,所提出的前端可以帮助广东话TTS系统在使用普通话输入时实现0.27的MOS改进。摘要:Chinese dialect text-to-speech(TTS) system usually can only be utilized by native linguists, because the written form of Chinese dialects has different characters, idioms, grammar and usage from Mandarin, and even the local speaker cannot input a correct sentence. For Mandarin text inputs, Chinese dialect TTS can only generate partly-meaningful speech with relatively poor prosody and naturalness. To lower the bar of use and make it more practical in commercial, we propose a novel Chinese dialect TTS frontend with a translation module. It helps to convert Mandarin text into idiomatic expressions with correct orthography and grammar, so that the intelligibility and naturalness of the synthesized speech can be improved. A non-autoregressive neural machine translation model with a glancing sampling strategy is proposed for the translation task. It is the first known work to incorporate translation with TTS frontend. Our experiments on Cantonese approve that the proposed frontend can help Cantonese TTS system achieve a 0.27 improvement in MOS with Mandarin inputs.
【6】 Motif Mining and Unsupervised Representation Learning for BirdCLEF 2022
标题:BirdCLEF 2022的Motif挖掘和无监督表示学习
链接:https://arxiv.org/abs/2206.04805
作者:Anthony Miyaguchi,Jiangyue Yu,Bryan Cheungvivatpant,Dakota Dudley,Aniketh Swain机构:Georgia Institute of Technology, North Ave NW, Atlanta, GA备注:Submitted to CEUR-WS under LifeCLEF for the BirdCLEF 2022 challenge as a working note摘要:我们使用无监督的方法为BirdCLEF 2022挑战建立了一个分类模型。我们使用音频基序的三重丢失谱图表示实现了训练数据集的无监督表示。我们的最佳模型在公共排行榜上的得分为0.48。摘要:We build a classification model for the BirdCLEF 2022 challenge using unsupervised methods. We implement an unsupervised representation of the training dataset using a triplet loss on spectrogram representation of audio motifs. Our best model performs with a score of 0.48 on the public leaderboard.
【7】 Speak Like a Dog: Human to Non-human creature Voice Conversion
标题:像狗一样说话:人类到非人类生物的声音转换
链接:https://arxiv.org/abs/2206.04780
作者:Kohei Suzuki,Shoki Sakamoto,Tadahiro Taniguchi,Hirokazu Kameoka构:College of Information Science and Engineering, Ritsumeikan University, Japan, NTT Communication Science Laboratories, NTT Corporation, Japan摘要:以人到非人的语音转换(H2NH-VC)任务为例,提出了一种新的语音转换(VC)任务,在保留语言信息的同时,将人类语音转换为类狗语音。虽然大多数VC研究都涉及人与人之间的VC,但H2NH-VC旨在将人类的语音转换为类似于非人类生物的语音。非并行VC允许我们开发H2NH-VC,因为我们无法收集非人类生物讲人类语言的并行数据集。在这项研究中,我们建议将狗作为非人生物目标域的一个示例,并定义“像狗一样说话”任务。为了阐明“像狗一样说话”任务的可能性和特点,我们在声学特征(Mel倒谱系数和Mel频谱图)、网络结构(五种不同的内核大小设置)和,和训练标准(基于变分自动编码器(VAE)和基于生成对抗网络)。最后,使用平均意见分数评估转换后的声音:狗的相似性、音质和清晰度以及字符错误率(CER)。实验表明,Mel谱图的使用改善了转换语音的狗的相似性,但保留语言信息却很困难。重点介绍了H2NH-VC当前VC方法的挑战和局限性。摘要:This paper proposes a new voice conversion (VC) task from human speech to dog-like speech while preserving linguistic information as an example of human to non-human creature voice conversion (H2NH-VC) tasks. Although most VC studies deal with human to human VC, H2NH-VC aims to convert human speech into non-human creature-like speech. Non-parallel VC allows us to develop H2NH-VC, because we cannot collect a parallel dataset that non-human creatures speak human language. In this study, we propose to use dogs as an example of a non-human creature target domain and define the "speak like a dog" task. To clarify the possibilities and characteristics of the "speak like a dog" task, we conducted a comparative experiment using existing representative non-parallel VC methods in acoustic features (Mel-cepstral coefficients and Mel-spectrograms), network architectures (five different kernel-size settings), and training criteria (variational autoencoder (VAE)- based and generative adversarial network-based). Finally, the converted voices were evaluated using mean opinion scores: dog-likeness, sound quality and intelligibility, and character error rate (CER). The experiment showed that the employment of the Mel-spectrogram improved the dog-likeness of the converted speech, while it is challenging to preserve linguistic information. Challenges and limitations of the current VC methods for H2NH-VC are highlighted.
【8】 CLAP: Learning Audio Concepts From Natural Language Supervision
标题:CLAP:从自然语言监督中学习音频概念
链接:https://arxiv.org/abs/2206.04769
作者:Benjamin Elizalde,Soham Deshmukh,Mahmoud Al Ismail,Huaming Wang摘要:主流音频分析模型经过训练,可以在一个类别标签到多个录音的范式下学习,重点是一项任务。在这种受限的监督下学习限制了模型的灵活性,因为它们需要标记音频进行训练,并且只能预测预定义的类别。相反,我们建议从自然语言监控中学习音频概念。我们称之为对比语言音频预训练(CRAP),该方法通过使用两个编码器和对比学习将语言和音频连接起来,将音频和文本描述带入一个联合的多模态空间。我们使用128k音频和文本对训练CLAP,并在8个领域的16个下游任务中对其进行评估,如声音事件分类、音乐任务和语音相关任务。虽然与类似的计算机视觉模型相比,CLAP的训练成对数明显减少,但它为零炮性能建立了SoTA。此外,我们在监督学习环境中评估了CLAP,并在5项任务中实现了SoTA。因此,CLAP的零炮能力消除了使用类标签进行训练的需要,在推理时实现了灵活的类预测,并推广到多个下游任务。摘要:Mainstream Audio Analytics models are trained to learn under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the predefined categories. Instead, we propose to learn audio concepts from natural language supervision. We call our approach Contrastive Language-Audio Pretraining (CLAP), which learns to connect language and audio by using two encoders and a contrastive learning to bring audio and text descriptions into a joint multimodal space. We trained CLAP with 128k audio and text pairs and evaluated it on 16 downstream tasks across 8 domains, such as Sound Event Classification, Music tasks, and Speech-related tasks. Although CLAP was trained with significantly less pairs than similar computer vision models, it establishes SoTA for Zero-Shot performance. Additionally, we evaluated CLAP in a supervised learning setup and achieve SoTA in 5 tasks. Hence, CLAP's Zero-Shot capability removes the need of training with class labels, enables flexible class prediction at inference time, and generalizes to multiple downstream tasks.
【9】 Feature-informed Embedding Space Regularization For Audio Classification
标题:基于特征信息的嵌入空间正则化音频分类
链接:https://arxiv.org/abs/2206.04850
作者:Yun-Ning Hung,Alexander Lerch机构:Speech, Audio & Music Intelligence Team, ByteDance, Music Informatics Group, Georgia Institute of Technology摘要:从在大规模数据集上预先训练的模型中得到的特征表示已显示出其在各种音频分析任务中的通用性。尽管有这种普遍性,但是,如果有足够的训练数据可用,任务特定的特性可以表现得更好,因为可以学习特定的任务相关属性。此外,复杂的预训练模型在推理过程中会带来相当大的计算负担。我们建议通过引入两种整合两种特征类信息的正则化方法,利用来自谱图输入的详细任务特定特征和一般预训练特征。在推理过程中,工作量保持较低,因为预训练的特征只在训练时需要。在对预训练特征VGGish、OpenL3以及两者的结合进行的实验中,我们表明,所提出的方法不仅优于基线方法,而且可以改进多个音频分类任务的最新模型。结果还表明,使用混合特征比使用单个特征表现更好。摘要:Feature representations derived from models pre-trained on large-scale datasets have shown their generalizability on a variety of audio analysis tasks. Despite this generalizability, however, task-specific features can outperform if sufficient training data is available, as specific task-relevant properties can be learned. Furthermore, the complex pre-trained models bring considerable computational burdens during inference. We propose to leverage both detailed task-specific features from spectrogram input and generic pre-trained features by introducing two regularization methods that integrate the information of both feature classes. The workload is kept low during inference as the pre-trained features are only necessary for training. In experiments with the pre-trained features VGGish, OpenL3, and a combination of both, we show that the proposed methods not only outperform baseline methods, but also can improve state-of-the-art models on several audio classification tasks. The results also suggest that using the mixture of features performs better than using individual features.
【1】 Feature-informed Embedding Space Regularization For Audio Classification标题:基于特征信息的嵌入空间正则化音频分类
链接:https://arxiv.org/abs/2206.04850
作者:Yun-Ning Hung,Alexander Lerch机构:Speech, Audio & Music Intelligence Team, ByteDance, Music Informatics Group, Georgia Institute of Technology摘要:从在大规模数据集上预先训练的模型中得到的特征表示已显示出其在各种音频分析任务中的通用性。尽管有这种普遍性,但是,如果有足够的训练数据可用,任务特定的特性可以表现得更好,因为可以学习特定的任务相关属性。此外,复杂的预训练模型在推理过程中会带来相当大的计算负担。我们建议通过引入两种整合两种特征类信息的正则化方法,利用来自谱图输入的详细任务特定特征和一般预训练特征。在推理过程中,工作量保持较低,因为预训练的特征只在训练时需要。在对预训练特征VGGish、OpenL3以及两者的结合进行的实验中,我们表明,所提出的方法不仅优于基线方法,而且可以改进多个音频分类任务的最新模型。结果还表明,使用混合特征比使用单个特征表现更好。摘要:Feature representations derived from models pre-trained on large-scale datasets have shown their generalizability on a variety of audio analysis tasks. Despite this generalizability, however, task-specific features can outperform if sufficient training data is available, as specific task-relevant properties can be learned. Furthermore, the complex pre-trained models bring considerable computational burdens during inference. We propose to leverage both detailed task-specific features from spectrogram input and generic pre-trained features by introducing two regularization methods that integrate the information of both feature classes. The workload is kept low during inference as the pre-trained features are only necessary for training. In experiments with the pre-trained features VGGish, OpenL3, and a combination of both, we show that the proposed methods not only outperform baseline methods, but also can improve state-of-the-art models on several audio classification tasks. The results also suggest that using the mixture of features performs better than using individual features.
【2】 Coswara: A website application enabling COVID-19 screening by analysing respiratory sound samples and health symptoms
标题:Coswara:一个网站应用程序,可以通过分析呼吸音样本和健康症状进行新冠肺炎筛查
链接:https://arxiv.org/abs/2206.05053
作者:Debarpan Bhattacharya,Debottam Dutta,Neeraj Kumar Sharma,Srikanth Raj Chetupalli,Pravin Mote,Sriram Ganapathy,Chandrakiran C,Sahiti Nori,Suhail K K,Sadhana Gonuguntla,Murali Alagesan机构:LEAP lab, Indian Institute of Science, Bangalore, India,Erlangen, Germany,Ramaiah Medical, College Hospital, Bangalore, India,General Hospital, Hoskote, Bangalore, India,PSG Institute of, Medical Sciences and Research, India摘要:2019冠状病毒疾病大流行加速了对2019冠状病毒疾病替代、快速和有效诊断方法设计的研究。在本文中,我们描述了Coswara工具,这是一个网站应用程序,旨在通过分析呼吸音样本和健康症状来检测2019冠状病毒疾病。使用此服务的用户可以使用连接到internet的任何设备登录网站,提供当前的健康症状信息,并记录与呼吸、咳嗽和说话相对应的少量声音采样。在云服务器上分析此信息后的一分钟内,网站工具将向用户输出2019冠状病毒疾病概率得分。由于2019冠状病毒疾病大流行继续需要大规模和可扩展的人群水平检测,我们假设拟议的工具提供了一个潜在的解决方案。摘要:The COVID-19 pandemic has accelerated research on design of alternative, quick and effective COVID-19 diagnosis approaches. In this paper, we describe the Coswara tool, a website application designed to enable COVID-19 detection by analysing respiratory sound samples and health symptoms. A user using this service can log into a website using any device connected to the internet, provide there current health symptom information and record few sound sampled corresponding to breathing, cough, and speech. Within a minute of analysis of this information on a cloud server the website tool will output a COVID-19 probability score to the user. As the COVID-19 pandemic continues to demand massive and scalable population level testing, we hypothesize that the proposed tool provides a potential solution towards this.
【3】 Going Beyond the Cookie Theft Picture Test: Detecting Cognitive Impairments using Acoustic Features
标题:超越曲奇偷窃图片测试:利用声学特征检测认知障碍
链接:https://arxiv.org/abs/2206.05018
作者:Franziska Braun,Andreas Erzigkeit,Hartmut Lehfeld,Thomas Hillemacher,Korbinian Riedhammer,Sebastian P. Bayerl机构:Technische Hochschule N¨urnberg Georg Simon Ohm, Germany, Psychiatrische Klinik und Psychotherapie, Universit¨atsklinikum Erlangen, Germany, Klinik f¨ur Psychiatrie und Psychotherapie, Universit¨atsklinik der Paracelsus备注:Accepted at the 25th International Conference on Text, Speech and Dialogue (TSD 2022)摘要:标准化测试在认知障碍的检测中起着至关重要的作用。先前的研究表明,使用标准化图片描述任务中的音频数据可以自动检测认知障碍。本研究超越了这一点,评估了我们的方法,这些数据来自两个标准化的神经心理学测试,即德国SKT和德国版的CERAD-NB,以及患者和心理学家之间的半结构化临床访谈。在测试中,我们重点关注三个子测试的语音记录:阅读数字(SKT 3)、干扰(SKT 7)和语言流利性(CERAD-NB 1)。我们表明,标准化测试的声学特征可以可靠地区分认知障碍个体和非认知障碍个体。此外,我们提供的证据表明,即使是从随机的访谈语音样本中提取的特征,也可以作为认知障碍的鉴别指标。在我们的基线实验中,我们使用OpenSMILE特征和支持向量机分类器。在一个改进的设置中,我们展示了使用wav2vec 2.0功能,我们可以实现高达85%的精度。摘要:Standardized tests play a crucial role in the detection of cognitive impairment. Previous work demonstrated that automatic detection of cognitive impairment is possible using audio data from a standardized picture description task. The presented study goes beyond that, evaluating our methods on data taken from two standardized neuropsychological tests, namely the German SKT and a German version of the CERAD-NB, and a semi-structured clinical interview between a patient and a psychologist. For the tests, we focus on speech recordings of three sub-tests: reading numbers (SKT 3), interference (SKT 7), and verbal fluency (CERAD-NB 1). We show that acoustic features from standardized tests can be used to reliably discriminate cognitively impaired individuals from non-impaired ones. Furthermore, we provide evidence that even features extracted from random speech samples of the interview can be a discriminator of cognitive impairment. In our baseline experiments, we use OpenSMILE features and Support Vector Machine classifiers. In an improved setup, we show that using wav2vec 2.0 features instead, we can achieve an accuracy of up to 85%.
【4】 Zero-Shot Audio Classification using Image Embeddings
标题:基于图像嵌入的Zero-Shot音频分类
链接:https://arxiv.org/abs/2206.04984
作者:Duygu Dogan,Huang Xie,Toni Heittola,Tuomas Virtanen机构:∗Unit of Computing Sciences, Tampere University, Finland备注:Accepted to the European Signal Processing Conference (EUSIPCO) 2022摘要:监督学习方法可以在存在大量标记数据的情况下解决给定的问题。然而,获取覆盖所有目标类的数据集通常需要手动标记,这既昂贵又耗时。零炮学习模型能够利用不可见概念的语义信息对其进行分类。本研究利用非线性声学语义投影将图像嵌入作为零炮音频分类的辅助信息。我们从开放的图像数据集中提取语义图像表示,并使用不同领域的语义信息评估模型在音频集的音频子集上的性能;图像、音频和文本。我们证明了图像嵌入可以作为语义信息来执行Zero-Shot音频分类。实验结果表明,图像嵌入和文本嵌入无论是单独嵌入还是共同嵌入都表现出相似的性能。我们还从测试样本中计算语义声学嵌入,以提供性能上限。结果表明,分类性能对测试类和训练类之间的语义关系非常敏感,当看到的类和看不到的类语义相似时,文本和图像嵌入可以达到语义声学嵌入。摘要:Supervised learning methods can solve the given problem in the presence of a large set of labeled data. However, the acquisition of a dataset covering all the target classes typically requires manual labeling which is expensive and time-consuming. Zero-shot learning models are capable of classifying the unseen concepts by utilizing their semantic information. The present study introduces image embeddings as side information on zero-shot audio classification by using a nonlinear acoustic-semantic projection. We extract the semantic image representations from the Open Images dataset and evaluate the performance of the models on an audio subset of AudioSet using semantic information in different domains; image, audio, and textual. We demonstrate that the image embeddings can be used as semantic information to perform zero-shot audio classification. The experimental results show that the image and textual embeddings display similar performance both individually and together. We additionally calculate the semantic acoustic embeddings from the test samples to provide an upper limit to the performance. The results show that the classification performance is highly sensitive to the semantic relation between test and training classes and textual and image embeddings can reach up to the semantic acoustic embeddings when the seen and unseen classes are semantically similar.
【5】 Feature Learning and Ensemble Pre-Tasks Based Self-Supervised Speech Denoising and Dereverberation
标题:基于特征学习和集成预任务的自监督语音去噪和混响
链接:https://arxiv.org/abs/2206.04962
作者:Yi Li,ShuangLin Li,Yang Sun,Syed Mohsen Naqvi备注:arXiv admin note: text overlap with arXiv:2112.11142摘要:自监督学习(SSL)在单耳语音增强方面取得了巨大的成功,而目标语音估计的准确性,尤其是对于看不见的说话人,在现有的预任务中仍然不够。由于语音信号包含说话人身份、副语言学和语音内容等多方面的信息,语音增强的潜在表征成为一项艰巨的任务。在本文中,我们研究了语音增强中常用的每个特征的有效性,并在SSL情况下利用了这些特征的组合。此外,我们还提出了一种集成训练策略。学习干净语音信号的潜在表示,同时利用去冗余掩模和估计比率掩模对混合信号进行去噪和去冗余。在训练阶段,将潜在表征学习和掩码估计作为两个预处理任务。此外,为了研究预任务之间的有效性,我们比较了不同的训练例程来训练模型并进一步完善性能。使用NOISEX和DAPS语料库对所提方法的有效性进行了评估,其效果也优于最新的方法。摘要:Self-supervised learning (SSL) achieves great success in monaural speech enhancement, while the accuracy of the target speech estimation, particularly for unseen speakers, remains inadequate with existing pre-tasks. As speech signal contains multi-faceted information including speaker identity, paralinguistics, and spoken content, the latent representation for speech enhancement becomes a tough task. In this paper, we study the effectiveness of each feature which is commonly used in speech enhancement and exploit the feature combination in the SSL case. Besides, we propose an ensemble training strategy. The latent representation of the clean speech signal is learned, meanwhile, the dereverberated mask and the estimated ratio mask are exploited to denoise and dereverberate the mixture. The latent representation learning and the masks estimation are considered as two pre-tasks in the training stage. In addition, to study the effectiveness between the pre-tasks, we compare different training routines to train the model and further refine the performance. The NOISEX and DAPS corpora are used to evaluate the efficacy of the proposed method, which also outperforms the state-of-the-art methods.
【6】 A Novel Chinese Dialect TTS Frontend with Non-Autoregressive Neural Machine Translation
标题:一种新的非自回归神经机器翻译汉语方言TTS前端
链接:https://arxiv.org/abs/2206.04922
作者:Wudi Bao,Junhui Zhang,Junjie Pan,Xiang Yin备注:Submitted to INTERSPEECH 2022, 5 pages,5 figures摘要:汉语方言文语转换(TTS)系统通常只能由母语语言学家使用,因为汉语方言的书面形式与普通话有着不同的字符、习语、语法和用法,甚至当地人也无法输入正确的句子。对于普通话文本输入,汉语方言TTS只能生成部分有意义的语音,韵律和自然度相对较差。为了降低使用门槛,使其在商业上更加实用,我们提出了一种带有翻译模块的汉语方言TTS前端。它有助于将汉语文本转换成具有正确拼写和语法的惯用表达,从而提高合成语音的清晰度和自然度。针对翻译任务,提出了一种具有扫描采样策略的非自回归神经机器翻译模型。这是第一个将翻译与TTS前端相结合的已知作品。我们在广东话上的实验表明,所提出的前端可以帮助广东话TTS系统在使用普通话输入时实现0.27的MOS改进。摘要:Chinese dialect text-to-speech(TTS) system usually can only be utilized by native linguists, because the written form of Chinese dialects has different characters, idioms, grammar and usage from Mandarin, and even the local speaker cannot input a correct sentence. For Mandarin text inputs, Chinese dialect TTS can only generate partly-meaningful speech with relatively poor prosody and naturalness. To lower the bar of use and make it more practical in commercial, we propose a novel Chinese dialect TTS frontend with a translation module. It helps to convert Mandarin text into idiomatic expressions with correct orthography and grammar, so that the intelligibility and naturalness of the synthesized speech can be improved. A non-autoregressive neural machine translation model with a glancing sampling strategy is proposed for the translation task. It is the first known work to incorporate translation with TTS frontend. Our experiments on Cantonese approve that the proposed frontend can help Cantonese TTS system achieve a 0.27 improvement in MOS with Mandarin inputs.
【7】 Motif Mining and Unsupervised Representation Learning for BirdCLEF 2022
标题:BirdCLEF 2022的Motif挖掘和无监督表示学习
链接:https://arxiv.org/abs/2206.04805
作者:Anthony Miyaguchi,Jiangyue Yu,Bryan Cheungvivatpant,Dakota Dudley,Aniketh Swain机构:Georgia Institute of Technology, North Ave NW, Atlanta, GA备注:Submitted to CEUR-WS under LifeCLEF for the BirdCLEF 2022 challenge as a working note摘要:我们使用无监督的方法为BirdCLEF 2022挑战建立了一个分类模型。我们使用音频基序的三重丢失谱图表示实现了训练数据集的无监督表示。我们的最佳模型在公共排行榜上的得分为0.48。摘要:We build a classification model for the BirdCLEF 2022 challenge using unsupervised methods. We implement an unsupervised representation of the training dataset using a triplet loss on spectrogram representation of audio motifs. Our best model performs with a score of 0.48 on the public leaderboard.
【8】 Speak Like a Dog: Human to Non-human creature Voice Conversion
标题:像狗一样说话:人类到非人类生物的声音转换
链接:https://arxiv.org/abs/2206.04780
作者:Kohei Suzuki,Shoki Sakamoto,Tadahiro Taniguchi,Hirokazu Kameoka机构:College of Information Science and Engineering, Ritsumeikan University, Japan, NTT Communication Science Laboratories, NTT Corporation, Japan摘要:以人到非人的语音转换(H2NH-VC)任务为例,提出了一种新的语音转换(VC)任务,在保留语言信息的同时,将人类语音转换为类狗语音。虽然大多数VC研究都涉及人与人之间的VC,但H2NH-VC旨在将人类的语音转换为类似于非人类生物的语音。非并行VC允许我们开发H2NH-VC,因为我们无法收集非人类生物讲人类语言的并行数据集。在这项研究中,我们建议将狗作为非人生物目标域的一个示例,并定义“像狗一样说话”任务。为了阐明“像狗一样说话”任务的可能性和特点,我们在声学特征(Mel倒谱系数和Mel频谱图)、网络结构(五种不同的内核大小设置)和,和训练标准(基于变分自动编码器(VAE)和基于生成对抗网络)。最后,使用平均意见分数评估转换后的声音:狗的相似性、音质和清晰度以及字符错误率(CER)。实验表明,Mel谱图的使用改善了转换语音的狗的相似性,但保留语言信息却很困难。重点介绍了H2NH-VC当前VC方法的挑战和局限性。摘要:This paper proposes a new voice conversion (VC) task from human speech to dog-like speech while preserving linguistic information as an example of human to non-human creature voice conversion (H2NH-VC) tasks. Although most VC studies deal with human to human VC, H2NH-VC aims to convert human speech into non-human creature-like speech. Non-parallel VC allows us to develop H2NH-VC, because we cannot collect a parallel dataset that non-human creatures speak human language. In this study, we propose to use dogs as an example of a non-human creature target domain and define the "speak like a dog" task. To clarify the possibilities and characteristics of the "speak like a dog" task, we conducted a comparative experiment using existing representative non-parallel VC methods in acoustic features (Mel-cepstral coefficients and Mel-spectrograms), network architectures (five different kernel-size settings), and training criteria (variational autoencoder (VAE)- based and generative adversarial network-based). Finally, the converted voices were evaluated using mean opinion scores: dog-likeness, sound quality and intelligibility, and character error rate (CER). The experiment showed that the employment of the Mel-spectrogram improved the dog-likeness of the converted speech, while it is challenging to preserve linguistic information. Challenges and limitations of the current VC methods for H2NH-VC are highlighted.
【9】 CLAP: Learning Audio Concepts From Natural Language Supervision
标题:CLAP:从自然语言监督中学习音频概念
链接:https://arxiv.org/abs/2206.04769
作者:Benjamin Elizalde,Soham Deshmukh,Mahmoud Al Ismail,Huaming Wang摘要:主流音频分析模型经过训练,可以在一个类别标签到多个录音的范式下学习,重点是一项任务。在这种受限的监督下学习限制了模型的灵活性,因为它们需要标记音频进行训练,并且只能预测预定义的类别。相反,我们建议从自然语言监控中学习音频概念。我们称之为对比语言音频预训练(CRAP),该方法通过使用两个编码器和对比学习将语言和音频连接起来,将音频和文本描述带入一个联合的多模态空间。我们使用128k音频和文本对训练CLAP,并在8个领域的16个下游任务中对其进行评估,如声音事件分类、音乐任务和语音相关任务。虽然与类似的计算机视觉模型相比,CLAP的训练成对数明显减少,但它为零炮性能建立了SoTA。此外,我们在监督学习环境中评估了CLAP,并在5项任务中实现了SoTA。因此,CLAP的零炮能力消除了使用类标签进行训练的需要,在推理时实现了灵活的类预测,并推广到多个下游任务。摘要:Mainstream Audio Analytics models are trained to learn under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the predefined categories. Instead, we propose to learn audio concepts from natural language supervision. We call our approach Contrastive Language-Audio Pretraining (CLAP), which learns to connect language and audio by using two encoders and a contrastive learning to bring audio and text descriptions into a joint multimodal space. We trained CLAP with 128k audio and text pairs and evaluated it on 16 downstream tasks across 8 domains, such as Sound Event Classification, Music tasks, and Speech-related tasks. Although CLAP was trained with significantly less pairs than similar computer vision models, it establishes SoTA for Zero-Shot performance. Additionally, we evaluated CLAP in a supervised learning setup and achieve SoTA in 5 tasks. Hence, CLAP's Zero-Shot capability removes the need of training with class labels, enables flexible class prediction at inference time, and generalizes to multiple downstream tasks.
机器翻译,仅供参考