今天跟大家分享一篇语音相关的论文合集:cs.SD语音5篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily
cs.SD语音

【1】 Predicting non-native speech perception using the Perceptual  Assimilation Model and state-of-the-art acoustic models

标题:使用知觉同化模型和最新声学模型预测非母语语音感知
链接:https://arxiv.org/abs/2205.15823
作者:Juliette Millet,Ioana Chitoran,Ewan Dunbar
机构:LLF, University of Paris, CNRS, CRI, FAN, IIFR, University of Paris, Department of Linguistics, University of Toronto, Toronto, Canada, CoML, ENSCNRSEHESSINRIAPSL, Paris, France
备注:None
摘要:我们的母语影响我们感知语音的方式,影响我们辨别非母语声音的能力。我们比较了两种关于母语对语音感知影响的观点:一种是感知同化模型,它呼吁在心理上将声音分类为母语音素类别,另一种观点是根据母语的统计数据调整丰富、细粒度的语音表征就够了。我们使用两种最先进的语音模型(Dirichlet过程高斯混合模型和最近的wav2vec 2.0模型)的表示来实现这一想法。我们针对来自六种语言的61个元音,提供了一个新的、开放的法语和英语参与者言语感知行为数据集。我们表明,音位同化比细粒度语音建模更能预测整体的辨别行为,也能预测与母语背景差异相关的辨别性差异。我们还表明,wav2vec 2.0虽然不善于捕捉母语对语音感知的影响,但它补充了有关母语音素同化的信息,并提供了一个良好的低水平语音表征模型,支持在语音感知过程中同时使用范畴感知和细粒度感知的观点。
摘要:Our native language influences the way we perceive speech sounds, affecting our ability to discriminate non-native sounds. We compare two ideas about the influence of the native language on speech perception: the Perceptual Assimilation Model, which appeals to a mental classification of sounds into native phoneme categories, versus the idea that rich, fine-grained phonetic representations tuned to the statistics of the native language, are sufficient. We operationalize this idea using representations from two state-of-the-art speech models, a Dirichlet process Gaussian mixture model and the more recent wav2vec 2.0 model. We present a new, open dataset of French- and English-speaking participants' speech perception behaviour for 61 vowel sounds from six languages. We show that phoneme assimilation is a better predictor than fine-grained phonetic modelling, both for the discrimination behaviour as a whole, and for predicting differences in discriminability associated with differences in native language background. We also show that wav2vec 2.0, while not good at capturing the effects of native language on speech perception, is complementary to information about native phoneme assimilation, and provides a good model of low-level phonetic representations, supporting the idea that both categorical and fine-grained perception are used during speech perception.


【2】 Do self-supervised speech models develop human-like perception biases?

标题:自我监督的语音模型会产生类似人类的感知偏差吗?

链接:https://arxiv.org/abs/2205.15819

作者:Juliette Millet,Ewan Dunbar
机构:CoML, ENSCNRSEHESSINRIAPSL, LLF, University of Paris, CNRS, CRI, FAN, IIFR, University of Paris, Paris, France, University of Toronto, Toronto, Canada
备注:None
摘要:语音处理的自监督模型形成表征空间,无需使用任何外部标签。它们似乎越来越成为一种可行的方法,至少可以部分消除昂贵的手动注释,这是低资源语言特别关注的问题。但是这些模型构建了什么样的表征空间呢?人类的感知专门研究听者母语的声音。在自我监督模型中是否也会发生同样的情况?我们研究了三种最先进的自我监督模型:wav2vec 2.0、HuBERT和对比预测编码(CPC)的表征空间,并将其与法语和英语人类听众的感知空间进行了比较,无论是在全球范围内,还是考虑到两种语言群体之间的行为差异。我们发现,CPC模型显示出很小的母语效应,但wav2vec 2.0和HuBERT似乎开发出了一个通用的语音感知空间,它不是特定于语言的。与有监督的电话识别器的预测进行比较表明,所有三个自监督模型都能捕捉到相对细粒度的感知现象,而有监督模型则更能捕捉到听者母语对感知的粗糙、电话水平的影响。
摘要:Self-supervised models for speech processing form representational spaces without using any external labels. Increasingly, they appear to be a feasible way of at least partially eliminating costly manual annotations, a problem of particular concern for low-resource languages. But what kind of representational spaces do these models construct? Human perception specializes to the sounds of listeners' native languages. Does the same thing happen in self-supervised models? We examine the representational spaces of three kinds of state-of-the-art self-supervised models: wav2vec 2.0, HuBERT and contrastive predictive coding (CPC), and compare them with the perceptual spaces of French-speaking and English-speaking human listeners, both globally and taking account of the behavioural differences between the two language groups. We show that the CPC model shows a small native language effect, but that wav2vec 2.0 and HuBERT seem to develop a universal speech perception space which is not language specific. A comparison against the predictions of supervised phone recognisers suggests that all three self-supervised models capture relatively fine-grained perceptual phenomena, while supervised models are better at capturing coarser, phone-level, effects of listeners' native language, on perception.


【3】 Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech  with Untranscribed Data

标题:Guided-TTS 2:一种高质量非转录数据自适应文语转换扩散模型

链接:https://arxiv.org/abs/2205.15370

作者:Sungwon Kim,Heeseung Kim,Sungroh Yoon
机构:Data Science & AI Lab., Seoul National University
摘要:我们提出了引导TTS 2,这是一种基于扩散的生成模型,用于使用未翻译数据的高质量自适应TTS。Guided TTS 2将说话人条件扩散模型与说话人相关音素分类器相结合,实现自适应文语转换。我们在大规模未翻译数据集上训练说话人条件扩散模型,以获得无分类器的引导方法,并进一步微调目标说话人参考语音上的扩散模型以进行自适应,只需40秒。我们证明,引导TTS 2在语音质量和说话人相似度方面的性能与高质量的单说话人TTS基线相当,只有10秒的未翻译数据。我们进一步表明,即使在零炮自适应设置下,引导TTS 2在多说话人数据集上的性能也优于自适应TTS基线。引导TTS 2只能使用未翻译的语音来适应范围广泛的语音,这使得自适应TTS能够适应非人类角色的语音,例如《指环王》中的咕噜。
摘要:We propose Guided-TTS 2, a diffusion-based generative model for high-quality adaptive TTS using untranscribed data. Guided-TTS 2 combines a speaker-conditional diffusion model with a speaker-dependent phoneme classifier for adaptive text-to-speech. We train the speaker-conditional diffusion model on large-scale untranscribed datasets for a classifier-free guidance method and further fine-tune the diffusion model on the reference speech of the target speaker for adaptation, which only takes 40 seconds. We demonstrate that Guided-TTS 2 shows comparable performance to high-quality single-speaker TTS baselines in terms of speech quality and speaker similarity with only a ten-second untranscribed data. We further show that Guided-TTS 2 outperforms adaptive TTS baselines on multi-speaker datasets even with a zero-shot adaptation setting. Guided-TTS 2 can adapt to a wide range of voices only using untranscribed speech, which enables adaptive TTS with the voice of non-human characters such as Gollum in \textit{"The Lord of the Rings"}.


【4】 Revisiting Audio Pattern Recognition for Asthma Medication Adherence:  Evaluation with the RDA Benchmark Suite

标题:用RDA Benchmark Suite评估音频模式识别对哮喘用药依从性的影响

链接:https://arxiv.org/abs/2205.15360

作者:Nikos D. Fakotakis,Stavros Nousias,Gerasimos Arvanitis,Evangelia I. Zacharaki,Konstantinos Moustakas
机构:∗Department of Electrical and Computer Engineering, University of Patras, Greece
摘要:哮喘是一种常见的、通常是长期的呼吸系统疾病,对全世界的社会和经济都有负面影响。治疗包括使用医疗设备(吸入器)将药物分配到呼吸道,其效率取决于吸入技术的精度。配备传感器并嵌入声音信号检测的健康监测系统能够识别药物致动,并可能成为可靠音频内容分析的有力工具。本文回顾了用于哮喘药物依从性评估的音频模式识别和机器学习技术,并介绍了呼吸和药物驱动(RDA)套件(https://gitlab.com/vvr/monitoring-medication-adherence/rda-benchmark)用于基准测试和进一步研究。RDA套件包括一套用于音频处理、特征提取和分类的工具,并随附由呼吸和药物驱动声音组成的数据集。RDA中的分类模型是基于传统和先进的机器学习以及深度网络体系结构实现的。本研究对已实施的方法进行了比较评估,检查了潜在的改进,并讨论了挑战和未来趋势。
摘要:Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide. Treatment involves using medical devices (inhalers) that distribute medication to the airways, and its efficiency depends on the precision of the inhalation technique. Health monitoring systems equipped with sensors and embedded with sound signal detection enable the recognition of drug actuation and could be powerful tools for reliable audio content analysis. This paper revisits audio pattern recognition and machine learning techniques for asthma medication adherence assessment and presents the Respiratory and Drug Actuation (RDA) Suite(https://gitlab.com/vvr/monitoring-medication-adherence/rda-benchmark) for benchmarking and further research. The RDA Suite includes a set of tools for audio processing, feature extraction and classification and is provided along with a dataset consisting of respiratory and drug actuation sounds. The classification models in RDA are implemented based on conventional and advanced machine learning and deep network architectures. This study provides a comparative evaluation of the implemented approaches, examines potential improvements and discusses challenges and future tendencies.


【5】 StyleTTS: A Style-Based Generative Model for Natural and Diverse  Text-to-Speech Synthesis

标题:StyleTTS:一种基于风格的自然多样文语合成生成模型

链接:https://arxiv.org/abs/2205.15439

作者:Yinghao Aaron Li,Cong Han,Nima Mesgarani
机构:Columbia University
摘要:由于并行TTS系统的快速发展,文本到语音(TTS)最近在合成高质量语音方面取得了巨大进展,但生成具有自然韵律变化、说话风格和情感音调的语音仍然具有挑战性。此外,由于持续时间和语音是分别生成的,并行TTS模型仍然难以找到对自然语音合成至关重要的最佳单调对齐。在这里,我们提出了StyleTTS,这是一种基于风格的平行TTS生成模型,可以从参考语音中合成具有自然韵律的多样化语音。通过新的可转移单调对齐器(TMA)和持续时间不变的数据增强方案,我们的方法在语音自然度和说话人相似性的主观测试中,在单说话人和多说话人数据集上都显著优于最新的模型。通过对说话风格的自我监督学习,我们的模型可以合成与任何给定参考语音具有相同韵律和情感基调的语音,而无需明确标记这些类别。
摘要:Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturalistic prosodic variations, speaking styles and emotional tones remains challenging. Moreover, since duration and speech are generated separately, parallel TTS models still have problems finding the best monotonic alignments that are crucial for naturalistic speech synthesis. Here, we propose StyleTTS, a style-based generative model for parallel TTS that can synthesize diverse speech with natural prosody from a reference speech utterance. With novel Transferable Monotonic Aligner (TMA) and duration-invariant data augmentation schemes, our method significantly outperforms state-of-the-art models on both single and multi-speaker datasets in subjective tests of speech naturalness and speaker similarity. Through self-supervised learning of the speaking styles, our model can synthesize speech with the same prosodic and emotional tone as any given reference speech without the need for explicitly labeling these categories.


eess.AS音频处理

【1】 Adversarial synthesis based data-augmentation for code-switched spoken  language identification

标题:基于对抗性合成的语码转换口语识别数据增强

链接:https://arxiv.org/abs/2205.15747

作者:Parth Shastri,Chirag Patil,Poorval Wanere,Dr. Shrinivas Mahajan,Dr. Abhishek Bhatt,Dr. Hardik Sailor
机构:Dept. of Electronics and telecommunincation, College of Engineering, Pune., ∥Samsung Research and Development, Bangalore.
备注:9 pages, 8 figures
摘要:口语识别(LID)是自动语音识别(ASR)的一个重要子任务,用于对音频段中的语言进行分类。自动LID在多语种国家发挥着有益的作用。在不同的国家,识别一种语言变得很困难,因为在多语言场景中,两种或两种以上的语言在对话中混合在一起。这种语音现象被称为语码混合或语码转换。这种性质不仅在印度,在许多亚洲国家也是如此。这样的代码混合数据很难找到,这进一步降低了语音LID的能力。由于这种代码混合数据缺乏可用性,它成为LID任务中的少数类。因此,这项工作主要解决这个问题,使用数据扩充作为少数代码交换类的解决方案。本研究主要研究印度语代码与英语的混合。口语LID使用印地语,代码与英语混合。本研究提出了基于生成对抗网络(GAN)的数据增强技术,该技术使用Mel频谱图对音频数据进行增强。GANs已经被证明能够准确地表示图像域中的真实数据分布。拟议的研究利用了GANs在语音领域的这些功能,如语音分类、自动语音识别等。GANs经过训练,生成少数代码混合类的Mel谱图,然后用于为分类器增加数据。与作为基线参考的卷积递归神经网络(CRNN)分类器相比,使用GANs可将未加权平均召回率整体提高3.5%。
摘要:Spoken Language Identification (LID) is an important sub-task of Automatic Speech Recognition(ASR) that is used to classify the language(s) in an audio segment. Automatic LID plays an useful role in multilingual countries. In various countries, identifying a language becomes hard, due to the multilingual scenario where two or more than two languages are mixed together during conversation. Such phenomenon of speech is called as code-mixing or code-switching. This nature is followed not only in India but also in many Asian countries. Such code-mixed data is hard to find, which further reduces the capabilities of the spoken LID. Due to the lack of avalibility of this code-mixed data, it becomes a minority class in LID task. Hence, this work primarily addresses this problem using data augmentation as a solution on the minority code-switched class. This study focuses on Indic language code-mixed with English. Spoken LID is performed on Hindi, code-mixed with English. This research proposes Generative Adversarial Network (GAN) based data augmentation technique performed using Mel spectrograms for audio data. GANs have already been proven to be accurate in representing the real data distribution in the image domain. Proposed research exploits these capabilities of GANs in speech domains such as speech classification, automatic speech recognition,etc. GANs are trained to generate Mel spectrograms of the minority code-mixed class which are then used to augment data for the classifier. Utilizing GANs give an overall improvement on Unweighted Average Recall by an amount of 3.5\% as compared to a Convolutional Recurrent Neural Network (CRNN) classifier used as the baseline reference.



【2】 Conversational Speech Separation: an Evaluation Study for Streaming  Applications

标题:会话语音分离:流媒体应用的评估研究

链接:https://arxiv.org/abs/2205.15700

作者:Giovanni Morrone,Samuele Cornell,Enrico Zovato,Alessio Brutti,Stefano Squartini
机构:Department of Information Engineering, Ancona, Italy, PerVoice S.p.A., Trento, Italy, Fondazione Bruno Kessler, Trento, Italy
备注:Audio Engineering Society Convention 152, May 2022, The Hague, Netherlands
摘要:连续语音分离(CSS)是最近提出的一种框架,旨在以流式方式将每个说话人从输入混合信号中分离出来。此后,我们将对CSS系统的实际设计考虑进行评估研究,以解决在最近的工作中被忽视的重要方面。我们特别关注分离性能、计算要求和输出延迟之间的权衡,展示了如何使用离线分离算法以期望的延迟执行CSS。我们对稀疏重叠数据上CSS处理窗口大小和跃点大小的选择进行了广泛的分析。我们发现,在计算量和性能之间的最佳权衡是在5s的窗口内实现的。
摘要:Continuous speech separation (CSS) is a recently proposed framework which aims at separating each speaker from an input mixture signal in a streaming fashion. Hereafter we perform an evaluation study on practical design considerations for a CSS system, addressing important aspects which have been neglected in recent works. In particular, we focus on the trade-off between separation performance, computational requirements and output latency showing how an offline separation algorithm can be used to perform CSS with a desired latency. We carry out an extensive analysis on the choice of CSS processing window size and hop size on sparsely overlapped data. We find out that the best trade-off between computational burden and performance is obtained for a window of 5 s.



【3】 StyleTTS: A Style-Based Generative Model for Natural and Diverse  Text-to-Speech Synthesis

标题:StyleTTS:一种基于风格的自然多样文语合成生成模型

链接:https://arxiv.org/abs/2205.15439

作者:Yinghao Aaron Li,Cong Han,Nima Mesgarani
机构:Columbia University
摘要:由于并行TTS系统的快速发展,文本到语音(TTS)最近在合成高质量语音方面取得了巨大进展,但生成具有自然韵律变化、说话风格和情感音调的语音仍然具有挑战性。此外,由于持续时间和语音是分别生成的,并行TTS模型仍然难以找到对自然语音合成至关重要的最佳单调对齐。在这里,我们提出了StyleTTS,这是一种基于风格的平行TTS生成模型,可以从参考语音中合成具有自然韵律的多样化语音。通过新的可转移单调对齐器(TMA)和持续时间不变的数据增强方案,我们的方法在语音自然度和说话人相似性的主观测试中,在单说话人和多说话人数据集上都显著优于最新的模型。通过对说话风格的自我监督学习,我们的模型可以合成与任何给定参考语音具有相同韵律和情感基调的语音,而无需明确标记这些类别。
摘要:Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturalistic prosodic variations, speaking styles and emotional tones remains challenging. Moreover, since duration and speech are generated separately, parallel TTS models still have problems finding the best monotonic alignments that are crucial for naturalistic speech synthesis. Here, we propose StyleTTS, a style-based generative model for parallel TTS that can synthesize diverse speech with natural prosody from a reference speech utterance. With novel Transferable Monotonic Aligner (TMA) and duration-invariant data augmentation schemes, our method significantly outperforms state-of-the-art models on both single and multi-speaker datasets in subjective tests of speech naturalness and speaker similarity. Through self-supervised learning of the speaking styles, our model can synthesize speech with the same prosodic and emotional tone as any given reference speech without the need for explicitly labeling these categories.


【4】 Predicting non-native speech perception using the Perceptual  Assimilation Model and state-of-the-art acoustic models

标题:使用知觉同化模型和最新声学模型预测非母语语音感知

链接:https://arxiv.org/abs/2205.15823

作者:Juliette Millet,Ioana Chitoran,Ewan Dunbar
机构:LLF, University of Paris, CNRS, CRI, FAN, IIFR, University of Paris, Department of Linguistics, University of Toronto, Toronto, Canada, CoML, ENSCNRSEHESSINRIAPSL, Paris, France
备注:None
摘要:我们的母语影响我们感知语音的方式,影响我们辨别非母语声音的能力。我们比较了两种关于母语对语音感知影响的观点:一种是感知同化模型,它呼吁在心理上将声音分类为母语音素类别,另一种观点是根据母语的统计数据调整丰富、细粒度的语音表征就够了。我们使用两种最先进的语音模型(Dirichlet过程高斯混合模型和最近的wav2vec 2.0模型)的表示来实现这一想法。我们针对来自六种语言的61个元音,提供了一个新的、开放的法语和英语参与者言语感知行为数据集。我们表明,音位同化比细粒度语音建模更能预测整体的辨别行为,也能预测与母语背景差异相关的辨别性差异。我们还表明,wav2vec 2.0虽然不善于捕捉母语对语音感知的影响,但它补充了有关母语音素同化的信息,并提供了一个良好的低水平语音表征模型,支持在语音感知过程中同时使用范畴感知和细粒度感知的观点。
摘要:Our native language influences the way we perceive speech sounds, affecting our ability to discriminate non-native sounds. We compare two ideas about the influence of the native language on speech perception: the Perceptual Assimilation Model, which appeals to a mental classification of sounds into native phoneme categories, versus the idea that rich, fine-grained phonetic representations tuned to the statistics of the native language, are sufficient. We operationalize this idea using representations from two state-of-the-art speech models, a Dirichlet process Gaussian mixture model and the more recent wav2vec 2.0 model. We present a new, open dataset of French- and English-speaking participants' speech perception behaviour for 61 vowel sounds from six languages. We show that phoneme assimilation is a better predictor than fine-grained phonetic modelling, both for the discrimination behaviour as a whole, and for predicting differences in discriminability associated with differences in native language background. We also show that wav2vec 2.0, while not good at capturing the effects of native language on speech perception, is complementary to information about native phoneme assimilation, and provides a good model of low-level phonetic representations, supporting the idea that both categorical and fine-grained perception are used during speech perception.


【5】 Do self-supervised speech models develop human-like perception biases?

标题:自我监督的语音模型会产生类似人类的感知偏差吗?

链接:https://arxiv.org/abs/2205.15819

作者:Juliette Millet,Ewan Dunbar
机构:CoML, ENSCNRSEHESSINRIAPSL, LLF, University of Paris, CNRS, CRI, FAN, IIFR, University of Paris, Paris, France, University of Toronto, Toronto, Canada
备注:None
摘要:语音处理的自监督模型形成表征空间,无需使用任何外部标签。它们似乎越来越成为一种可行的方法,至少可以部分消除昂贵的手动注释,这是低资源语言特别关注的问题。但是这些模型构建了什么样的表征空间呢?人类的感知专门研究听者母语的声音。在自我监督模型中是否也会发生同样的情况?我们研究了三种最先进的自我监督模型:wav2vec 2.0、HuBERT和对比预测编码(CPC)的表征空间,并将其与法语和英语人类听众的感知空间进行了比较,无论是在全球范围内,还是考虑到两种语言群体之间的行为差异。我们发现,CPC模型显示出很小的母语效应,但wav2vec 2.0和HuBERT似乎开发出了一个通用的语音感知空间,它不是特定于语言的。与有监督的电话识别器的预测进行比较表明,所有三个自监督模型都能捕捉到相对细粒度的感知现象,而有监督模型则更能捕捉到听者母语对感知的粗糙、电话水平的影响。
摘要:Self-supervised models for speech processing form representational spaces without using any external labels. Increasingly, they appear to be a feasible way of at least partially eliminating costly manual annotations, a problem of particular concern for low-resource languages. But what kind of representational spaces do these models construct? Human perception specializes to the sounds of listeners' native languages. Does the same thing happen in self-supervised models? We examine the representational spaces of three kinds of state-of-the-art self-supervised models: wav2vec 2.0, HuBERT and contrastive predictive coding (CPC), and compare them with the perceptual spaces of French-speaking and English-speaking human listeners, both globally and taking account of the behavioural differences between the two language groups. We show that the CPC model shows a small native language effect, but that wav2vec 2.0 and HuBERT seem to develop a universal speech perception space which is not language specific. A comparison against the predictions of supervised phone recognisers suggests that all three self-supervised models capture relatively fine-grained perceptual phenomena, while supervised models are better at capturing coarser, phone-level, effects of listeners' native language, on perception.


【6】 Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech  with Untranscribed Data

标题:Guided-TTS 2:一种高质量非转录数据自适应文语转换扩散模型

链接:https://arxiv.org/abs/2205.15370

作者:Sungwon Kim,Heeseung Kim,Sungroh Yoon
机构:Data Science & AI Lab., Seoul National University
摘要:我们提出了引导TTS 2,这是一种基于扩散的生成模型,用于使用未翻译数据的高质量自适应TTS。Guided TTS 2将说话人条件扩散模型与说话人相关音素分类器相结合,实现自适应文语转换。我们在大规模未翻译数据集上训练说话人条件扩散模型,以获得无分类器的引导方法,并进一步微调目标说话人参考语音上的扩散模型以进行自适应,只需40秒。我们证明,引导TTS 2在语音质量和说话人相似度方面的性能与高质量的单说话人TTS基线相当,只有10秒的未翻译数据。我们进一步表明,即使在零炮自适应设置下,引导TTS 2在多说话人数据集上的性能也优于自适应TTS基线。引导TTS 2只能使用未翻译的语音来适应范围广泛的语音,这使得自适应TTS能够适应非人类角色的语音,例如《指环王》中的咕噜。
摘要:We propose Guided-TTS 2, a diffusion-based generative model for high-quality adaptive TTS using untranscribed data. Guided-TTS 2 combines a speaker-conditional diffusion model with a speaker-dependent phoneme classifier for adaptive text-to-speech. We train the speaker-conditional diffusion model on large-scale untranscribed datasets for a classifier-free guidance method and further fine-tune the diffusion model on the reference speech of the target speaker for adaptation, which only takes 40 seconds. We demonstrate that Guided-TTS 2 shows comparable performance to high-quality single-speaker TTS baselines in terms of speech quality and speaker similarity with only a ten-second untranscribed data. We further show that Guided-TTS 2 outperforms adaptive TTS baselines on multi-speaker datasets even with a zero-shot adaptation setting. Guided-TTS 2 can adapt to a wide range of voices only using untranscribed speech, which enables adaptive TTS with the voice of non-human characters such as Gollum in \textit{"The Lord of the Rings"}.


【7】 Revisiting Audio Pattern Recognition for Asthma Medication Adherence:  Evaluation with the RDA Benchmark Suite
标题:用RDA Benchmark Suite评估音频模式识别对哮喘用药依从性的影响
链接:https://arxiv.org/abs/2205.15360
作者:Nikos D. Fakotakis,Stavros Nousias,Gerasimos Arvanitis,Evangelia I. Zacharaki,Konstantinos Moustakas
机构:∗Department of Electrical and Computer Engineering, University of Patras, Greece
摘要:哮喘是一种常见的、通常是长期的呼吸系统疾病,对全世界的社会和经济都有负面影响。治疗包括使用医疗设备(吸入器)将药物分配到呼吸道,其效率取决于吸入技术的精度。配备传感器并嵌入声音信号检测的健康监测系统能够识别药物致动,并可能成为可靠音频内容分析的有力工具。本文回顾了用于哮喘药物依从性评估的音频模式识别和机器学习技术,并介绍了呼吸和药物驱动(RDA)套件(https://gitlab.com/vvr/monitoring-medication-adherence/rda-benchmark)用于基准测试和进一步研究。RDA套件包括一套用于音频处理、特征提取和分类的工具,并随附由呼吸和药物驱动声音组成的数据集。RDA中的分类模型是基于传统和先进的机器学习以及深度网络体系结构实现的。本研究对已实施的方法进行了比较评估,检查了潜在的改进,并讨论了挑战和未来趋势。
摘要:Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide. Treatment involves using medical devices (inhalers) that distribute medication to the airways, and its efficiency depends on the precision of the inhalation technique. Health monitoring systems equipped with sensors and embedded with sound signal detection enable the recognition of drug actuation and could be powerful tools for reliable audio content analysis. This paper revisits audio pattern recognition and machine learning techniques for asthma medication adherence assessment and presents the Respiratory and Drug Actuation (RDA) Suite(https://gitlab.com/vvr/monitoring-medication-adherence/rda-benchmark) for benchmarking and further research. The RDA Suite includes a set of tools for audio processing, feature extraction and classification and is provided along with a dataset consisting of respiratory and drug actuation sounds. The classification models in RDA are implemented based on conventional and advanced machine learning and deep network architectures. This study provides a comparative evaluation of the implemented approaches, examines potential improvements and discusses challenges and future tendencies.


机器翻译,仅供参考