今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理4篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily
cs.SD语音
【1】 LPCSE: Neural Speech Enhancement through Linear Predictive Coding

标题:LPCSE:基于线性预测编码的神经语音增强

链接:https://arxiv.org/abs/2206.06908

作者:Yang Liu,Na Tang,Xiaoli Chu,Yang Yang,Jun Wang
机构:The University of Sheffield, Sheffield, United Kingdom, University of Chinese Academy of Sciences, Shanghai, China, Terminus Group, China, University College London, London, United Kingdom
摘要:5G/B5G通信系统对体验质量的要求越来越严格,这导致了新兴的神经语音增强技术,但这些技术是在与现有基于专家规则的语音发音和失真模型分离的情况下开发的,例如经典的线性预测编码(LPC)语音模型,因为很难将这些模型与自动可微机器学习框架集成。在本文中,为了提高神经语音增强的效率,我们引入了一种基于LPC的语音增强(LPCSE)体系结构,该体系结构利用了LPC语音模型中的强归纳偏差以及神经网络的表达能力。LPCSE通过两个新块实现可微端到端学习:一个块在将LPC语音模型集成到神经网络时利用专家规则减少计算开销,另一个块通过将线性预测系数映射到滤波器极点来确保模型的稳定性并避免端到端训练中爆发梯度。实验结果表明,LPCSE成功地恢复了因传输损耗而失真的语音共振峰,并且在LJ语音语料库上的语音质量感知评价(PESQ)和短时目标可懂度(STOI)方面优于两种现有的神经网络规模相当的神经语音增强方法。
摘要:The increasingly stringent requirement on quality-of-experience in 5G/B5G communication systems has led to the emerging neural speech enhancement techniques, which however have been developed in isolation from the existing expert-rule based models of speech pronunciation and distortion, such as the classic Linear Predictive Coding (LPC) speech model because it is difficult to integrate the models with auto-differentiable machine learning frameworks. In this paper, to improve the efficiency of neural speech enhancement, we introduce an LPC-based speech enhancement (LPCSE) architecture, which leverages the strong inductive biases in the LPC speech model in conjunction with the expressive power of neural networks. Differentiable end-to-end learning is achieved in LPCSE via two novel blocks: a block that utilizes the expert rules to reduce the computational overhead when integrating the LPC speech model into neural networks, and a block that ensures the stability of the model and avoids exploding gradients in end-to-end training by mapping the Linear prediction coefficients to the filter poles. The experimental results show that LPCSE successfully restores the formants of the speeches distorted by transmission loss, and outperforms two existing neural speech enhancement methods of comparable neural network sizes in terms of the Perceptual evaluation of speech quality (PESQ) and Short-Time Objective Intelligibility (STOI) on the LJ Speech corpus.


【2】 Exploring speaker enrolment for few-shot personalisation in emotional  vocalisation prediction

标题:在情感发声预测中探索说话人注册的少数几次个性化

链接:https://arxiv.org/abs/2206.06680

作者:Andreas Triantafyllopoulos,Meishu Song,Zijiang Yang,Xin Jing,Björn W. Schuller
机构:University of Augsburg
备注:Proceedings of the ICML Expressive Vocalizations Workshop and Competition held in conjunction with the $\mathit{39}^{th}$ International Conference on Machine Learning, Copyright 2022 by the author(s)
摘要:在这项工作中,我们探索了一种新颖的用于情感发声预测的少数镜头个性化架构。核心贡献是一个“注册”编码器,它利用目标说话人的两个未标记样本来调整情感编码器的输出;调整基于dot产品关注度,因此有效地作为“软”特征选择的一种形式。情感和注册编码器基于两种标准音频架构:CNN14和CNN10。进一步引导这两个编码器忘记或学习辅助情绪和/或说话人信息。我们的最佳方法实现了美元的CCC。ExVo少数镜头开发集650美元,比我们的基线CNN14 CCC美元增加2.5\%。634$.
摘要:In this work, we explore a novel few-shot personalisation architecture for emotional vocalisation prediction. The core contribution is an `enrolment' encoder which utilises two unlabelled samples of the target speaker to adjust the output of the emotion encoder; the adjustment is based on dot-product attention, thus effectively functioning as a form of `soft' feature selection. The emotion and enrolment encoders are based on two standard audio architectures: CNN14 and CNN10. The two encoders are further guided to forget or learn auxiliary emotion and/or speaker information. Our best approach achieves a CCC of $.650$ on the ExVo Few-Shot dev set, a $2.5\%$ increase over our baseline CNN14 CCC of $.634$.


【3】 WHIS: Hearing impairment simulator based on the gammachirp auditory  filterbank

标题:WHIS:基于Gammachirp听觉滤波器组的听力障碍模拟器

链接:https://arxiv.org/abs/2206.06604

作者:Toshio Irino
机构:Wakayama University, Sakaedani Wakayama, -, Japan
备注:This paper was submitted to Trends in Hearing on Jun 5, 2022
摘要:基于gammachirp filterbank(GCFB)的修订版,实现了新版本的听力损伤模拟器(WHIS),其中包括基于帧的快速处理、绝对阈值(AT)、听力受损(HI)监听器的听力图和控制耳蜗输入输出(IO)功能的参数。被称为压缩健康$\ alpha$的参数控制IO功能的斜率,使其范围从正常听力(NH)听者到HI听者,而不会在很大程度上改变总听力损失(HL)。新的WHIS旨在为NH监听器提供与目标HI监听器相同的EPs。WHIS的分析部分与修改后的GCFB几乎相同,只是使用IO函数代替增益函数。我们提出了两种合成方法:用于感知小失真的直接时变滤波器和用于进一步HI模拟(包括时间涂抹)的滤波器组分析合成。我们评估了WHIS系列和剑桥版HL模拟器(CamHLS)在IO功能和光谱距离方面的差异。IO函数在$\ alpha$小于0.5时模拟得相当好,但在$\ alpha$等于1时模拟得不太好。因此,当IO功能足够健康时,很难模拟HL。这是任何现有HL模拟器以及WHIS的基本限制。新的WHIS产生了比CamHLS更小的光谱失真,并且与以前的版本相当兼容。
摘要:A new version of a hearing impairment simulator (WHIS) was implemented based on a revised version of the gammachirp filterbank (GCFB), which incorporates fast frame-based processing, absolute threshold (AT), an audiogram of a hearing-impaired (HI) listener, and a parameter to control the cochlear input-output (IO) function. The parameter referred to as the compression health $\alpha$ controlled the slope of the IO function to range from normal hearing (NH) listeners to HI listeners, without largely changing the total hearing loss (HL). The new WHIS was designed provide an NH listener the same EPs as those of a target HI listener.The analysis part of WHIS was almost the same as that of the revised GCFB, except that the IO function was used instead of the gain function. We proposed two synthesis methods: a direct time-varying filter for perceptually small distortion and a filterbank analysis-synthesis for further HI simulations including temporal smearing. We evaluated the WHIS family and a Cambridge version of the HL simulator (CamHLS) in terms of differences in the IO function and spectral distance. The IO functions were simulated fairly well at $\alpha$ less than 0.5 but not at $\alpha$ equal to 1. Thus, it is difficult to simulate the HL when the IO function is sufficiently healthy. This is a fundamental limit of any existing HL simulator as well as WHIS. The new WHIS yielded a smaller spectral distortion than CamHLS and was fairly compatible with the previous version.


【4】 Speech intelligibility of simulated hearing loss sounds and its  prediction using the Gammachirp Envelope Similarity Index (GESI)

标题:模拟听力损失的语音清晰度及其Gammachirp包络相似性指数预测

链接:https://arxiv.org/abs/2206.06573

作者:Toshio Irino,Honoka Tamaru,Ayako Yamamoto
机构:Wakayama University, Japan
备注:This paper was submitted to Interspeech 2022
摘要:在本研究中,在实验室和远程环境中使用模拟听力损失(HL)声音进行了语音清晰度(SI)实验,以阐明外周功能障碍的影响。使用Wadai听力障碍模拟器(WHIS),对有噪声的语音进行处理,以模拟70岁和80岁儿童的平均HL。这些声音被呈现给认知功能正常的正常听力(NH)听者。结果表明,远程实验的发散度大于实验室实验。然而,远程结果可以与实验室结果相等,主要是通过使用实验网页上准备的音调pip测试结果进行数据筛选。此外,一种新提出的被称为Gammachirp包络相似指数(GESI)的客观可懂度度量(OIM)很好地解释了实验室和远程实验中的心理测量功能。通过适当设置HL参数,GESI有可能解释HI监听器的SI。
摘要:In the present study, speech intelligibility (SI) experiments were performed using simulated hearing loss (HL) sounds in laboratory and remote environments to clarify the effects of peripheral dysfunction. Noisy speech sounds were processed to simulate the average HL of 70- and 80-year-olds using Wadai Hearing Impairment Simulator (WHIS). These sounds were presented to normal hearing (NH) listeners whose cognitive function could be assumed to be normal. The results showed that the divergence was larger in the remote experiments than in the laboratory ones. However, the remote results could be equalized to the laboratory ones, mostly through data screening using the results of tone pip tests prepared on the experimental web page. In addition, a newly proposed objective intelligibility measure (OIM) called the Gammachirp Envelope Similarity Index (GESI) explained the psychometric functions in the laboratory and remote experiments fairly well. GESI has the potential to explain the SI of HI listeners by properly setting HL parameters.


eess.AS音频处理

【1】 Adversarial Audio Synthesis with Complex-valued Polynomial Networks

标题:基于复值多项式网络的对抗性音频合成

链接:https://arxiv.org/abs/2206.06811

作者:Yongtao Wu,Grigorios G Chrysos,Volkan Cevher
机构:KTH Royal Institute of Technology, EPFL, Switzerland
备注:Accepted as oral presentation in Workshop on Machine Learning for Audio Synthesis at ICML 2022
摘要:音频合成中的时频(TF)表示越来越多地采用实值网络建模。然而,忽视TF表示的复杂值性质可能会导致次优性能,并需要额外的模块(例如,用于建模阶段)。为此,我们引入了复数多项式网络,称为APOLLO,它以一种自然的方式集成了这些复数表示。具体来说,阿波罗利用高阶张量作为标度参数捕捉输入元素的高阶相关性。通过利用标准的张量分解,我们导出了不同的体系结构,并能够对更丰富的相关性进行建模。我们概述了这些架构,并展示了它们在四个基准测试中的音频生成性能。作为亮点,阿波罗在音频生成方面比对抗性方法提高了17.5美元,比SC09数据集上最先进的扩散模型提高了8.2美元。我们的模型可以鼓励在复杂领域对其他高效体系结构进行系统化设计。
摘要:Time-frequency (TF) representations in audio synthesis have been increasingly modeled with real-valued networks. However, overlooking the complex-valued nature of TF representations can result in suboptimal performance and require additional modules (e.g., for modeling the phase). To this end, we introduce complex-valued polynomial networks, called APOLLO, that integrate such complex-valued representations in a natural way. Concretely, APOLLO captures high-order correlations of the input elements using high-order tensors as scaling parameters. By leveraging standard tensor decompositions, we derive different architectures and enable modeling richer correlations. We outline such architectures and showcase their performance in audio generation across four benchmarks. As a highlight, APOLLO results in $17.5\%$ improvement over adversarial methods and $8.2\%$ over the state-of-the-art diffusion models on SC09 dataset in audio generation. Our models can encourage the systematic design of other efficient architectures on the complex field.


【2】 LPCSE: Neural Speech Enhancement through Linear Predictive Coding

标题:LPCSE:基于线性预测编码的神经语音增强

链接:https://arxiv.org/abs/2206.06908

作者:Yang Liu,Na Tang,Xiaoli Chu,Yang Yang,Jun Wang
机构:The University of Sheffield, Sheffield, United Kingdom, University of Chinese Academy of Sciences, Shanghai, China, Terminus Group, China, University College London, London, United Kingdom
摘要:5G/B5G通信系统对体验质量的要求越来越严格,这导致了新兴的神经语音增强技术,但这些技术是在与现有基于专家规则的语音发音和失真模型分离的情况下开发的,例如经典的线性预测编码(LPC)语音模型,因为很难将这些模型与自动可微机器学习框架集成。在本文中,为了提高神经语音增强的效率,我们引入了一种基于LPC的语音增强(LPCSE)体系结构,该体系结构利用了LPC语音模型中的强归纳偏差以及神经网络的表达能力。LPCSE通过两个新块实现可微端到端学习:一个块在将LPC语音模型集成到神经网络时利用专家规则减少计算开销,另一个块通过将线性预测系数映射到滤波器极点来确保模型的稳定性并避免端到端训练中爆发梯度。实验结果表明,LPCSE成功地恢复了因传输损耗而失真的语音共振峰,并且在LJ语音语料库上的语音质量感知评价(PESQ)和短时目标可懂度(STOI)方面优于两种现有的神经网络规模相当的神经语音增强方法。
摘要:The increasingly stringent requirement on quality-of-experience in 5G/B5G communication systems has led to the emerging neural speech enhancement techniques, which however have been developed in isolation from the existing expert-rule based models of speech pronunciation and distortion, such as the classic Linear Predictive Coding (LPC) speech model because it is difficult to integrate the models with auto-differentiable machine learning frameworks. In this paper, to improve the efficiency of neural speech enhancement, we introduce an LPC-based speech enhancement (LPCSE) architecture, which leverages the strong inductive biases in the LPC speech model in conjunction with the expressive power of neural networks. Differentiable end-to-end learning is achieved in LPCSE via two novel blocks: a block that utilizes the expert rules to reduce the computational overhead when integrating the LPC speech model into neural networks, and a block that ensures the stability of the model and avoids exploding gradients in end-to-end training by mapping the Linear prediction coefficients to the filter poles. The experimental results show that LPCSE successfully restores the formants of the speeches distorted by transmission loss, and outperforms two existing neural speech enhancement methods of comparable neural network sizes in terms of the Perceptual evaluation of speech quality (PESQ) and Short-Time Objective Intelligibility (STOI) on the LJ Speech corpus.


【3】 WHIS: Hearing impairment simulator based on the gammachirp auditory  filterbank

标题:WHIS:基于Gammachirp听觉滤波器组的听力障碍模拟器

链接:https://arxiv.org/abs/2206.06604

作者:Toshio Irino
机构:Wakayama University, Sakaedani Wakayama, -, Japan
备注:This paper was submitted to Trends in Hearing on Jun 5, 2022
摘要:基于gammachirp filterbank(GCFB)的修订版,实现了新版本的听力损伤模拟器(WHIS),其中包括基于帧的快速处理、绝对阈值(AT)、听力受损(HI)监听器的听力图和控制耳蜗输入输出(IO)功能的参数。被称为压缩健康$\ alpha$的参数控制IO功能的斜率,使其范围从正常听力(NH)听者到HI听者,而不会在很大程度上改变总听力损失(HL)。新的WHIS旨在为NH监听器提供与目标HI监听器相同的EPs。WHIS的分析部分与修改后的GCFB几乎相同,只是使用IO函数代替增益函数。我们提出了两种合成方法:用于感知小失真的直接时变滤波器和用于进一步HI模拟(包括时间涂抹)的滤波器组分析合成。我们评估了WHIS系列和剑桥版HL模拟器(CamHLS)在IO功能和光谱距离方面的差异。IO函数在$\ alpha$小于0.5时模拟得相当好,但在$\ alpha$等于1时模拟得不太好。因此,当IO功能足够健康时,很难模拟HL。这是任何现有HL模拟器以及WHIS的基本限制。新的WHIS产生了比CamHLS更小的光谱失真,并且与以前的版本相当兼容。
摘要:A new version of a hearing impairment simulator (WHIS) was implemented based on a revised version of the gammachirp filterbank (GCFB), which incorporates fast frame-based processing, absolute threshold (AT), an audiogram of a hearing-impaired (HI) listener, and a parameter to control the cochlear input-output (IO) function. The parameter referred to as the compression health $\alpha$ controlled the slope of the IO function to range from normal hearing (NH) listeners to HI listeners, without largely changing the total hearing loss (HL). The new WHIS was designed provide an NH listener the same EPs as those of a target HI listener.The analysis part of WHIS was almost the same as that of the revised GCFB, except that the IO function was used instead of the gain function. We proposed two synthesis methods: a direct time-varying filter for perceptually small distortion and a filterbank analysis-synthesis for further HI simulations including temporal smearing. We evaluated the WHIS family and a Cambridge version of the HL simulator (CamHLS) in terms of differences in the IO function and spectral distance. The IO functions were simulated fairly well at $\alpha$ less than 0.5 but not at $\alpha$ equal to 1. Thus, it is difficult to simulate the HL when the IO function is sufficiently healthy. This is a fundamental limit of any existing HL simulator as well as WHIS. The new WHIS yielded a smaller spectral distortion than CamHLS and was fairly compatible with the previous version.


【4】 Speech intelligibility of simulated hearing loss sounds and its  prediction using the Gammachirp Envelope Similarity Index (GESI)

标题:模拟听力损失的语音清晰度及其Gammachirp包络相似性指数预测

链接:https://arxiv.org/abs/2206.06573

作者:Toshio Irino,Honoka Tamaru,Ayako Yamamoto
机构:Wakayama University, Japan
备注:This paper was submitted to Interspeech 2022
摘要:在本研究中,在实验室和远程环境中使用模拟听力损失(HL)声音进行了语音清晰度(SI)实验,以阐明外周功能障碍的影响。使用Wadai听力障碍模拟器(WHIS),对有噪声的语音进行处理,以模拟70岁和80岁儿童的平均HL。这些声音被呈现给认知功能正常的正常听力(NH)听者。结果表明,远程实验的发散度大于实验室实验。然而,远程结果可以与实验室结果相等,主要是通过使用实验网页上准备的音调pip测试结果进行数据筛选。此外,一种新提出的被称为Gammachirp包络相似指数(GESI)的客观可懂度度量(OIM)很好地解释了实验室和远程实验中的心理测量功能。通过适当设置HL参数,GESI有可能解释HI监听器的SI。
摘要:In the present study, speech intelligibility (SI) experiments were performed using simulated hearing loss (HL) sounds in laboratory and remote environments to clarify the effects of peripheral dysfunction. Noisy speech sounds were processed to simulate the average HL of 70- and 80-year-olds using Wadai Hearing Impairment Simulator (WHIS). These sounds were presented to normal hearing (NH) listeners whose cognitive function could be assumed to be normal. The results showed that the divergence was larger in the remote experiments than in the laboratory ones. However, the remote results could be equalized to the laboratory ones, mostly through data screening using the results of tone pip tests prepared on the experimental web page. In addition, a newly proposed objective intelligibility measure (OIM) called the Gammachirp Envelope Similarity Index (GESI) explained the psychometric functions in the laboratory and remote experiments fairly well. GESI has the potential to explain the SI of HI listeners by properly setting HL parameters.


机器翻译,仅供参考