今天跟大家分享一篇语音相关的论文合集:cs.SD语音5篇,eess.AS音频处理6篇。

cs.SD语音

【1】 Machine-learning applied to classify flow-induced sound parameters from  simulated human voice

标题:基于机器学习的流声参数分类方法

链接:https://arxiv.org/abs/2207.09265

作者:Florian Kraxberger,Andreas Wurzinger,Stefan Schoder
机构:Institute of Fundamentals and Theory in Electrical Engineering (IGTE), Graz University of Technology,  Graz, Austria
备注:17 pages, 11 figures, v0.1, work in progress, working paper
摘要:发声障碍严重影响患者的生活质量。本文采用模拟方法研究了发声过程中的因果链,表现出典型的声音特征,如声门下压力和功能性发声障碍,如声门闭合不全和左右不对称。据此,利用已发表的混合气动声学仿真模型,对24种不同的声结构进行了参数研究,在此基础上,选择了声学参数(HNR,CPP,...)与这些仿真配置细节相关联,以基于仿真结果得出对人发声的流诱导声音生成的特征认识。最近,一些机构研究了流和声学特性的实验数据,并将其与健康和紊乱的语音信号相关联。基于此,该研究是进一步对数据集进行详细定义的一个步骤,数据集较小,但相关特征的定义是精确的,基于现有的simVoice仿真方法,对小数据集进行相关性分析,并使用具有RBF核的支持向量机分类器对这些表示进行分类。通过线性判别分析,可以直观地显示各个研究的维度。这可以绘制相关性,并根据嘴前的声音信号确定最重要的评估特征。基于CPP和箱形图可视化,可以最好地区分GC类型。此外,使用LDA降维特征空间,可以以91.7%的准确率对声门下压力进行最佳分类,而与健康或紊乱的语音模拟参数无关。
摘要:Disorders of voice production have severe effects on the quality of life of the affected individuals. A simulation approach is used to investigate the cause-effect chain in voice production showing typical characteristics of voice such as sub-glottal pressure and of functional voice disorders as glottal closure insufficiency and left-right asymmetry. Therewith, 24 different voice configurations are simulated in a parameter study using a previously published hybrid aeroacoustic simulation model. Based on these 24 simulation configurations, selected acoustic parameters (HNR, CPP, ...) at simulation evaluation points are correlated with these simulation configuration details to derive characteristic insight in the flow-induced sound generation of human phonation based on simulation results. Recently, several institutions studied experimental data, of flow and acoustic properties and correlated it with healthy and disordered voice signals. Upon this, the study is a next step towards a detailed dataset definition, the dataset is small, but the definition of relevant characteristics are precise based on the existing simulation methodology of simVoice. The small datasets are studied by correlation analysis, and a Support Vector Machine classifier with RBF kernel is used to classify the representations. With the use of Linear Discriminant Analysis the dimensions of the individual studies are visualized. This allows to draw correlations and determine the most important features evaluated from the acoustic signals in front of the mouth. The GC type can be best discriminated based on CPP and boxplot visualizations. Furthermore and using the LDA-dimensionality-reduced feature space, one can best classify subglottal pressure with 91.7\% accuracy, independent of healthy or disordered voice simulation parameters.


【2】 Realistic sources, receivers and walls improve the generalisability of  virtually-supervised blind acoustic parameter estimators

标题:真实的声源、接收器和墙壁提高了虚拟监督盲声学参数估计器的通用性

链接:https://arxiv.org/abs/2207.09133

作者:Prerak Srivastava,Antoine Deleforge,Emmanuel Vincent
机构:Universit´e de Lorraine, CNRS, Inria, Loria, F-, Nancy, France
摘要:声学参数盲估计是指从未知声源的记录中推断出环境的声学特性.由于实际测量数据的有限性,最近的研究工作使用了部分或完全在模拟数据上训练的深度神经网络.本文提出了一种基于深度神经网络的声学参数盲估计算法,我们研究了使用快速图像源房间脉冲响应模拟器纯粹训练的模型是否可以推广到真实数据。我们提出了一个精心制作的模拟训练集的消融研究,接收器和壁响应。真实性的程度由墙壁吸收系数的采样和将测量的方向性模式应用于麦克风和声源来控制。对在这些数据集上训练的art模型进行了评估,以联合估计房间的体积、总表面积和来自多个多通道语音记录的倍频带混响时间。结果表明,在训练时每增加一层仿真真实性,都显著提高了对真实信号的所有量的估计。
摘要:Blind acoustic parameter estimation consists in inferring the acoustic properties of an environment from recordings of unknown sound sources. Recent works in this area have utilized deep neural networks trained either partially or exclusively on simulated data, due to the limited availability of real annotated measurements. In this paper, we study whether a model purely trained using a fast image-source room impulse response simulator can generalize to real data. We present an ablation study on carefully crafted simulated training sets that account for different levels of realism in source, receiver and wall responses. The extent of realism is controlled by the sampling of wall absorption coefficients and by applying measured directivity patterns to microphones and sources. A state-of-the-art model trained on these datasets is evaluated on the task of jointly estimating the room's volume, total surface area, and octave-band reverberation times from multiple, multichannel speech recordings. Results reveal that every added layer of simulation realism at train time significantly improves the estimation of all quantities on real signals.


【3】 Contrastive Environmental Sound Representation Learning

标题:对比环境声音表征学习

链接:https://arxiv.org/abs/2207.08825

作者:Peter Ochieng,Dennis Kaburu
机构:Department of Computer Science and Technology, University of Cambridge, Department of Information Technology, Jomo Kenyatta University of Agriculture and Technology
摘要:环境声音的机器听觉是音频识别领域的一个重要研究课题,它赋予机器对不同输入声音进行区分的能力,从而指导其决策.本文利用自监督对比技术和浅层一维CNN提取不同的音频特征我们使用给定音频的原始音频波形和声谱图两者来生成其表示,并且评估所提出的学习者是否对音频输入的类型是不可知的。我们进一步使用典型相关分析在ESC-50和UrbanSound8K上对所提出的技术进行了评估,结果表明,与单独的音频表示相比,融合后的全局特征能够更好地表示音频信号。实验结果表明,该方法能够提取环境音频的大部分特征,在ESC-50和UrbanSound8K数据集上的性能分别提高了12.8%和0.9%.
摘要:Machine hearing of the environmental sound is one of the important issues in the audio recognition domain. It gives the machine the ability to discriminate between the different input sounds that guides its decision making. In this work we exploit the self-supervised contrastive technique and a shallow 1D CNN to extract the distinctive audio features (audio representations) without using any explicit annotations.We generate representations of a given audio using both its raw audio waveform and spectrogram and evaluate if the proposed learner is agnostic to the type of audio input. We further use canonical correlation analysis (CCA) to fuse representations from the two types of input of a given audio and demonstrate that the fused global feature results in robust representation of the audio signal as compared to the individual representations. The evaluation of the proposed technique is done on both ESC-50 and UrbanSound8K. The results show that the proposed technique is able to extract most features of the environmental audio and gives an improvement of 12.8% and 0.9% on the ESC-50 and UrbanSound8K datasets respectively.


【4】 Audio Input Generates Continuous Frames to Synthesize Facial Video Using  Generative Adiversarial Networks

标题:音频输入生成连续帧以使用生成式对偶网络合成面部视频

链接:https://arxiv.org/abs/2207.08813

作者:Hanhaodi Zhang
机构:School of Information Technology and Electrical Engineering The University of Queensland
备注:5 pages, 5 figures
摘要:本文提出了一种基于音频的语音视频生成的简单方法:给定一段音频,我们可以生成一个目标人脸说这段音频的视频。(GAN)以截断语音音频输入为条件并使用卷积门循环单元(GRU)在发生器和鉴别器中,我们的模型通过利用短音频和在这个持续时间内的帧来训练,为了训练,我们设计了一个简单的编码器,并比较了使用GAN和不使用GRU产生的帧,我们使用GRU产生时间相关的帧,结果表明短音频可以产生相对真实的输出结果。
摘要:This paper presents a simple method for speech videos generation based on audio: given a piece of audio, we can generate a video of the target face speaking this audio. We propose Generative Adversarial Networks (GAN) with cut speech audio input as condition and use Convolutional Gate Recurrent Unit (GRU) in generator and discriminator. Our model is trained by exploiting the short audio and the frames in this duration. For training, we cut the audio and extract the face in the corresponding frames. We designed a simple encoder and compare the generated frames using GAN with and without GRU. We use GRU for temporally coherent frames and the results show that short audio can produce relatively realistic output results.


【5】 GAFX: A General Audio Feature eXtractor

标题:GAFX:一种通用的音频特征提取器

链接:https://arxiv.org/abs/2207.09145

作者:Zhaoyang Bu,Hanhaodi Zhang,Xiaohu Zhu
机构:Center for Safe AGI, Ohla AI Group
摘要:大多数音频任务的机器学习模型处理的是一个人工特征——声谱图。然而,基于深度学习的特征是否能取代声谱图仍然是未知数。本文通过比较不同的可学习神经网络提取特征与成功的声谱图模型,回答了这个问题,并提出了一个通用音频特征提取器(GAFX)的双U网(GAFX-U),资源网络(GAFX-R)和注意事项我们设计实验在GTZAN数据集上的音乐体裁分类任务上评估该模型,并对我们的框架和我们的模型GAFX-U的不同配置进行详细的消融研究,跟随音频频谱图Transformer(AST)分类器实现了有竞争力的性能。
摘要:Most machine learning models for audio tasks are dealing with a handcrafted feature, the spectrogram. However, it is still unknown whether the spectrogram could be replaced with deep learning based features. In this paper, we answer this question by comparing the different learnable neural networks extracting features with a successful spectrogram model and proposed a General Audio Feature eXtractor (GAFX) based on a dual U-Net (GAFX-U), ResNet (GAFX-R), and Attention (GAFX-A) modules. We design experiments to evaluate this model on the music genre classification task on the GTZAN dataset and perform a detailed ablation study of different configurations of our framework and our model GAFX-U, following the Audio Spectrogram Transformer (AST) classifier achieves competitive performance.


eess.AS音频处理

【1】 Do uHear? Validation of uHear App for Preliminary Screening of Hearing  Ability in Soundscape Studies

标题:在声景研究中,uHear App用于听力初步筛查的确认

链接:https://arxiv.org/abs/2207.09221

作者:Zhen-Ting Ong,Bhan Lam,Kenneth Ooi,Karn N. Watcharasupat,Trevor Wong,Woon-Seng Gan
机构:School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
备注:Full paper submitted to 24th International Congress on Acoustics
摘要:涉及声景感知的研究通常会排除听力损失的参与者,以防止感知受损影响实验结果。参与者通常会接受纯音测听(识别和量化特定频率听力损失的“黄金标准”)筛选,如果不符合研究相关阈值,则会被排除。然而,为声景研究采购专业测听设备可能成本效益不高。手动进行听力测试是一项劳动密集型的工作。此外,声景研究的测试要求可能不需要像医疗诊断环境中那样高的灵敏度和特异性。因此,在本研究中,我们调查了uHear应用程序(一款iOS应用程序)的有效性。作为常规听力计的负担得起的和自动的替代物,在为了声景研究或一般的听力测试的目的而筛选参与者的听力损失中。根据与163名参与者的听力计进行的听力测试比较,uHear应用程序被发现具有较高的精确度(98.04%)使用世界卫生组织时(WHO)听力正常评估分级方案。精确度进一步提高(98.69%),这进一步支持了这种具有成本效益的自动化替代方案,以筛查正常听力。
摘要:Studies involving soundscape perception often exclude participants with hearing loss to prevent impaired perception from affecting experimental results. Participants are typically screened with pure tone audiometry, the "gold standard" for identifying and quantifying hearing loss at specific frequencies, and excluded if a study-dependent threshold is not met. However, procuring professional audiometric equipment for soundscape studies may be cost-ineffective, and manually performing audiometric tests is labour-intensive. Moreover, testing requirements for soundscape studies may not require sensitivities and specificities as high as that in a medical diagnosis setting. Hence, in this study, we investigate the effectiveness of the uHear app, an iOS application, as an affordable and automatic alternative to a conventional audiometer in screening participants for hearing loss for the purpose of soundscape studies or listening tests in general. Based on audiometric comparisons with the audiometer of 163 participants, the uHear app was found to have high precision (98.04%) when using the World Health Organization (WHO) grading scheme for assessing normal hearing. Precision is further improved (98.69%) when all frequencies assessed with the uHear app is considered in the grading, which lends further support to this cost-effective, automated alternative to screen for normal hearing.


【2】 GAFX: A General Audio Feature eXtractor

标题:GAFX:一种通用的音频特征提取器

链接:https://arxiv.org/abs/2207.09145

* 与cs.SD语音【5】为同一篇

作者:Zhaoyang Bu,Hanhaodi Zhang,Xiaohu Zhu
机构:Center for Safe AGI, Ohla AI Group
摘要:大多数音频任务的机器学习模型处理的是一个人工特征——声谱图。然而,基于深度学习的特征是否能取代声谱图仍然是未知数。本文通过比较不同的可学习神经网络提取特征与成功的声谱图模型,回答了这个问题,并提出了一个通用音频特征提取器(GAFX)的双U网(GAFX-U),资源网络(GAFX-R)和注意事项我们设计实验在GTZAN数据集上的音乐体裁分类任务上评估该模型,并对我们的框架和我们的模型GAFX-U的不同配置进行详细的消融研究,跟随音频频谱图Transformer(AST)分类器实现了有竞争力的性能。
摘要:Most machine learning models for audio tasks are dealing with a handcrafted feature, the spectrogram. However, it is still unknown whether the spectrogram could be replaced with deep learning based features. In this paper, we answer this question by comparing the different learnable neural networks extracting features with a successful spectrogram model and proposed a General Audio Feature eXtractor (GAFX) based on a dual U-Net (GAFX-U), ResNet (GAFX-R), and Attention (GAFX-A) modules. We design experiments to evaluate this model on the music genre classification task on the GTZAN dataset and perform a detailed ablation study of different configurations of our framework and our model GAFX-U, following the Audio Spectrogram Transformer (AST) classifier achieves competitive performance.


【3】 Machine-learning applied to classify flow-induced sound parameters from  simulated human voice

标题:基于机器学习的流声参数分类方法

链接:https://arxiv.org/abs/2207.09265

* 与cs.SD语音【1】为同一篇

作者:Florian Kraxberger,Andreas Wurzinger,Stefan Schoder
机构:Institute of Fundamentals and Theory in Electrical Engineering (IGTE), Graz University of Technology,  Graz, Austria
备注:17 pages, 11 figures, v0.1, work in progress, working paper
摘要:发声障碍严重影响患者的生活质量。本文采用模拟方法研究了发声过程中的因果链,表现出典型的声音特征,如声门下压力和功能性发声障碍,如声门闭合不全和左右不对称。据此,利用已发表的混合气动声学仿真模型,对24种不同的声结构进行了参数研究,在此基础上,选择了声学参数(HNR,CPP,...)与这些仿真配置细节相关联,以基于仿真结果得出对人发声的流诱导声音生成的特征认识。最近,一些机构研究了流和声学特性的实验数据,并将其与健康和紊乱的语音信号相关联。基于此,该研究是进一步对数据集进行详细定义的一个步骤,数据集较小,但相关特征的定义是精确的,基于现有的simVoice仿真方法,对小数据集进行相关性分析,并使用具有RBF核的支持向量机分类器对这些表示进行分类。通过线性判别分析,可以直观地显示各个研究的维度。这可以绘制相关性,并根据嘴前的声音信号确定最重要的评估特征。基于CPP和箱形图可视化,可以最好地区分GC类型。此外,使用LDA降维特征空间,可以以91.7%的准确率对声门下压力进行最佳分类,而与健康或紊乱的语音模拟参数无关。
摘要:Disorders of voice production have severe effects on the quality of life of the affected individuals. A simulation approach is used to investigate the cause-effect chain in voice production showing typical characteristics of voice such as sub-glottal pressure and of functional voice disorders as glottal closure insufficiency and left-right asymmetry. Therewith, 24 different voice configurations are simulated in a parameter study using a previously published hybrid aeroacoustic simulation model. Based on these 24 simulation configurations, selected acoustic parameters (HNR, CPP, ...) at simulation evaluation points are correlated with these simulation configuration details to derive characteristic insight in the flow-induced sound generation of human phonation based on simulation results. Recently, several institutions studied experimental data, of flow and acoustic properties and correlated it with healthy and disordered voice signals. Upon this, the study is a next step towards a detailed dataset definition, the dataset is small, but the definition of relevant characteristics are precise based on the existing simulation methodology of simVoice. The small datasets are studied by correlation analysis, and a Support Vector Machine classifier with RBF kernel is used to classify the representations. With the use of Linear Discriminant Analysis the dimensions of the individual studies are visualized. This allows to draw correlations and determine the most important features evaluated from the acoustic signals in front of the mouth. The GC type can be best discriminated based on CPP and boxplot visualizations. Furthermore and using the LDA-dimensionality-reduced feature space, one can best classify subglottal pressure with 91.7\% accuracy, independent of healthy or disordered voice simulation parameters.


【4】 Realistic sources, receivers and walls improve the generalisability of  virtually-supervised blind acoustic parameter estimators

标题:真实的声源、接收器和墙壁提高了虚拟监督盲声学参数估计器的通用性

链接:https://arxiv.org/abs/2207.09133

* 与cs.SD语音【2】为同一篇

作者:Prerak Srivastava,Antoine Deleforge,Emmanuel Vincent
机构:Universit´e de Lorraine, CNRS, Inria, Loria, F-, Nancy, France
摘要:声学参数盲估计是指从未知声源的记录中推断出环境的声学特性.由于实际测量数据的有限性,最近的研究工作使用了部分或完全在模拟数据上训练的深度神经网络.本文提出了一种基于深度神经网络的声学参数盲估计算法,我们研究了使用快速图像源房间脉冲响应模拟器纯粹训练的模型是否可以推广到真实数据。我们提出了一个精心制作的模拟训练集的消融研究,接收器和壁响应。真实性的程度由墙壁吸收系数的采样和将测量的方向性模式应用于麦克风和声源来控制。对在这些数据集上训练的art模型进行了评估,以联合估计房间的体积、总表面积和来自多个多通道语音记录的倍频带混响时间。结果表明,在训练时每增加一层仿真真实性,都显著提高了对真实信号的所有量的估计。
摘要:Blind acoustic parameter estimation consists in inferring the acoustic properties of an environment from recordings of unknown sound sources. Recent works in this area have utilized deep neural networks trained either partially or exclusively on simulated data, due to the limited availability of real annotated measurements. In this paper, we study whether a model purely trained using a fast image-source room impulse response simulator can generalize to real data. We present an ablation study on carefully crafted simulated training sets that account for different levels of realism in source, receiver and wall responses. The extent of realism is controlled by the sampling of wall absorption coefficients and by applying measured directivity patterns to microphones and sources. A state-of-the-art model trained on these datasets is evaluated on the task of jointly estimating the room's volume, total surface area, and octave-band reverberation times from multiple, multichannel speech recordings. Results reveal that every added layer of simulation realism at train time significantly improves the estimation of all quantities on real signals.


【5】 Contrastive Environmental Sound Representation Learning

标题:对比环境声音表征学习

链接:https://arxiv.org/abs/2207.08825

* 与cs.SD语音【3】为同一篇

作者:Peter Ochieng,Dennis Kaburu
机构:Department of Computer Science and Technology, University of Cambridge, Department of Information Technology, Jomo Kenyatta University of Agriculture and Technology
摘要:环境声音的机器听觉是音频识别领域的一个重要研究课题,它赋予机器对不同输入声音进行区分的能力,从而指导其决策.本文利用自监督对比技术和浅层一维CNN提取不同的音频特征我们使用给定音频的原始音频波形和声谱图两者来生成其表示,并且评估所提出的学习者是否对音频输入的类型是不可知的。我们进一步使用典型相关分析在ESC-50和UrbanSound8K上对所提出的技术进行了评估,结果表明,与单独的音频表示相比,融合后的全局特征能够更好地表示音频信号。实验结果表明,该方法能够提取环境音频的大部分特征,在ESC-50和UrbanSound8K数据集上的性能分别提高了12.8%和0.9%.
摘要:Machine hearing of the environmental sound is one of the important issues in the audio recognition domain. It gives the machine the ability to discriminate between the different input sounds that guides its decision making. In this work we exploit the self-supervised contrastive technique and a shallow 1D CNN to extract the distinctive audio features (audio representations) without using any explicit annotations.We generate representations of a given audio using both its raw audio waveform and spectrogram and evaluate if the proposed learner is agnostic to the type of audio input. We further use canonical correlation analysis (CCA) to fuse representations from the two types of input of a given audio and demonstrate that the fused global feature results in robust representation of the audio signal as compared to the individual representations. The evaluation of the proposed technique is done on both ESC-50 and UrbanSound8K. The results show that the proposed technique is able to extract most features of the environmental audio and gives an improvement of 12.8% and 0.9% on the ESC-50 and UrbanSound8K datasets respectively.


【6】 Audio Input Generates Continuous Frames to Synthesize Facial Video Using  Generative Adiversarial Networks

标题:音频输入生成连续帧以使用生成式对偶网络合成面部视频

链接:https://arxiv.org/abs/2207.08813

* 与cs.SD语音【4】为同一篇

作者:Hanhaodi Zhang
机构:School of Information Technology and Electrical Engineering The University of Queensland
备注:5 pages, 5 figures
摘要:本文提出了一种基于音频的语音视频生成的简单方法:给定一段音频,我们可以生成一个目标人脸说这段音频的视频。(GAN)以截断语音音频输入为条件并使用卷积门循环单元(GRU)在发生器和鉴别器中,我们的模型通过利用短音频和在这个持续时间内的帧来训练,为了训练,我们设计了一个简单的编码器,并比较了使用GAN和不使用GRU产生的帧,我们使用GRU产生时间相关的帧,结果表明短音频可以产生相对真实的输出结果。
摘要:This paper presents a simple method for speech videos generation based on audio: given a piece of audio, we can generate a video of the target face speaking this audio. We propose Generative Adversarial Networks (GAN) with cut speech audio input as condition and use Convolutional Gate Recurrent Unit (GRU) in generator and discriminator. Our model is trained by exploiting the short audio and the frames in this duration. For training, we cut the audio and extract the face in the corresponding frames. We designed a simple encoder and compare the generated frames using GAN with and without GRU. We use GRU for temporally coherent frames and the results show that short audio can produce relatively realistic output results.


机器翻译,仅供参考